Skip to content

What is…

Pre-training

The first, most expensive phase: teaching a model language by having it predict the next token across trillions of tokens.

Pre-training is where a model absorbs grammar, facts, reasoning patterns and coding skills from an enormous dataset. It takes thousands of GPUs for weeks or months. The result — a "base model" — is knowledgeable but not yet a helpful assistant; it just continues text.

Everything after this (fine-tuning, RLHF) shapes that raw knowledge into a model that follows instructions and behaves well.

💡 Think of it like

Reading an entire library before anyone teaches you manners or how to answer questions.

🧠 Test yourself

Which of these describes Pre-training?

Related terms

🎮 Learn AI by playing

40 bite-size missions, boss battles and a certificate. Free.

Start the bootcamp →

☀️ One AI term every morning

Plus the day's top stories, in your inbox by 8am.