Skip to content

What is…

Latency

How long you wait for a response — time to first token plus generation speed.

Latency is the delay between hitting Enter and seeing a useful answer. Bigger models and reasoning modes are slower; smaller or distilled models are faster. For voice assistants and autocomplete, latency matters more than raw smarts.

Picking the smallest model that does the job well is one of the most practical skills in applied AI.

🧠 Test yourself

Which of these describes Latency?

Related terms

🎮 Learn AI by playing

40 bite-size missions, boss battles and a certificate. Free.

Start the bootcamp →

☀️ One AI term every morning

Plus the day's top stories, in your inbox by 8am.