Skip to content

What is…

Multimodal

A model that handles more than text — images, audio, video or files — as input, output, or both.

Early chatbots were text-only. Multimodal models can read a screenshot, a chart, a photo of a whiteboard or a PDF; listen and talk back in real time; and in some cases generate images or video.

Practically, this means you can snap a photo of an error message, a receipt or a broken appliance and just ask.

📌 Example

Uploading a photo of your fridge and asking "what can I make for dinner?"

🧠 Test yourself

Which of these describes Multimodal?

Related terms

🎮 Learn AI by playing

40 bite-size missions, boss battles and a certificate. Free.

Start the bootcamp →

☀️ One AI term every morning

Plus the day's top stories, in your inbox by 8am.