What is…
Multimodal
A model that handles more than text — images, audio, video or files — as input, output, or both.
Early chatbots were text-only. Multimodal models can read a screenshot, a chart, a photo of a whiteboard or a PDF; listen and talk back in real time; and in some cases generate images or video.
Practically, this means you can snap a photo of an error message, a receipt or a broken appliance and just ask.
📌 Example
Uploading a photo of your fridge and asking "what can I make for dinner?"
🧠 Test yourself
Which of these describes Multimodal?
Related terms
🎮 Learn AI by playing
40 bite-size missions, boss battles and a certificate. Free.
Start the bootcamp →☀️ One AI term every morning
Plus the day's top stories, in your inbox by 8am.