AI that processes more than one type of input — text, images, audio, video, documents — and can produce output in multiple formats. Leading multimodal models include the AI, GPT-4o, and Gemini. Enables document intelligence, image analysis, audio transcription, and other tasks that go beyond text.
Multimodal AI
Related terms
This definition is part of a free, structured course on how AI actually works.
Start learning