Multimodal
Models that process and generate multiple types of data (text, images, audio, video) within a single architecture instead of using separate specialized models for each modality.
Multimodal models accept inputs beyond text: images, audio clips, video frames, and documents. Rather than using separate specialized models for each modality, multimodal architectures process them in a unified way, allowing the model to reason across modalities simultaneously.
Input modalities in 2025-2026
- Text: Universal baseline for all frontier models
- Images: GPT-4o, Claude 3.5 Sonnet, Gemini 2.0 Pro, Llama 3.2, Grok-2, Qwen-VL - all support image understanding
- Video: Gemini 2.0 Pro, GPT-4o (limited frame rates), Claude 3.5 Sonnet - analyze video frames in sequence
- Audio: GPT-4o, Gemini 2.0 Flash - real-time audio understanding and generation
- Documents (PDF/layout): Most frontier models parse document structure via vision, though OCR reliability varies
How vision works in LLMs
Images are typically processed by a vision encoder (like ViT) that converts image patches into embeddings. These embeddings are projected into the LLM's token embedding space and concatenated with text tokens. The LLM then processes text and image tokens together using the same attention mechanism.
Limitations
Spatial reasoning (determining which object is to the left of another) remains weak in most models. Fine-grained OCR in complex layouts is unreliable. Video understanding is limited to relatively short clips and lower frame rates compared to specialized video models. Audio understanding in non-English languages lags behind text-only performance significantly. Most multimodal models process video by sampling frames rather than analyzing continuous temporal information. Image resolution handling also remains a bottleneck: most models compress or resize high-resolution images, losing fine details needed for tasks like reading small text in screenshots.
Related terms
Models relevant to Multimodal
Gemini 2.5 Pro
Google's advanced thinking model for complex reasoning, coding, and long context.
View model →GPT-4o
OpenAI's versatile, fast multimodal workhorse (text + image)
View model →GPT-5
OpenAI's landmark August 2025 flagship: strong reasoning at a low price
View model →Amazon Nova Pro
Amazon's balanced multimodal Bedrock model for text, image, and video at scale.
View model →Llama 4
Meta's natively multimodal open MoE herd with industry-leading context length.
View model →