Multimodal AI
Models that see, hear, and read at once.
Multimodal AI describes models that take in and reason across more than one kind of data — text, images, audio, and video — in a single system, aligning those modalities so the model can answer about an image, transcribe and act on speech, or ground a response in a chart.
6 episodes
- Every AI Agent Has an Evaluation Gap | Alex Ratner, Snorkel AI
- Why LLMs Are Plausibility Engines, Not Truth Engines | Dan Klein
- Gemini 3 & Robot Dogs: Inside Google DeepMind's AI Experiments | Paige Bailey
- Breaking the Language Barrier: Smartling's AI Translation Pipeline | Olga Beregovaya
- The Making of Gemini 2.0: DeepMind's Approach to AI Development and Deployment | Logan Kilpatrick
- Practical Lessons for GenAI Evals | Chip Huyen & Vivienne Zhang
Explainer on this topic
Term on this topic
Guests on this topic
Alex RatnerDan KleinPaige BaileyOlga BeregovayaLogan KilpatrickChip HuyenVivienne Zhang