AI Evaluation & Reliability
Measuring, testing, and trusting AI systems.
AI evaluation is the practice of measuring whether an AI system actually works — through evals, benchmarks, and reliability testing that catch hallucinations and regressions before they reach production.
35 episodes
- The First Fully Autonomous AI Attack Is 18 Months Away | Kristin Lovejoy
- We Built Agents, Nobody Built HR | Tyler Akidau, Redpanda
- How Superhuman Built AI Into a 100ms Product | Loïc Houssier
- Every AI Agent Has an Evaluation Gap | Alex Ratner, Snorkel AI
- Hallucinations Are a Data Architecture Problem | Sudhir Hasbe, Neo4j
- Why LLMs Are Plausibility Engines, Not Truth Engines | Dan Klein
- How Intercom Cut $250K/Month by Ditching GPT for Qwen
- Explaining Eval Engineering | Galileo's Vikram Chatterji
- Beyond Transformers: How Liquid AI Is Rethinking LLM Architecture | Maxime Labonne
- Architecting AI Agents: The Shift from Models to Systems | Aishwarya Srinivasan
- Vercel's Playbook for AI Agents: From Vibe Check to Production | Malte Ubl
- From Demo to Defensibility: How to Build an AI Business that Lasts | Aurimas Griciūnas
- Mindset Over Metrics: How to Approach AI Engineering | Hamel Husain
- Mastering Multi-Agent Systems | MongoDB’s Mikiko Chandrasekhar
- The AI Agent Trust Gap: Bridging Risk to Reliability | Elastic’s Philipp Krenn
- Architecting Reliable Agentic AI | Cisco’s Giovanna Carofiglio on the AGNTCY Collective
- The Emerging AI Agent Stack | CrewAI’s João Moura
- Your Key to AI Success is Hiding in Plain Sight | Cohesity's Greg Statton
- The 2025 AI Shift: From Chat to Task Completion & Reliable Action | Galileo Founders
- Amplitude's AI Playbook: How Wade Chambers Builds for the Agentic Future
- AI's Two Extremes – Foundations & The Frontier | Databricks’ Denny Lee
- Why Enterprises Need a Different Approach to AI Agents | Lyzr’s Siva Surendira
- Breaking the Language Barrier: Smartling's AI Translation Pipeline | Olga Beregovaya
- Low-Code AI: From Requirements to Apps in Minutes | OutSystems' Rodrigo Coutinho
- Inside IBM's watsonx: Building Enterprise AI That Ships | Dr. Maryam Ashoori
- Using AI to Modernize Your Legacy Applications | MongoDB’s Rachelle Palmer
- AI in 2025: Agents & The Rise of Evaluation-Driven Development
- The Making of Gemini 2.0: DeepMind's Approach to AI Development and Deployment | Logan Kilpatrick
- How DeepSeek Changed the AI Race Overnight
- AI in 2025: Agents & The Rise of Evaluation Driven Development
- Beyond Chatbots: How Twilio Uses AI to Strengthen Human Connection | Vinnie Giarrusso
- The Enterprise AI Deployment Playbook | ServiceTitan, Indeed & Twilio
- Practical Lessons for GenAI Evals | Chip Huyen & Vivienne Zhang
- GenAI Predictions for 2025 | Databricks & Cohere
- The State of AI: Open-Source Models & Enterprise Trust | May Habib
Explainers on this topic
- How to Measure Whether AI Is Actually Paying Off
- The 3 Levels of Evaluating an AI Agent
- When AI Hallucinations Are Good (and When They're Dangerous)
- LLM-as-a-Judge vs. Human Evaluation: When to Use Which
- AI Agent Metrics Beyond Accuracy
- Are AI Hallucinations a Data Problem?
- How AI Agents Use Tools
- How to Evaluate a RAG System
- How to Test an AI System
- Observability vs. Evaluation vs. Benchmarking
- What Is AI Observability
- What Is LLM-as-a-Judge
- Why Enterprise AI Projects Fail to Show ROI
- Why Multi-Agent Systems Fail
- When AI Hallucinations Are Good (and When They're Dangerous)
- The 3 Levels of Evaluating an AI Agent
Terms on this topic
- Accuracy
- AI Benchmark
- AI Evaluation
- AI Hallucination
- AI Red Teaming
- Answer Relevance
- AUC-ROC
- BERTScore
- BLEU Score
- Cohen's Kappa
- Context Relevance
- Data Poisoning
- Explainability
- F1 Score
- Faithfulness
- Human in the Loop
- Instruction Adherence
- Latency
- LLM as a Judge
- Mean Reciprocal Rank (MRR)
- METEOR
- Model Drift
- Perplexity
- Precision and Recall
- ROUGE
- Synthetic Data
- Word Error Rate (WER)
Guests on this topic
Kristin LovejoyTyler AkidauLoïc HoussierAlex RatnerSudhir HasbeDan KleinFergal ReidVikram ChatterjiMaxime LabonneAishwarya SrinivasanMalte UblAurimas GriciūnasHamel HusainMikiko ChandrasekharPhilipp KrennGiovanna CarofiglioJoão MouraGreg StattonAtindriyo SanyalWade ChambersDenny LeeSiva SurendiraOlga BeregovayaRodrigo CoutinhoMaryam AshooriRachelle PalmerAndrew ZiglerLogan KilpatrickYash ShethVinnie GiarrussoMehmet Murat EzbiderliGrant LedfordChip HuyenVivienne ZhangSara HookerCraig WileyMay Habib