Topics

AI Evaluation & Reliability

Measuring, testing, and trusting AI systems.

AI evaluation is the practice of measuring whether an AI system actually works — through evals, benchmarks, and reliability testing that catch hallucinations and regressions before they reach production.

35 episodes

Explainers on this topic

Terms on this topic

Guests on this topic

Kristin LovejoyTyler AkidauLoïc HoussierAlex RatnerSudhir HasbeDan KleinFergal ReidVikram ChatterjiMaxime LabonneAishwarya SrinivasanMalte UblAurimas GriciūnasHamel HusainMikiko ChandrasekharPhilipp KrennGiovanna CarofiglioJoão MouraGreg StattonAtindriyo SanyalWade ChambersDenny LeeSiva SurendiraOlga BeregovayaRodrigo CoutinhoMaryam AshooriRachelle PalmerAndrew ZiglerLogan KilpatrickYash ShethVinnie GiarrussoMehmet Murat EzbiderliGrant LedfordChip HuyenVivienne ZhangSara HookerCraig WileyMay Habib