Microsoft Research on AI failures
What you need to know
Microsoft Research's Jennifer Neville says evaluation should push AI systems beyond traditional benchmarks, and that "surprising failures" emerge when models are tested against real user needs. She offers practical guidance for working with current AI systems and says examining data matters when results defy expectations.
Generated by AI from the reporting · updated 3 days ago
Timeline
Follow the reports to see every side of the story.
- Microsoft ResearchWhat AI gets wrong and what failure teaches us
Microsoft Research's Jennifer Neville discusses how evaluation pushes AI systems beyond traditional benchmarks and why "surprising failures" emerge when models are tested on real user needs. She offers practical guidance for working with current AI systems and explains why examining data matters when results defy expectations.