Skip to content

Content Type

Benchmarks Latest News

Which model is actually stronger: ongoing coverage of benchmark scores, evaluation-methodology disputes, and leaderboard shifts.

3 selectedLast 30 days: 2Total indexed: 8

Updated

Benchmarks picks

Oct 5Mon1–3
Oct 3Sat
  1. Hugging Face Blog69

    Microsoft and Hugging Face release ThinkingBox, a benchmark that grades AI agents on backend state across 507 workflows

    Microsoft and Hugging Face released ThinkingBox, an agent benchmark that grades terminal backend state and side effects rather than final responses or tool-call validity.

    Why it matters: The paper's 20-run repeat metric and failure breakdown show why a clean tool-call trace can still leave the wrong database state.

Aug 20Thu