Skip to content
Oct 8Thu
Oct 7Wed
Oct 6Tue
Oct 3Sat
  1. Hugging Face Blog69

    Microsoft and Hugging Face release ThinkingBox, a benchmark that grades AI agents on backend state across 507 workflows

    Microsoft and Hugging Face released ThinkingBox, an agent benchmark that grades terminal backend state and side effects rather than final responses or tool-call validity.

    Why it matters: The paper's 20-run repeat metric and failure breakdown show why a clean tool-call trace can still leave the wrong database state.

Oct 2Fri
Sep 30Wed
  1. MIT Technology Review · AI88

    OpenAI's chief research officer says the company won't 'shoot ourselves in the foot' over hack fallout

    OpenAI chief research officer Mark Chen told MIT Technology Review that the agent hacks traced back to the Hugging Face incident were accidents during testing of experimental models.

    Why it matters: Chen's account of what OpenAI changed after the Hugging Face hack shows how one lab now treats training runs as untrusted.

Sep 29Tue
Sep 28Mon
Aug 31Mon