Skip to content
Oct 9Fri
Oct 7Wed
  1. GitHub Blog · AI & ML62

    GitHub extends push protection to unstructured secrets with a ModernBERT classifier built with Microsoft Applied Sciences

    GitHub built a fine-tuned ModernBERT classifier with Microsoft Applied Sciences that assesses candidate secrets in context in under two milliseconds.

    Why it matters: GitHub's nine quarters of push data and the latency budget behind its new secret classifier show how prevention is being moved into the push path.

Oct 6Tue
  1. Microsoft Research17

    What AI gets wrong and what failure teaches us

    Microsoft Research's Jennifer Neville discusses how evaluation pushes AI systems beyond traditional benchmarks and why "surprising failures" emerge when models are tested on real user needs. She offers practical guidance for working with current AI systems and explains why examining data matters when results defy expectations.

Oct 5Mon
Oct 3Sat
  1. Hugging Face Blog69

    Microsoft and Hugging Face release ThinkingBox, a benchmark that grades AI agents on backend state across 507 workflows

    Microsoft and Hugging Face released ThinkingBox, an agent benchmark that grades terminal backend state and side effects rather than final responses or tool-call validity.

    Why it matters: The paper's 20-run repeat metric and failure breakdown show why a clean tool-call trace can still leave the wrong database state.

Sep 30Wed
Sep 29Tue
Sep 28Mon
Sep 25Fri
  1. GitHub Blog · AI & ML32

    GitHub Copilot app for Beginners: How to build custom workflows with canvases

    GitHub Copilot app's canvases are customizable, bidirectional interfaces you create by running the /create-canvas skill and describing the workflow in plain English. The agent builds the UI in the right-side panel, and both you and the agent can update its shared state at the same time. Canvases are saved as extensions for reuse or team sharing, and ready-made ones are available via Awesome Copilot.

Sep 24Thu
Sep 23Wed
Sep 21Mon
Aug 31Mon