Skip to content
  1. Hugging Face Blog66

    NVIDIA fine-tunes Nemotron 3 into IOI and IMO gold-level specialists

    NVIDIA reports that fine-tuned Nemotron 3 systems reached gold-medal level at both IOI 2026 and IMO 2026.

    Why it matters: The post lays out a four-part specialization recipe and the SFT, RL and inference-loop split behind two gold-level competition results.

  1. GitHub Blog · AI & ML62

    GitHub releases ReviewBench, an open benchmark for AI code review

    GitHub released ReviewBench, an open offline benchmark for AI code review agents.

    Why it matters: GitHub's own numbers show how an offline code-review benchmark tracked a production A/B test, useful for teams weighing offline signals.

  1. Anthropic Blog71

    Anthropic releases Claude Opus 5.5 for longer, context-heavy coding sessions

    Anthropic released Claude Opus 5.5, which it estimates costs about 40% less to run than Opus 5 for typical token-billed workloads. Input and output token prices were cut 20% and cached token reads 60%, and the company says Opus 5.5 generates output more than 30% faster than Opus 5. Anthropic also published Claude Code usage data from March to September 2026 showing context per request grew 2.6x, Claude works 3.3x longer per prompt with over 40% more model calls, and cache-missing input fell by more than 50%.

    Why it matters: The post pairs Claude Code usage data with the pricing and cache mechanics behind Opus 5.5, useful for judging cost on long coding sessions.

You've reached the end.