Google's Gemini 4 "Carbon" model reportedly feels like Anthropic's Opus 5.5 coding performance
Google is internally testing Gemini 4 variants Argon, Barium, and Carbon.
Google is internally testing Gemini 4 variants Argon, Barium, and Carbon.
A study of coding practices across hundreds of firms finds that human code review acts as a significant bottleneck for AI coding tools.
Simon Willison shipped a Newsletters index page for his blog.
LegalOn cut estimated daily Codex costs by 65% while maintaining development speed. It matched Astra, Sol, and Luna to tasks and managed budgets strategically.
Oracle is using ChatGPT Work and Codex to turn specialist knowledge into fast, repeatable workflows across recruiting, engineering, and operations, cutting days of work down to minutes.
Anthropic shipped Claude Haiku 5.5, its first Haiku-tier update in about a year.
Same story, featured as“Anthropic releases Claude Haiku 5.5 with up to 90 percent price cuts and large benchmark gains”
Artcraft, which began last year as a controllable AI tool for artists, has pivoted to a suite of seven open source apps recreating the interfaces and tools of Adobe Photoshop.
Mistral released Mistral Large 4, a 1 trillion-parameter open-weight model nicknamed Le Chonk.
NVIDIA reports that fine-tuned Nemotron 3 systems reached gold-medal level at both IOI 2026 and IMO 2026.
Why it matters: The post lays out a four-part specialization recipe and the SFT, RL and inference-loop split behind two gold-level competition results.
Mistral AI launched a public preview of Mistral Large 4, a 1 trillion-parameter natively multimodal model with 52 billion active parameters.
Why it matters: The post gives the parameter layout, benchmark scores and the weight-release timing, so readers can judge where an open-weight European model now sits.
Reflection announced Beam, a text-only 501B-total / 23B-active MoE for coding, agentic and scientific work.
GitHub released ReviewBench, an open offline benchmark for AI code review agents.
Why it matters: GitHub's own numbers show how an offline code-review benchmark tracked a production A/B test, useful for teams weighing offline signals.
GitHub outlines three skills developers need as AI reshapes their work: directing AI agents rather than just using them.
Google DeepMind announced Gemini 4 Argon, a frontier model built for deep reasoning across long-horizon workflows.
Why it matters: The post details Argon's internal Google results and its phased rollout through the Fairwind Program, giving a concrete picture of frontier capability and access limits.
Claude for Government is now generally available to federal and state agencies.
Why it matters: Details the FedRAMP High environment, spend caps and ATO-oriented audit controls agencies get before adopting Claude.
Anthropic released Claude Opus 5.5, which it estimates costs about 40% less to run than Opus 5 for typical token-billed workloads. Input and output token prices were cut 20% and cached token reads 60%, and the company says Opus 5.5 generates output more than 30% faster than Opus 5. Anthropic also published Claude Code usage data from March to September 2026 showing context per request grew 2.6x, Claude works 3.3x longer per prompt with over 40% more model calls, and cache-missing input fell by more than 50%.
Why it matters: The post pairs Claude Code usage data with the pricing and cache mechanics behind Opus 5.5, useful for judging cost on long coding sessions.
GitHub rebuilt the pull request view in its Copilot app to keep reviews fast on enormous diffs.
A GitHub Podcast episode pushes back on five common AI hot takes: you still need to read AI-generated code.
Mistral helped a European energy operator migrate 40,000 lines of Fortran 77 to C++.