Google Research ran a three-month field experiment with 133 lawyers at eleven IP firms.
Why it matters: The three-month field experiment separates AI-assisted drafting gains from unassisted redlining skill, showing where juniors stall and seniors improve.
NVIDIA reports that fine-tuned Nemotron 3 systems reached gold-medal level at both IOI 2026 and IMO 2026.
Why it matters: The post lays out a four-part specialization recipe and the SFT, RL and inference-loop split behind two gold-level competition results.
OpenAI published new results on open problems in mathematics produced by an internal frontier model, and shared Lean proof formalizations and research details on GitHub.
GitHub released ReviewBench, an open offline benchmark for AI code review agents.
Why it matters: GitHub's own numbers show how an offline code-review benchmark tracked a production A/B test, useful for teams weighing offline signals.
Microsoft and Hugging Face released ThinkingBox, an agent benchmark that grades terminal backend state and side effects rather than final responses or tool-call validity.
Why it matters: The paper's 20-run repeat metric and failure breakdown show why a clean tool-call trace can still leave the wrong database state.
Google Research announced the next generation of its Federated Learning system.
Why it matters: Google's TEE-based federated learning design shows how verifiable execution and differential privacy are combined in a production system.
Google DeepMind introduced SynthID Bio, a family of watermarking methods that embeds a verifiable signature into AI-generated biological code while preserving protein function in laboratory testing. In wet-lab tests across VEGF-A, the SARS-CoV-2 spike protein RBD and PD-L1, watermarked binder designs matched the hit rate, binding affinity and natural sequence diversity of unwatermarked versions, and for protein folding the method fine-tunes part of AlphaFold 3's diffusion network so predicted 3D coordinates carry a detectable signature. DeepMind is publishing the methods paper and open-sourcing the code, in vitro data and model weights, and says key challenges include making the watermark more robust against deliberate tampering.
Why it matters: The post details how a watermark is embedded into protein sequences and structures and what wet-lab tests showed about function.
Microsoft Research introduced Quine, a research effort combining a multimodal world model of biology with a harness that connects models.
Why it matters: The original gives the system's design and a concrete wet-lab validation, so readers can judge how a multimodal world model fits into real experimental loops.
Microsoft Research published a systematic study of mobile robotic manipulation workloads showing that running physical AI inference only on onboard GPUs limits robot performance.
Why it matters: The measurement study quantifies how onboard GPU limits hurt task success and battery life, and what offloading changes.
Anthropic announced a new life sciences research group and lab.
Why it matters: The post gives the agent count, token budget and search time behind one autonomous discovery, useful for judging AI-driven hypothesis generation.
Google DeepMind introduced AlphaGenome Atlas, a platform with precomputed AlphaGenome predictions for the effects of 9 billion single-nucleotide variants.
Why it matters: The 1-petabyte scale and the AVI score show how a precomputed variant map changes what geneticists can screen without lab work.