AWS September 2026 recap: Bedrock adds GPT-6 Astra, Claude Opus 5.5, Kimi K3; AgentCore and Strands updates
AWS's September 2026 Bedrock updates added OpenAI's GPT-6 Astra, Sol, and Luna models.
AWS's September 2026 Bedrock updates added OpenAI's GPT-6 Astra, Sol, and Luna models.
Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic.
Why it matters: The post details Haiku 5.5's effort controls and subagent role, useful for judging cost and routing tradeoffs.
At a Microsoft event in San Francisco, NVIDIA and Microsoft announced RTX Spark.
Why it matters: NVIDIA and Microsoft lay out the hardware and OS primitives for running agents locally on Windows, with preorder timing and specs.
Amazon Quick and Amazon Bedrock Knowledge Bases add real-time ACL enforcement on top of pre-retrieval filtering.
GitHub built a fine-tuned ModernBERT classifier with Microsoft Applied Sciences that assesses candidate secrets in context in under two milliseconds.
Why it matters: GitHub's nine quarters of push data and the latency budget behind its new secret classifier show how prevention is being moved into the push path.
Microsoft Research Asia open-sourced Agent Lightning v1.0, a roughly 3.
Why it matters: The original gives the framework's design choices and a measured SWE-bench gain, so readers can judge whether to reuse their existing harness for RL.
Qlik built Qlik Answers on Amazon Bedrock to give employees grounded, sourced answers from knowledge bases.
Microsoft Research's Jennifer Neville discusses how evaluation pushes AI systems beyond traditional benchmarks and why "surprising failures" emerge when models are tested on real user needs. She offers practical guidance for working with current AI systems and explains why examining data matters when results defy expectations.
GitHub released ReviewBench, an open offline benchmark for AI code review agents.
Why it matters: GitHub's own numbers show how an offline code-review benchmark tracked a production A/B test, useful for teams weighing offline signals.
Microsoft and Hugging Face released ThinkingBox, an agent benchmark that grades terminal backend state and side effects rather than final responses or tool-call validity.
Why it matters: The paper's 20-run repeat metric and failure breakdown show why a clean tool-call trace can still leave the wrong database state.
Microsoft Research built a machine learning pipeline that forecasts geomagnetic storm risk for 66.
Microsoft Research introduced Quine, a research effort combining a multimodal world model of biology with a harness that connects models.
Why it matters: The original gives the system's design and a concrete wet-lab validation, so readers can judge how a multimodal world model fits into real experimental loops.
Microsoft Research Asia – Singapore marks one year since opening as Microsoft's first research lab in Southeast Asia.
GitHub Copilot app's canvases are customizable, bidirectional interfaces you create by running the /create-canvas skill and describing the workflow in plain English. The agent builds the UI in the right-side panel, and both you and the agent can update its shared state at the same time. Canvases are saved as extensions for reuse or team sharing, and ready-made ones are available via Awesome Copilot.
GitHub argues chat is often the wrong interface for AI.
GitHub rebuilt the pull request view in its Copilot app to keep reviews fast on enormous diffs.
Microsoft Research published a systematic study of mobile robotic manipulation workloads showing that running physical AI inference only on onboard GPUs limits robot performance.
Why it matters: The measurement study quantifies how onboard GPU limits hurt task success and battery life, and what offloading changes.
Microsoft Research published RetroChimera in Nature.
Microsoft Research released GigaPath-Flash and GigaTIME-Flash.