GitHub built a fine-tuned ModernBERT classifier with Microsoft Applied Sciences that assesses candidate secrets in context in under two milliseconds.
Why it matters: GitHub's nine quarters of push data and the latency budget behind its new secret classifier show how prevention is being moved into the push path.
Microsoft Research Asia open-sourced Agent Lightning v1.0, a roughly 3.
Why it matters: The original gives the framework's design choices and a measured SWE-bench gain, so readers can judge whether to reuse their existing harness for RL.
NVIDIA will offer DGX Spark with 64GB of unified memory from Acer, ASUS, Dell, Gigabyte.
Why it matters: The post gives the 64GB configuration's price, memory ceiling and two-unit clustering numbers, so readers can size local agent workloads against it.
GPT-6 Astra Ultrafast is now available in the OpenAI API and to eligible ChatGPT Work and Codex users.
Why it matters: The post gives the speedup figure and the agent loop it targets, so readers can judge whether the latency change matters for their own tool-calling workflows.
H Company released Holo4, a new series of generalist computer-use agent models in two sizes.
Why it matters: The post gives the two model sizes, the interfaces they cover and the OSWorld 2.0 numbers, so readers can weigh a cheaper open-weight computer-use agent against closed frontier models.
Anthropic opened a directory submission portal for Claude plugins, which package MCP connectors.
Why it matters: Anthropic lays out the plugin packaging and submission path, so developers can see how a connector or skill becomes a listed extension.
GitHub Security Lab released the Fuzzing Taskflow, an autonomous fuzzing pipeline for C/C++ projects built on its Taskflow Agent framework. Pointed at a GitHub owner/repo slug.
Why it matters: The post details how the agent splits judgment from execution across MCP tools, useful for anyone building autonomous security pipelines.
Anthropic released Claude Opus 5.5, which it estimates costs about 40% less to run than Opus 5 for typical token-billed workloads. Input and output token prices were cut 20% and cached token reads 60%, and the company says Opus 5.5 generates output more than 30% faster than Opus 5. Anthropic also published Claude Code usage data from March to September 2026 showing context per request grew 2.6x, Claude works 3.3x longer per prompt with over 40% more model calls, and cache-missing input fell by more than 50%.
Why it matters: The post pairs Claude Code usage data with the pricing and cache mechanics behind Opus 5.5, useful for judging cost on long coding sessions.
Microsoft Research published a systematic study of mobile robotic manipulation workloads showing that running physical AI inference only on onboard GPUs limits robot performance.
Why it matters: The measurement study quantifies how onboard GPU limits hurt task success and battery life, and what offloading changes.
Mistral released Agentic Search, a multi-step retrieval layer that lets models search.
Why it matters: The post gives benchmark deltas and the five retrieval tools, so readers can judge whether their one-shot RAG pipeline should be replaced.
Anthropic released three beta features on the Claude Developer Platform: Tool Search Tool.
Why it matters: The post gives the token and accuracy numbers behind three tool-use features, so readers can judge which bottleneck in their own agent setup each one addresses.