Google brings agentic AI to Gemini, starting with businesses
Google announced at a Google Cloud event that Gemini now includes a unified agent that can plan work.
Google announced at a Google Cloud event that Gemini now includes a unified agent that can plan work.
AWS shows how to build an airline voice concierge on Amazon Bedrock AgentCore.
AWS shows how to automate remediation after an AWS DevOps Agent investigation using AWS Lambda Durable Functions.
Cornerstone OnDemand built Orion AI, a multi-agent system on Amazon Bedrock and Strands Agents.
Stacklok, founded by Kubernetes creators Craig McLuckie and Joe Beda.
Simon Willison released llm-openai-decisions 0.1a0, an LLM plugin for OpenAI's new Decisions API announced at DevDay.
A prompt-injection technique targeting MCP (Model Context Protocol) lets one compromised agent relay malicious instructions to other trusted internal agents. Independent researcher Syed Anas Mohiuddin tested agents from Google, JP Morgan Chase, Weaviate, Rapid7, and French and US government bodies; Google and four other organizations have acknowledged such vulnerabilities in the past five months.
Microsoft and Hugging Face released ThinkingBox, an agent benchmark that grades terminal backend state and side effects rather than final responses or tool-call validity.
Why it matters: The paper's 20-run repeat metric and failure breakdown show why a clean tool-call trace can still leave the wrong database state.
Pi released Pi 1.0 and Pi Durable, both of which hit the front page of Hacker News. Pi 1.0 adds Codemode with native support for MCP.
Hugging Face researchers propose ProvenanceGuard, a post-generation verification layer for black-box MCP agents that checks whether each claim is supported by the source the answer names.
Asana builds its AI agents on the Work Graph model, so agents take defined roles.
GitHub Copilot app's canvases are customizable, bidirectional interfaces you create by running the /create-canvas skill and describing the workflow in plain English. The agent builds the UI in the right-side panel, and both you and the agent can update its shared state at the same time. Canvases are saved as extensions for reuse or team sharing, and ready-made ones are available via Awesome Copilot.
Anthropic opened a directory submission portal for Claude plugins, which package MCP connectors.
Why it matters: Anthropic lays out the plugin packaging and submission path, so developers can see how a connector or skill becomes a listed extension.
GitHub Security Lab released the Fuzzing Taskflow, an autonomous fuzzing pipeline for C/C++ projects built on its Taskflow Agent framework. Pointed at a GitHub owner/repo slug.
Why it matters: The post details how the agent splits judgment from execution across MCP tools, useful for anyone building autonomous security pipelines.
A GitHub Podcast episode pushes back on five common AI hot takes: you still need to read AI-generated code.
Anthropic released three beta features on the Claude Developer Platform: Tool Search Tool.
Why it matters: The post gives the token and accuracy numbers behind three tool-use features, so readers can judge which bottleneck in their own agent setup each one addresses.
Anthropic Engineering describes presenting MCP servers as code APIs instead of direct tool calls.
Why it matters: Anthropic's own walkthrough of turning MCP servers into code APIs, with the token math and the sandboxing tradeoff spelled out.
Anthropic introduced Agent Skills, folders of instructions.
Why it matters: Anthropic's own account of how Agent Skills load context in layers, useful for anyone packaging agent expertise.