Asana cuts model costs 76x in browser tests with GPT-6.1 Sol
Asana used GPT-6 Astra in Codex to make its browser agent 76x cheaper and 5x faster in tests. The goal is to offer customers more capable models.
Asana used GPT-6 Astra in Codex to make its browser agent 76x cheaper and 5x faster in tests. The goal is to offer customers more capable models.
Postman detailed how its Agent Mode serves 40 million developers on Amazon Bedrock.
AWS's September 2026 Bedrock updates added OpenAI's GPT-6 Astra, Sol, and Luna models.
A Hugging Face author used ML Intern in HuggingChat to build six models over a few days for about USD 103 in total compute.
Why it matters: A first-hand account of prompting an agent to train six small models, with per-project budgets and costs.
NVIDIA developers are using frontier AI models such as GPT-6 Astra with Omniverse libraries to turn simulation ideas into working applications. Projects include a humanoid warehouse simulator, an autonomous-driving test workflow on San Francisco's Market Street, and Robo Olympics, where a simulated Unitree G1 humanoid cleared a hurdle in 64 of 100 trials.
Amazon Bedrock AgentCore payments lets agents pay for services on demand.
AWS shows how to build an airline voice concierge on Amazon Bedrock AgentCore.
Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic.
Why it matters: The post details Haiku 5.5's effort controls and subagent role, useful for judging cost and routing tradeoffs.
At a Microsoft event in San Francisco, NVIDIA and Microsoft announced RTX Spark.
Why it matters: NVIDIA and Microsoft lay out the hardware and OS primitives for running agents locally on Windows, with preorder timing and specs.
Amazon Quick and Amazon Bedrock Knowledge Bases add real-time ACL enforcement on top of pre-retrieval filtering.
Microsoft Research Asia open-sourced Agent Lightning v1.0, a roughly 3.
Why it matters: The original gives the framework's design choices and a measured SWE-bench gain, so readers can judge whether to reuse their existing harness for RL.
AWS outlines an Agentic Value Model for justifying agentic automation, arguing the RPA-era ROI formula of hours saved times labor cost misses most agent value. It adds exception handling.
Qlik built Qlik Answers on Amazon Bedrock to give employees grounded, sourced answers from knowledge bases.
AWS shows how to automate remediation after an AWS DevOps Agent investigation using AWS Lambda Durable Functions.
AWS shares a six-week program that pairs non-engineering business professionals with mentors and production-grade tools to build working AI prototypes. A team of four built WealthWise.
Cornerstone OnDemand built Orion AI, a multi-agent system on Amazon Bedrock and Strands Agents.
OpenAI is bringing College Planner to ChatGPT for Teens to help students manage college applications, alongside new flashcards and quizzes. The company is also forming a teen AI council.
Radisson Hotel Group partnered with Accenture to build a ChatGPT plugin on OpenAI technology, letting travelers find, compare, and book hotels while planning trips.
This post shows how to build a personal assistant with persistent memory using OpenClaw.
Microsoft Research's Jennifer Neville discusses how evaluation pushes AI systems beyond traditional benchmarks and why "surprising failures" emerge when models are tested on real user needs. She offers practical guidance for working with current AI systems and explains why examining data matters when results defy expectations.
Mistral AI launched a public preview of Mistral Large 4, a 1 trillion-parameter natively multimodal model with 52 billion active parameters.
Why it matters: The post gives the parameter layout, benchmark scores and the weight-release timing, so readers can judge where an open-weight European model now sits.
OpenAI and Ironclad are training and evaluating AI agents on complex contracting workflows to advance computer use for professional work.
Google released a CAPS workshop report on agentic privacy and security.
Anthropic is launching an expanded Cyber Verification Program that merges Project Glasswing and the earlier CVP into three access tiers.
Why it matters: The tier structure and CyScenarioBench block rates show how safeguard levels are traded against defensive access.
Cresta built Conductor, a natural-language agent builder on the Claude Agent SDK.
Microsoft and Hugging Face released ThinkingBox, an agent benchmark that grades terminal backend state and side effects rather than final responses or tool-call validity.
Why it matters: The paper's 20-run repeat metric and failure breakdown show why a clean tool-call trace can still leave the wrong database state.
GitHub outlines three skills developers need as AI reshapes their work: directing AI agents rather than just using them.
ServiceNow CoreAI built AutoSynthData, a pipeline that turns a target model's failures and a stronger teacher's successes into new training tasks for enterprise agents. It generates tasks as system specification, user prompt, and verifier, then validates them in the environment and uses accepted samples for post-training, with the curriculum shifting toward remaining weaknesses. The pipeline is illustrated with EnterpriseOps Gym.
Anthropic introduced mods, small TypeScript functions that change how Claude Code works by hooking into events such as tool calls.
Why it matters: The post details how mods hook Claude Code events and what admins can restrict, useful for judging control over an existing workflow.
Google DeepMind announced Gemini 4 Argon, a frontier model built for deep reasoning across long-horizon workflows.
Why it matters: The post details Argon's internal Google results and its phased rollout through the Fairwind Program, giving a concrete picture of frontier capability and access limits.
Claude for Government is now generally available to federal and state agencies.
Why it matters: Details the FedRAMP High environment, spend caps and ATO-oriented audit controls agencies get before adopting Claude.
Anthropic's sales team built a buying agent on Claude Managed Agents (beta) that handles thousands of conversations daily.
Hugging Face researchers propose ProvenanceGuard, a post-generation verification layer for black-box MCP agents that checks whether each claim is supported by the source the answer names.
Asana builds its AI agents on the Work Graph model, so agents take defined roles.
Microsoft Research Asia – Singapore marks one year since opening as Microsoft's first research lab in Southeast Asia.
H Company released Holo4, a new series of generalist computer-use agent models in two sizes.
Why it matters: The post gives the two model sizes, the interfaces they cover and the OSWorld 2.0 numbers, so readers can weigh a cheaper open-weight computer-use agent against closed frontier models.
Anthropic and NVIDIA collaborated to add security and control layers to the agent stack.
Why it matters: The post lays out the split-brain sandbox architecture and the policy rules that decide what an agent may reach.
GitHub Copilot app's canvases are customizable, bidirectional interfaces you create by running the /create-canvas skill and describing the workflow in plain English. The agent builds the UI in the right-side panel, and both you and the agent can update its shared state at the same time. Canvases are saved as extensions for reuse or team sharing, and ready-made ones are available via Awesome Copilot.
Anthropic opened a directory submission portal for Claude plugins, which package MCP connectors.
Why it matters: Anthropic lays out the plugin packaging and submission path, so developers can see how a connector or skill becomes a listed extension.
GitHub argues chat is often the wrong interface for AI.