Microsoft's Decision-1 model enters the fast-growing AI decision model race
Microsoft released Decision-1, a decision model built for fast, structured decisions such as classifications.
Microsoft released Decision-1, a decision model built for fast, structured decisions such as classifications.
TII released Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic and the Emirati dialect.
Anthropic shipped Claude Haiku 5.5, its first Haiku-tier update in about a year.
Same story, featured as“Anthropic releases Claude Haiku 5.5 with up to 90 percent price cuts and large benchmark gains”
Simon Willison tested Mistral Large 4 alongside Claude Opus 5.5, GPT-6.1-sol.
Mistral released a preview of Mistral Large 4, a 1 trillion parameter model with 49 billion active parameters trained on its own cluster of 3.
Google released EmbeddingGemma 2, an open model that converts text, images, video.
Musubi announced PolicyLM-1.7B, a lightweight open-weights decision model built for real-time content moderation that applies a plain-English content policy to messages in under 50 milliseconds. The company says it is designed to be similar in cost and speed to the AI classifiers used by most social platforms, but can apply complex policies without special training and needs no retraining when a policy changes, letting policy-setters iterate. Co-founder and chief AI officer Filip Jankovic frames it as a way for platform managers to label content proactively, and the announcement positions it against TypeSafe AI's Jev, released in September and followed by competing decision models from OpenAI and Amazon.
Anthropic released Claude Haiku 5.5, a fast low-cost model priced at $0.10/$0.50 per million input/output tokens up to 100.
OpenAI is rolling out GPT-6 with a feature called Intelligent UI for all ChatGPT users.
Why it matters: The piece lays out GPT-6's Intelligent UI and its tiered rollout, useful for gauging how chat answers shift from text to interfaces.
Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic.
Why it matters: The post details Haiku 5.5's effort controls and subagent role, useful for judging cost and routing tradeoffs.
Anthropic released Claude Haiku 5.5, its fastest and most affordable small model.
Why it matters: The pricing table and benchmark jumps let readers weigh Haiku 5.5 against their own cost-sensitive workloads.
Mistral released Mistral Large 4, a 1 trillion-parameter open-weight model nicknamed Le Chonk.
GPT-6 is rolling out globally in ChatGPT with Intelligent UI, delivering faster responses with visuals and interactive experiences users can explore and use directly.
Google DeepMind launched EmbeddingGemma 2, a 740M-parameter open embedding model under Apache 2.0 that maps text.
Why it matters: The model card numbers let readers judge whether a 740M on-device embedder can replace their current retrieval stack.
Mistral AI launched a public preview of Mistral Large 4, a 1 trillion-parameter natively multimodal model with 52 billion active parameters.
Why it matters: The post gives the parameter layout, benchmark scores and the weight-release timing, so readers can judge where an open-weight European model now sits.
Reflection announced Beam, a text-only 501B-total / 23B-active MoE for coding, agentic and scientific work.
OpenAI released GPT-6.1 Sol at $2/$10 per million input/output tokens.
Ai2 open-sourced AstaBrief 8B, a model built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report.
Why it matters: The post gives the training recipe and filtering lessons behind an open-weights scientific report model, useful for anyone building grounded generation.
Google is rolling out Gemini 4 Argon with frontier-level benchmarks at $2/$10.
Google DeepMind introduced Gemini 4 Argon for coding, enterprise knowledge work and cyber defense.
Why it matters: The roundup lays out Argon's benchmark wins, its 1M-token output mechanism and its limited preview access side by side.
Google DeepMind announced Gemini 4 Argon, a frontier model built for deep reasoning across long-horizon workflows.
Why it matters: The post details Argon's internal Google results and its phased rollout through the Fairwind Program, giving a concrete picture of frontier capability and access limits.
H Company released Holo4, a new series of generalist computer-use agent models in two sizes.
Why it matters: The post gives the two model sizes, the interfaces they cover and the OSWorld 2.0 numbers, so readers can weigh a cheaper open-weight computer-use agent against closed frontier models.
Google DeepMind introduced Gemini 3.8 Live with Live Avatar, which pairs near real-time video generation with speech so Gemini's live dialogue models can listen.
Why it matters: The post details how low-latency video, async tool calls and 97-language lip-sync change what enterprise agents can do.
Anthropic released Claude Opus 5.5, which it estimates costs about 40% less to run than Opus 5 for typical token-billed workloads. Input and output token prices were cut 20% and cached token reads 60%, and the company says Opus 5.5 generates output more than 30% faster than Opus 5. Anthropic also published Claude Code usage data from March to September 2026 showing context per request grew 2.6x, Claude works 3.3x longer per prompt with over 40% more model calls, and cache-missing input fell by more than 50%.
Why it matters: The post pairs Claude Code usage data with the pricing and cache mechanics behind Opus 5.5, useful for judging cost on long coding sessions.
Google DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS.
Why it matters: The post details voice design, line-by-line direction and rollout channels, so readers can judge fit for dubbing or voice agents.
Google DeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking.
Why it matters: The post gives benchmark placements and rollout channels, so readers can weigh the two Live variants against their own voice-agent needs.
Microsoft Research released GigaPath-Flash and GigaTIME-Flash.