Skip to content
  1. AWS Machine Learning Blog71

    Claude Haiku 5.5 launches on Amazon Bedrock and Claude Platform on AWS

    Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic.

    Why it matters: The post details Haiku 5.5's effort controls and subagent role, useful for judging cost and routing tradeoffs.

  1. Google DeepMind71

    Google DeepMind releases EmbeddingGemma 2, a 740M open multimodal embedding model

    Google DeepMind launched EmbeddingGemma 2, a 740M-parameter open embedding model under Apache 2.0 that maps text.

    Why it matters: The model card numbers let readers judge whether a 740M on-device embedder can replace their current retrieval stack.

  2. Mistral AI82

    Mistral AI launches Mistral Large 4 public preview, a 1T-parameter multimodal model with weights due this month

    Mistral AI launched a public preview of Mistral Large 4, a 1 trillion-parameter natively multimodal model with 52 billion active parameters.

    Why it matters: The post gives the parameter layout, benchmark scores and the weight-release timing, so readers can judge where an open-weight European model now sits.

  1. Google DeepMind71

    Google DeepMind introduces Gemini 3.8 Live with Live Avatar

    Google DeepMind introduced Gemini 3.8 Live with Live Avatar, which pairs near real-time video generation with speech so Gemini's live dialogue models can listen.

    Why it matters: The post details how low-latency video, async tool calls and 97-language lip-sync change what enterprise agents can do.

  2. Anthropic Blog71

    Anthropic releases Claude Opus 5.5 for longer, context-heavy coding sessions

    Anthropic released Claude Opus 5.5, which it estimates costs about 40% less to run than Opus 5 for typical token-billed workloads. Input and output token prices were cut 20% and cached token reads 60%, and the company says Opus 5.5 generates output more than 30% faster than Opus 5. Anthropic also published Claude Code usage data from March to September 2026 showing context per request grew 2.6x, Claude works 3.3x longer per prompt with over 40% more model calls, and cache-missing input fell by more than 50%.

    Why it matters: The post pairs Claude Code usage data with the pricing and cache mechanics behind Opus 5.5, useful for judging cost on long coding sessions.

  1. Google DeepMind71

    Google DeepMind releases Gemini 3.8 Flash TTS and Flash-Lite TTS

    Google DeepMind introduced Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS.

    Why it matters: The post details voice design, line-by-line direction and rollout channels, so readers can judge fit for dubbing or voice agents.

You've reached the end.