Arena raises $200M Series B at $3.1B valuation and adds an alignment leaderboard
Arena, the UC Berkeley-originated crowdsourced AI model leaderboard, raised a $200 million Series B at a $3.1 billion valuation.
Arena, the UC Berkeley-originated crowdsourced AI model leaderboard, raised a $200 million Series B at a $3.1 billion valuation.
GitHub released ReviewBench, an open offline benchmark for AI code review agents.
Why it matters: GitHub's own numbers show how an offline code-review benchmark tracked a production A/B test, useful for teams weighing offline signals.
Microsoft and Hugging Face released ThinkingBox, an agent benchmark that grades terminal backend state and side effects rather than final responses or tool-call validity.
Why it matters: The paper's 20-run repeat metric and failure breakdown show why a clean tool-call trace can still leave the wrong database state.
OpenAI released GPT-6.1 Sol at $2/$10 per million input/output tokens.
Hugging Face releases the Open TTS Leaderboard, using objective metrics (WER/CER via Qwen3 ASR.
Nathan Lambert argues Chinese labs remain the clear leaders in open-weight models.
Mistral released Agentic Search, a multi-step retrieval layer that lets models search.
Why it matters: The post gives benchmark deltas and the five retrieval tools, so readers can judge whether their one-shot RAG pipeline should be replaced.
Import AI 469 covers DiG-bench, a 70-game benchmark testing whether AI can infer hidden rules through exploration.