Falcon-ASR: TII's 1.6B Arabic speech recognition model
TII released Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic and the Emirati dialect.
TII released Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic and the Emirati dialect.
Microsoft Research Asia open-sourced Agent Lightning v1.0, a roughly 3.
Why it matters: The original gives the framework's design choices and a measured SWE-bench gain, so readers can judge whether to reuse their existing harness for RL.
Mistral AI launched a public preview of Mistral Large 4, a 1 trillion-parameter natively multimodal model with 52 billion active parameters.
Why it matters: The post gives the parameter layout, benchmark scores and the weight-release timing, so readers can judge where an open-weight European model now sits.
OpenAI published new results on open problems in mathematics produced by an internal frontier model, and shared Lean proof formalizations and research details on GitHub.
GitHub released ReviewBench, an open offline benchmark for AI code review agents.
Why it matters: GitHub's own numbers show how an offline code-review benchmark tracked a production A/B test, useful for teams weighing offline signals.
Ai2 open-sourced AstaBrief 8B, a model built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report.
Why it matters: The post gives the training recipe and filtering lessons behind an open-weights scientific report model, useful for anyone building grounded generation.
Hugging Face releases the Open TTS Leaderboard, using objective metrics (WER/CER via Qwen3 ASR.
H Company released Holo4, a new series of generalist computer-use agent models in two sizes.
Why it matters: The post gives the two model sizes, the interfaces they cover and the OSWorld 2.0 numbers, so readers can weigh a cheaper open-weight computer-use agent against closed frontier models.
Anthropic and NVIDIA collaborated to add security and control layers to the agent stack.
Why it matters: The post lays out the split-brain sandbox architecture and the policy rules that decide what an agent may reach.
GitHub Security Lab released the Fuzzing Taskflow, an autonomous fuzzing pipeline for C/C++ projects built on its Taskflow Agent framework. Pointed at a GitHub owner/repo slug.
Why it matters: The post details how the agent splits judgment from execution across MCP tools, useful for anyone building autonomous security pipelines.
Microsoft Research published a systematic study of mobile robotic manipulation workloads showing that running physical AI inference only on onboard GPUs limits robot performance.
Why it matters: The measurement study quantifies how onboard GPU limits hurt task success and battery life, and what offloading changes.
Microsoft Research published RetroChimera in Nature.
Google Research introduces MilleMiglia, a C++ instance generator that creates realistic.
Microsoft Research released GigaPath-Flash and GigaTIME-Flash.
IBM Research extended K-Search, the evolutionary kernel search framework from UC Berkeley Sky Lab.
Why it matters: The post details how a structured CUDA-to-MLX translation layer, not the LLM itself, drives the kernel gains.
Berkeley AI Research introduces SPEX and ProxySPEX, algorithms that identify influential feature.
Berkeley AI Research researchers developed an information-based framework for evaluating and optimizing imaging systems.
Anthropic introduced two sandboxing features in Claude Code: a sandboxed bash tool.
Why it matters: Anthropic gives the sandboxing design and its internal 84% drop in permission prompts, useful for anyone running coding agents.