ttok 1.0 released, defaults to GPT-5/GPT-6 tokenizer
Simon Willison released ttok 1.0, switching the default tokenizer from GPT-4 to the GPT-5/GPT-6 family. The change follows an experiment by William Liu showing all seven GPT models (5.5.
Simon Willison released ttok 1.0, switching the default tokenizer from GPT-4 to the GPT-5/GPT-6 family. The change follows an experiment by William Liu showing all seven GPT models (5.5.
Simon Willison released ttok 0.4, updating his tiktoken-based token-counting CLI after a couple of years. The release fixes a Click warning.
TII released Falcon-ASR, a 1.6B-parameter speech recognition model focused on Arabic and the Emirati dialect.
Simon Willison documented running Parseable, a new OpenTelemetry-compatible observability tool.
OpenAI released 372 mathematical results generated by an internal frontier model.
Same story, featured as“OpenAI shares AI results on open mathematics problems with Lean proofs on GitHub”
Musubi announced PolicyLM-1.7B, a lightweight open-weights decision model built for real-time content moderation that applies a plain-English content policy to messages in under 50 milliseconds. The company says it is designed to be similar in cost and speed to the AI classifiers used by most social platforms, but can apply complex policies without special training and needs no retraining when a policy changes, letting policy-setters iterate. Co-founder and chief AI officer Filip Jankovic frames it as a way for platform managers to label content proactively, and the announcement positions it against TypeSafe AI's Jev, released in September and followed by competing decision models from OpenAI and Amazon.
Artcraft, which began last year as a controllable AI tool for artists, has pivoted to a suite of seven open source apps recreating the interfaces and tools of Adobe Photoshop.
Microsoft Research Asia open-sourced Agent Lightning v1.0, a roughly 3.
Why it matters: The original gives the framework's design choices and a measured SWE-bench gain, so readers can judge whether to reuse their existing harness for RL.
Mistral released Mistral Large 4, a 1 trillion-parameter open-weight model nicknamed Le Chonk.
OpenAI published 722 mathematical manuscripts from an unreleased internal frontier model in a public GitHub repo.
Why it matters: The release details how many open problems were addressed and how much compute each solution took, which frames how AI math results are produced.
Simon Willison released llm-openai-decisions 0.1a0, an LLM plugin for OpenAI's new Decisions API announced at DevDay.
Simon Willison released llm-mistral 0.16, a new version of the Mistral plugin for his LLM CLI tool. The post is dated 6th October 2026.
Mistral AI launched a public preview of Mistral Large 4, a 1 trillion-parameter natively multimodal model with 52 billion active parameters.
Why it matters: The post gives the parameter layout, benchmark scores and the weight-release timing, so readers can judge where an open-weight European model now sits.
OpenAI published new results on open problems in mathematics produced by an internal frontier model, and shared Lean proof formalizations and research details on GitHub.
Reflection announced Beam, a text-only 501B-total / 23B-active MoE for coding, agentic and scientific work.
GitHub released ReviewBench, an open offline benchmark for AI code review agents.
Why it matters: GitHub's own numbers show how an offline code-review benchmark tracked a production A/B test, useful for teams weighing offline signals.
Ai2 open-sourced AstaBrief 8B, a model built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report.
Why it matters: The post gives the training recipe and filtering lessons behind an open-weights scientific report model, useful for anyone building grounded generation.
Pi released Pi 1.0 and Pi Durable, both of which hit the front page of Hacker News. Pi 1.0 adds Codemode with native support for MCP.
Hugging Face releases the Open TTS Leaderboard, using objective metrics (WER/CER via Qwen3 ASR.
H Company released Holo4, a new series of generalist computer-use agent models in two sizes.
Why it matters: The post gives the two model sizes, the interfaces they cover and the OSWorld 2.0 numbers, so readers can weigh a cheaper open-weight computer-use agent against closed frontier models.
Anthropic and NVIDIA collaborated to add security and control layers to the agent stack.
Why it matters: The post lays out the split-brain sandbox architecture and the policy rules that decide what an agent may reach.
GitHub Security Lab released the Fuzzing Taskflow, an autonomous fuzzing pipeline for C/C++ projects built on its Taskflow Agent framework. Pointed at a GitHub owner/repo slug.
Why it matters: The post details how the agent splits judgment from execution across MCP tools, useful for anyone building autonomous security pipelines.
Microsoft Research published a systematic study of mobile robotic manipulation workloads showing that running physical AI inference only on onboard GPUs limits robot performance.
Why it matters: The measurement study quantifies how onboard GPU limits hurt task success and battery life, and what offloading changes.
Microsoft Research published RetroChimera in Nature.
Google Research introduces MilleMiglia, a C++ instance generator that creates realistic.
Microsoft Research released GigaPath-Flash and GigaTIME-Flash.
IBM Research extended K-Search, the evolutionary kernel search framework from UC Berkeley Sky Lab.
Why it matters: The post details how a structured CUDA-to-MLX translation layer, not the LLM itself, drives the kernel gains.
Berkeley AI Research introduces SPEX and ProxySPEX, algorithms that identify influential feature.
Berkeley AI Research researchers developed an information-based framework for evaluating and optimizing imaging systems.
Anthropic introduced two sandboxing features in Claude Code: a sandboxed bash tool.
Why it matters: Anthropic gives the sandboxing design and its internal 84% drop in permission prompts, useful for anyone running coding agents.