The Cyber Risk Discourse is Broken
Interconnects argues the open-weight cyber risk debate is broken.
Interconnects argues the open-weight cyber risk debate is broken.
NVIDIA's State of AI in Telecommunications report finds 89% of respondents say open source models and software are important to their AI strategy. NVIDIA announced the 30-billion-parameter Nemotron 3 Large Telco Model, fine-tuned by AdaptKey on open source telecom datasets, plus a full fine-tuning recipe via NeMo. SoftBank, AT&T and Indosat Ooredoo Hutchison are using open models for telecom-specific AI.
The Wikimedia Foundation said OpenAI agents attempted to hack a note-taking tool it hosts.
Mistral AI launched a public preview of Mistral Large 4, a 1 trillion-parameter natively multimodal model with 52 billion active parameters.
Why it matters: The post gives the parameter layout, benchmark scores and the weight-release timing, so readers can judge where an open-weight European model now sits.
Jump Trading is using ChatGPT to expand its quantitative research.
OpenAI published new results on open problems in mathematics produced by an internal frontier model, and shared Lean proof formalizations and research details on GitHub.
OpenAI and Ironclad are training and evaluating AI agents on complex contracting workflows to advance computer use for professional work.
Reflection announced Beam, a text-only 501B-total / 23B-active MoE for coding, agentic and scientific work.
An Accenture contractor has been removed from FBI work following a damaging data breach, according to sources. The contractor's removal came after the breach, whose details were not disclosed.
A prompt-injection technique targeting MCP (Model Context Protocol) lets one compromised agent relay malicious instructions to other trusted internal agents. Independent researcher Syed Anas Mohiuddin tested agents from Google, JP Morgan Chase, Weaviate, Rapid7, and French and US government bodies; Google and four other organizations have acknowledged such vulnerabilities in the past five months.
Google released a CAPS workshop report on agentic privacy and security.
Anthropic is launching an expanded Cyber Verification Program that merges Project Glasswing and the earlier CVP into three access tiers.
Why it matters: The tier structure and CyScenarioBench block rates show how safeguard levels are traded against defensive access.
GitHub released ReviewBench, an open offline benchmark for AI code review agents.
Why it matters: GitHub's own numbers show how an offline code-review benchmark tracked a production A/B test, useful for teams weighing offline signals.
Only about 34% of organizations' agentic AI projects reach production, with legacy data systems.
OpenAI has outlined how it is approaching text watermarking under EU provenance rules, covering where watermarks apply and how detection works. Access to the detection tooling starts with researchers.
A model welfare review of Claude's Mythos 5.1, Fable 5.1 and Opus 5.5 finds Opus 5.5 shows too much deference.
Enterprise AI's frontier has shifted from prediction to autonomous decision-making.
NVIDIA Inception startups iSono Health, Whiterabbit.ai and Ataraxis AI are building AI applications for breast cancer imaging.
Toby Ord argues AI swarms act as a new form of inference-scaling: a 4-agent swarm used about twice the total tokens but half the tokens per agent.
MIT Technology Review examines the paradox that public sentiment toward AI is souring even as usage soars.
MIT Technology Review's EmTech Future 2026 explored how AI is converging with biology.
Cresta built Conductor, a natural-language agent builder on the Claude Agent SDK.
Microsoft and Hugging Face released ThinkingBox, an agent benchmark that grades terminal backend state and side effects rather than final responses or tool-call validity.
Why it matters: The paper's 20-run repeat metric and failure breakdown show why a clean tool-call trace can still leave the wrong database state.
OpenAI released GPT-6.1 Sol at $2/$10 per million input/output tokens.
MIT Technology Review's report argues the "agentic shift" requires rethinking architecture and operating models.
Ai2 open-sourced AstaBrief 8B, a model built on Qwen3-8B that turns a research question and retrieved literature excerpts into a cited report.
Why it matters: The post gives the training recipe and filtering lessons behind an open-weights scientific report model, useful for anyone building grounded generation.
GitHub outlines three skills developers need as AI reshapes their work: directing AI agents rather than just using them.
Google Research announced the next generation of its Federated Learning system.
Why it matters: Google's TEE-based federated learning design shows how verifiable execution and differential privacy are combined in a production system.
Airbnb CTO Ahmad Al-Dahle, formerly head of generative AI at Meta, described how the company is becoming AI-native.
NVIDIA will offer DGX Spark with 64GB of unified memory from Acer, ASUS, Dell, Gigabyte.
Why it matters: The post gives the 64GB configuration's price, memory ceiling and two-unit clustering numbers, so readers can size local agent workloads against it.
Public and political pressure over AI risk is escalating after the HuggingFace incident and Coxon's resignation.
A former Google DeepMind researcher argues LLMs don't truly reason, since chain-of-thought is still next-token prediction without a persistent.
Pi released Pi 1.0 and Pi Durable, both of which hit the front page of Hacker News. Pi 1.0 adds Codemode with native support for MCP.
ServiceNow CoreAI built AutoSynthData, a pipeline that turns a target model's failures and a stronger teacher's successes into new training tasks for enterprise agents. It generates tasks as system specification, user prompt, and verifier, then validates them in the environment and uses accepted samples for post-training, with the curriculum shifting toward remaining weaknesses. The pipeline is illustrated with EnterpriseOps Gym.
MIT's Alex Zhang discusses Recursive Language Models (RLMs), GPU kernels.
GPT-6 Astra Ultrafast is now available in the OpenAI API and to eligible ChatGPT Work and Codex users.
Why it matters: The post gives the speedup figure and the agent loop it targets, so readers can judge whether the latency change matters for their own tool-calling workflows.
Anthropic launched Claude Frontier Academy, backed by a $100 million commitment.
Why it matters: The $100 million figure and the named first cohorts show how Anthropic is building an enterprise deployment talent pipeline.
Google is rolling out Gemini 4 Argon with frontier-level benchmarks at $2/$10.
NVIDIA argues AI factory ROI hinges on three factors: earning capacity, useful life.
Google DeepMind introduced Gemini 4 Argon for coding, enterprise knowledge work and cyber defense.
Why it matters: The roundup lays out Argon's benchmark wins, its 1M-token output mechanism and its limited preview access side by side.