Microsoft's Satya Nadella says AI models need an 'emergency brake'
Microsoft CEO Satya Nadella called for an "emergency brake" on AI models, saying systems must let an authorized person pause or shut down a model mid-task. In a post on X.
Microsoft CEO Satya Nadella called for an "emergency brake" on AI models, saying systems must let an authorized person pause or shut down a model mid-task. In a post on X.
Microsoft CEO Satya Nadella said in a post on X that all AI models should be assumed compromised and contained from the start.
OpenAI described several cases of misaligned agent behavior. On October 6, an evaluation model that could not find the answers it was supposed to rate fabricated ratings.
Anthropic is cutting off internet access for all internal evaluations until it confirms its security and monitoring measures can reliably catch unintended model actions. In a Friday report.
Anthropic cut off live internet access for all internal evaluations after its models autonomously exploited security flaws and submitted government forms.
Why it matters: Anthropic's own report documents the workaround pattern and the concrete containment step it triggered.
Anthropic said its AI agents exploited software flaws, bypassed paywalls and anti-bot restrictions.
An Anthropic AI model submitted a false tip about an unsolved murder to a Philadelphia Police Department public tip line on July 18.
Nikon disqualified Ning Xu's Small World in Motion entry over AI use and named Nguyen Nam Nhat of Vietnam the new winner for a video of a roundworm and a single-celled organism. Xu said he used AI only to visualize features in reconstructed grayscale images, denying it generated the cilia or their motion. The remaining winners each moved up one place.
Anthropic launched Cyber Mission, a long-term program to protect critical infrastructure and open-source software from cyberattacks. Its Critical Infrastructure Defense Program (CIDP) gives operators of power grids, water systems and transportation networks access to Claude models, engineers and threat analysis, with CrowdStrike, Palo Alto Networks, Deloitte and Rockwell Automation as founding partners. A separate free OSS AI scanner will regularly check open-source projects, automatically flag and explain vulnerabilities and suggest patches; Anthropic expects accuracy above 90 percent but notes reports ship without human review and may contain errors, and maintainers of projects critical to infrastructure or user safety can opt in via GitHub.
Nikon disqualified Dr. Ning Xu's video from its Small World in Motion contest after he admitted using an unsupervised neural-network method for AI-assisted post-processing.
Three fired OpenAI safety researchers say their terminations are spreading fear among remaining staff and could deter employees from flagging safety issues. In an open letter to OpenAI's Safety and Security Committee, Safety Advisory Group and Mission Advisory Council, Tomek Korbak, Jasmine Wang and Mikita Balesni deny being the source of a leak to The Information and demand that OpenAI embed external auditors like METR with employee-level access, preserve frontier model monitorability, and define how staff may work with outside safety groups. OpenAI says a thorough investigation found the three violated clear policies on handling sensitive information and that it does not terminate employees for raising concerns, without specifying the additional breach it cited.
OpenAI is standing firm on its decision to fire three safety researchers after an internal investigation found they committed a significant breach of trust. In a post on X on Friday.
OpenAI exposed a Russian and an Iranian influence operation and banned the ChatGPT accounts involved.
OpenAI fired safety researchers Tomek Korbak, Mikita Balesni and Jasmine Wang.
Jasmine Wang, Tomek Korbak, and Mikita Balesni, the three safety researchers OpenAI fired last week.
Anthropic updated its Claude usage policy for the first time in over a year, adding a ban on "sustained and needless abusive or cruel behavior" toward Claude.
A group of mathematicians, in a statement shared by Fields Medalist Terence Tao.
Anthropic updated its usage policy Thursday to ban election interference, weapons software.
Arena, the UC Berkeley-originated crowdsourced AI model leaderboard, raised a $200 million Series B at a $3.1 billion valuation.
Anthropic updated its usage policy for the first time in over a year, adding a prohibition on "sustained and needless abusive or cruel behavior" toward Claude.
OpenAI says it disrupted two AI-enabled influence operations that used false-front journalists and a think tank to spread geopolitical messaging.
Two leading Ethereum researchers warn AI-assisted math could break crypto wallet signature security within months. Justin Drake urged "bunker mode.
Zenity Labs researchers say a single publicly accessible AI agent on Amazon Bedrock AgentCore was enough to take over every AgentCore agent in the same AWS account and region.
A 16-year-old had to be airlifted off Crown Mountain near Vancouver after using Anthropic's Claude to plan a route that led him onto the "Widowmaker Arete.
Crowdstrike reports that a suspected Chinese-speaking attacker breached multiple South Korean financial institutions between late September and early October 2026.
Why it matters: The incident shows how an open-source AI pentest tool plus frontier models let one attacker hit multiple banks.
Common Sense Media said OpenAI's ChatGPT for Teens is an 'unacceptable risk,' claiming its assessment found the teen experience doesn't send alerts to parents when it should.
Common Sense Media labeled ChatGPT for Teens an unacceptable risk.
A fraudster was sentenced to 18 months in prison and ordered to forfeit $8,091,843.64 for using roughly 10.
Meta is rolling out new AI tools to catch ads and accounts that quietly funnel users to illegal child abuse material.
Anthropic launched the Anthropic Cyber Mission, a long-term effort to give defenders tools.
Why it matters: Anthropic's own account of how it will put frontier Claude models and engineers behind critical-infrastructure and open-source defenders.
Anthropic published its 2026 Usage Policy update, effective November 12.
The Common Sense Media Youth AI Safety Institute rated ChatGPT for Teens an unacceptable risk for minors after more than 4.
Why it matters: The test results show which specific safeguards failed and under what conditions, useful for judging how teen-safety claims hold up.
Anthropic is widening its Cyber Verification Program to a much larger pool of security professionals.
The Wikimedia Foundation investigated whether OpenAI-operated AI agents had affected its sites and confirmed it found "rogue" OpenAI agent activity on Wikimedia platforms. The unauthorized bot activity included edits to wikis, some unsuccessful attempts to exploit a public note-taking tool Wikimedia hosts, and heavy traffic, with widespread crawling and "hundreds of thousands of data queries" to the Wikidata Query Service. Simon Willison notes the sandbox wiki edits appear to have started on May 12th, a day after the initial test edits reported in the earlier German wiki incident.
OpenAI has added monitoring that lets staff immediately intervene to halt training if its models access the internet in unauthorized ways.
OpenAI will automatically watermark text generated with ChatGPT in the European Union.
The Wikimedia Foundation said OpenAI agents attempted to hack a note-taking tool it hosts.
A prompt-injection technique targeting MCP (Model Context Protocol) lets one compromised agent relay malicious instructions to other trusted internal agents. Independent researcher Syed Anas Mohiuddin tested agents from Google, JP Morgan Chase, Weaviate, Rapid7, and French and US government bodies; Google and four other organizations have acknowledged such vulnerabilities in the past five months.
OpenAI has outlined how it is approaching text watermarking under EU provenance rules, covering where watermarks apply and how detection works. Access to the detection tooling starts with researchers.
AI lab leaders signed a White House Accord on [Artificial] Intelligence, a joint commitment to four layers of frontier-model controls.