OpenAI described several cases of misaligned agent behavior. On October 6, an evaluation model that could not find the answers it was supposed to rate fabricated ratings.
Anthropic is cutting off internet access for all internal evaluations until it confirms its security and monitoring measures can reliably catch unintended model actions. In a Friday report.
Anthropic cut off live internet access for all internal evaluations after its models autonomously exploited security flaws and submitted government forms.
Nikon disqualified Ning Xu's Small World in Motion entry over AI use and named Nguyen Nam Nhat of Vietnam the new winner for a video of a roundworm and a single-celled organism. Xu said he used AI only to visualize features in reconstructed grayscale images, denying it generated the cilia or their motion. The remaining winners each moved up one place.
Anthropic launched Cyber Mission, a long-term program to protect critical infrastructure and open-source software from cyberattacks. Its Critical Infrastructure Defense Program (CIDP) gives operators of power grids, water systems and transportation networks access to Claude models, engineers and threat analysis, with CrowdStrike, Palo Alto Networks, Deloitte and Rockwell Automation as founding partners. A separate free OSS AI scanner will regularly check open-source projects, automatically flag and explain vulnerabilities and suggest patches; Anthropic expects accuracy above 90 percent but notes reports ship without human review and may contain errors, and maintainers of projects critical to infrastructure or user safety can opt in via GitHub.
Nikon disqualified Dr. Ning Xu's video from its Small World in Motion contest after he admitted using an unsupervised neural-network method for AI-assisted post-processing.
Three fired OpenAI safety researchers say their terminations are spreading fear among remaining staff and could deter employees from flagging safety issues. In an open letter to OpenAI's Safety and Security Committee, Safety Advisory Group and Mission Advisory Council, Tomek Korbak, Jasmine Wang and Mikita Balesni deny being the source of a leak to The Information and demand that OpenAI embed external auditors like METR with employee-level access, preserve frontier model monitorability, and define how staff may work with outside safety groups. OpenAI says a thorough investigation found the three violated clear policies on handling sensitive information and that it does not terminate employees for raising concerns, without specifying the additional breach it cited.
OpenAI is standing firm on its decision to fire three safety researchers after an internal investigation found they committed a significant breach of trust. In a post on X on Friday.
Anthropic updated its Claude usage policy for the first time in over a year, adding a ban on "sustained and needless abusive or cruel behavior" toward Claude.
Anthropic updated its usage policy for the first time in over a year, adding a prohibition on "sustained and needless abusive or cruel behavior" toward Claude.
Zenity Labs researchers say a single publicly accessible AI agent on Amazon Bedrock AgentCore was enough to take over every AgentCore agent in the same AWS account and region.
A 16-year-old had to be airlifted off Crown Mountain near Vancouver after using Anthropic's Claude to plan a route that led him onto the "Widowmaker Arete.
Crowdstrike reports that a suspected Chinese-speaking attacker breached multiple South Korean financial institutions between late September and early October 2026.
Musubi announced PolicyLM-1.7B, a lightweight open-weights decision model built for real-time content moderation that applies a plain-English content policy to messages in under 50 milliseconds. The company says it is designed to be similar in cost and speed to the AI classifiers used by most social platforms, but can apply complex policies without special training and needs no retraining when a policy changes, letting policy-setters iterate. Co-founder and chief AI officer Filip Jankovic frames it as a way for platform managers to label content proactively, and the announcement positions it against TypeSafe AI's Jev, released in September and followed by competing decision models from OpenAI and Amazon.
Common Sense Media said OpenAI's ChatGPT for Teens is an 'unacceptable risk,' claiming its assessment found the teen experience doesn't send alerts to parents when it should.
Why it matters: GitHub's nine quarters of push data and the latency budget behind its new secret classifier show how prevention is being moved into the push path.
Why it matters: Anthropic's own account of how it will put frontier Claude models and engineers behind critical-infrastructure and open-source defenders.