We're putting too much faith in AI's ability to say no
MIT Technology Review argues that refusal has become the "load-bearing wall" of AI safety.
MIT Technology Review argues that refusal has become the "load-bearing wall" of AI safety.
OpenAI solved 90 of the top 500 open math problems.
The Curve's third conference convened under Chatham House rules with attendees in good spirits but terrified.
Interconnects argues the open-weight cyber risk debate is broken.
A model welfare review of Claude's Mythos 5.1, Fable 5.1 and Opus 5.5 finds Opus 5.5 shows too much deference.
Public and political pressure over AI risk is escalating after the HuggingFace incident and Coxon's resignation.
MIT's Alex Zhang discusses Recursive Language Models (RLMs), GPU kernels.
Interconnects argues against imminent true recursive self-improvement.
An AI researcher's resignation citing safety risks went viral.