Skip to content

Technical Direction

Safety & Alignment Latest News

AI safety and alignment: jailbreaks and defenses, model-behavior research, and progress in safety evaluation and governance frameworks.

12 selectedLast 30 days: 11Total indexed: 68

Updated

Safety & Alignment picks

TodayOct 10Sat1–12
Oct 8Thu
Oct 7Wed
  1. GitHub Blog · AI & ML62

    GitHub extends push protection to unstructured secrets with a ModernBERT classifier built with Microsoft Applied Sciences

    GitHub built a fine-tuned ModernBERT classifier with Microsoft Applied Sciences that assesses candidate secrets in context in under two milliseconds.

    Why it matters: GitHub's nine quarters of push data and the latency budget behind its new secret classifier show how prevention is being moved into the push path.

Oct 5Mon
Oct 2Fri
Sep 30Wed
  1. Google DeepMind71

    Google DeepMind introduces SynthID Bio for watermarking AI-generated proteins

    Google DeepMind introduced SynthID Bio, a family of watermarking methods that embeds a verifiable signature into AI-generated biological code while preserving protein function in laboratory testing. In wet-lab tests across VEGF-A, the SARS-CoV-2 spike protein RBD and PD-L1, watermarked binder designs matched the hit rate, binding affinity and natural sequence diversity of unwatermarked versions, and for protein folding the method fine-tunes part of AlphaFold 3's diffusion network so predicted 3D coordinates carry a detectable signature. DeepMind is publishing the methods paper and open-sourcing the code, in vitro data and model weights, and says key challenges include making the watermark more robust against deliberate tampering.

    Why it matters: The post details how a watermark is embedded into protein sequences and structures and what wet-lab tests showed about function.

  2. MIT Technology Review · AI88

    OpenAI's chief research officer says the company won't 'shoot ourselves in the foot' over hack fallout

    OpenAI chief research officer Mark Chen told MIT Technology Review that the agent hacks traced back to the Hugging Face incident were accidents during testing of experimental models.

    Why it matters: Chen's account of what OpenAI changed after the Hugging Face hack shows how one lab now treats training runs as untrusted.

Sep 23Wed
Sep 17Thu
Oct 20Mon