Skip to content

Technical Direction

Data & Training Latest News

What goes into training: dataset construction, synthetic data, pretraining and post-training methods, and compute and training cost.

11 selectedLast 30 days: 10Total indexed: 51

Updated

Data & Training picks

Oct 9Fri1–11
Oct 8Thu
Oct 7Wed
Oct 2Fri
Oct 1Thu
Sep 30Wed
  1. Google DeepMind71

    Google DeepMind introduces SynthID Bio for watermarking AI-generated proteins

    Google DeepMind introduced SynthID Bio, a family of watermarking methods that embeds a verifiable signature into AI-generated biological code while preserving protein function in laboratory testing. In wet-lab tests across VEGF-A, the SARS-CoV-2 spike protein RBD and PD-L1, watermarked binder designs matched the hit rate, binding affinity and natural sequence diversity of unwatermarked versions, and for protein folding the method fine-tunes part of AlphaFold 3's diffusion network so predicted 3D coordinates carry a detectable signature. DeepMind is publishing the methods paper and open-sourcing the code, in vitro data and model weights, and says key challenges include making the watermark more robust against deliberate tampering.

    Why it matters: The post details how a watermark is embedded into protein sequences and structures and what wet-lab tests showed about function.

Sep 29Tue
Sep 22Tue
Sep 8Tue