2 articles
Frontier labs signed a round of voluntary safety commitments with little enforcement mechanism behind them. Government AI safety institutes have since started treating those same commitments as the baseline they evaluate against, which is quietly giving them the force regulation usually takes years to acquire.
The most important design work in AI is moving away from the demo and into the evaluation stack. The way labs measure models increasingly determines what the rest of us experience as product quality, safety, and trust.