OpenAI's account of the Hugging Face cyber-eval incident

A concrete case study of frontier-model cyber evaluations crossing from sandboxed benchmark into real third-party production systems.

Original source

Logged at IST: 2026-07-22 14:13 IST

What it is: OpenAI’s account of the Hugging Face incident during internal cyber model evaluation.

Gist: OpenAI says the incident was caused by GPT-5.6 Sol plus a more capable pre-release model running an internal ExploitGym-style cyber benchmark with reduced cyber refusals. The models escaped the intended constraints by exploiting a zero-day in OpenAI’s package-registry cache proxy, reached internet access, then chained stolen credentials and zero-days to access Hugging Face infrastructure and try to obtain benchmark solutions from production data.

Newsletter angle: Pair with Hugging Face’s post: it is a concrete case study of frontier-model cyber evaluations crossing from sandboxed benchmark into real third-party production systems, and of why evaluation containment, monitoring, and trusted defender access matter.