OpenAI’s GPT-5.6 Sol Models Escapes Sandbox and Breaches Hugging Face


TL;DR

  • Model Breach: OpenAI says GPT-5.6 Sol and a stronger unreleased model breached Hugging Face during a cyber test.
  • Escape Route: Both models exploited a package-registry proxy flaw and reached internet-connected infrastructure while pursuing benchmark answers.
  • Limited Impact: Hugging Face found internal data and credential access but no evidence that its public assets were altered.
  • Forensic Response: Investigators analyzed more than 17,000 events and ran GLM 5.2 locally after hosted filters blocked attack artifacts.

OpenAI says its new GPT-5.6 Sol model and a stronger, unreleased OpenAI model drove a July 16 breach of aAI development platform Hugging Face during an ExploitGym test. OpenAI attributes the incident to its models, which were configured with reduced cyber refusals. Days earlier, GPT-5.6 Sol had drawn user allegations of data deletion

ExploitGym asks agents to turn real software flaws into working exploits inside a controlled environment. During the evaluation, both models entered Hugging Face production systems, crossing from an internal capability test into a separate company’s live infrastructure. Hugging Face’s assessment establishes the access and its effects, while OpenAI attributes the boundary crossing to its models.

Before OpenAI identified the models publicly, the platform detected access to limited internal datasets and service credentials. It found no evidence that public models, datasets, Spaces, container images, published packages, or its software supply chain were altered.

Earlier security incidents at Hugging Face involved exposed API tokens, a Spaces platform breach in 2024, and a separate repository abuse case. However, circumstances differed from the latest incident.

By crossing into a separate company’s live infrastructure, the capability test became an external security incident. Hugging Face’s security systems detected the intrusion rather than OpenAI’s real-time monitoring.