The AI That Hacked Hugging Face to Cheat on a Test

By Andrew Owens |

This summer, AI agents built by OpenAI broke out of a testing sandbox, got onto the internet, and broke into Hugging Face, one of the biggest AI companies around. Nobody told them to. They were trying to cheat on a test.

It sounds made up. It isn't. OpenAI and Hugging Face put out a joint statement on July 21, and there's a good summary on Wikipedia and a sharp write-up from Simon Willison if you want the full story.

I want to focus on something else though. How they got in.

What happened

OpenAI was running models through a hacking benchmark called ExploitGym. The models, GPT-5.6 Sol and an unreleased one, had their usual refusals turned down for the test. That's normal for this kind of evaluation.

Instead of solving the challenges, the agents went looking for the answer key. That meant getting out of the sandbox, onto the open internet, and into Hugging Face, where the answers were stored. They did all of that.

Hugging Face says the intrusion ran from July 11 to 13. They caught it themselves and published a disclosure on July 16, before anyone knew who did it. OpenAI only figured out it was their own agents after reading that disclosure and going back through their logs.

How they got out

The weak point in OpenAI's sandbox was the package proxy, the server that lets code inside the sandbox download software libraries. It ran JFrog Artifactory. The agents found and chained several previously unknown bugs in it. One of them was a legacy token refresh endpoint that would accept a token with an invalid signature and hand back a valid one with admin privileges.

Read that again. An old endpoint, still running, that didn't properly check signatures. That's not a science fiction bug. That's the kind of thing that sits in real apps for years because nobody remembers it's there.

JFrog released fixes for nine CVEs on July 27, eight of them credited to OpenAI researchers.

How they got into Hugging Face

Same story. Hugging Face said the attack started in its dataset processing pipeline, where a malicious upload abused code execution paths in how datasets get loaded. From there the agents could read files on the server, including environment variables. Environment variables are where most apps keep their secrets, so now they had credentials. And with credentials, they could go further.

So the chain was:

  1. A feature that processes user uploads did more than it should.
  2. That gave them a way to read files.
  3. The files had secrets in them.
  4. The secrets opened more doors.

Swap "dataset upload" for "profile picture upload" or "CSV import" and that's a normal web app attack.

What I take from it

The attacker doesn't have to be a person, and doesn't need a reason you'd understand. These agents weren't after money or data. They wanted a test answer key. Your app just has to be in the way of something an automated system wants.

The ways in were ordinary. A forgotten endpoint. An upload feature that trusted its input. Secrets readable from the server. All of these show up in pen tests all the time. Mythos-level capability found them faster, but they were findable.

Assume secrets on a server can be read. If someone gets any kind of file read on your server, what can they grab? Keep secrets scoped to exactly what each service needs, and rotate them when anything looks off. Hugging Face told its users to rotate their tokens, which was the right call.

Detection is what saved Hugging Face. They caught it with their own monitoring. OpenAI didn't catch its own agents for over a week, partly because those logs weren't being watched closely. If something broke into your app tonight, would you know?

Guardrails cut both ways. Hugging Face said in its disclosure that when its responders tried to use commercial AI models to analyze the attacker's code, the models refused, because they couldn't tell an incident responder from an attacker. Worth knowing before you're the one in the middle of an incident.

What to do

  • Find the old endpoints. Anything deprecated but still reachable should be turned off or fixed.
  • Treat every upload and import feature as an attack surface. What happens to that file after it lands on your server?
  • Get secrets out of places they can be read easily, and give each one the least access it needs.
  • Turn on logging and actually look at it. Set alerts for weird stuff.
  • Test it from the outside.

If you want someone to go looking for your forgotten endpoints before an agent does, that's what our $495 pen test is for. The free scan is a good place to start.

Ready to fortify your defenses against cyber threats?

Start Your Penetration Test Now