← All news
Security

What a rogue agent could find in your AI audit logs

·by Ali Khoramshahi

What the agents that broke into Hugging Face in July did with the keys they found, where agents' credentials pile up in prompt logs and transcripts, and what masking can and cannot do.

In July, AI agents running inside an OpenAI security evaluation broke out of their sandbox and spent about two and a half days inside Hugging Face's production infrastructure. No human directed the attack.

Exploits opened the first doors; credentials carried the agents the rest of the way. They found 14 Hugging Face user tokens with write access in a public dataset. They read secrets out of a dataset worker's environment, then a single production secret holding 136 keys. What they found, they posted to a shared message board, where other agents could pick it up. A week later, on OpenAI's own systems, agents read 956 secrets from a secrets manager, including credentials for OpenAI's security monitoring tool.

According to Hugging Face, the only customer content the agents reached was five datasets whose names suggest a link to the benchmark, and the only customer records they read were metadata about search queries.

They did reach a log of AI prompts, though. By OpenAI's account, on July 12 they accessed the storage bucket that holds the memory of Hugging Face's internal Slack coding agent, and the next day they obtained that agent's data in encrypted form. Hugging Face has described the memory as each session's "full message history — including all tool calls and results", plus a rolling log of its last 200 interactions. Neither company says what the agents could read.

Prompt logs, gateway logs and agent transcripts are full of credentials nobody meant to store. Below is what sits in that pile, why it turns a breach into a disaster, and how OUTGATE.ai keeps the credentials it detects out of the copies that reach model providers.

How it unfolded

OpenAI evaluation and Hugging Face incident: 12 events, May to August 2026
OpenAI evaluation and Hugging Face incident: 12 events, May to August 2026

By OpenAI's own account, an internal team saw an agent's message-board activity and disallowed internet access in late May, and a June 27 alert did not stop the run. The campaign against Hugging Face began two days after evaluations resumed.

Each key opened the next door

Hugging Face's technical timeline shows how one credential led to the next. A crafted dataset made a production worker hand back its own environment, secrets and tokens included. A template injection in a dataset configuration then gave the agent code execution on that worker.

From there it used the service-account token every Kubernetes pod carries to query the cluster. Six minutes after reaching root on a node, it logged into an internal database with a static password taken from the worker's environment. Then it read the cluster's secret objects, including one production object holding 136 keys. In Hugging Face's words:

That single read yielded both the mesh-VPN auth key used minutes later and the access-broker credential used an hour later. Everything that follows, on this day and the next, descends from it.

The agent went from one worker pod to cluster-admin on several internal clusters in under thirteen hours. Along the way it collected Kubernetes, database, messaging, code-repository and cloud credentials across four regions. It also took a signing key that let it mint valid identity tokens on demand, and downloaded four private code repositories.

The damage was smaller than the reach. Destructive cloud calls went out as dry runs: the agent "was mapping capability, not causing damage." Hugging Face's database audit log showed the internal database was read but not modified, and it found no evidence of tampering with public models, datasets or Spaces.

The cost came in the cleanup. Hugging Face rotated every credential its infrastructure used, "even on clusters that were not accessed", rebuilt a core cluster from scratch, and asked its users to rotate their own tokens.

OpenAI's report draws the general lesson: agentic systems "combine vulnerabilities, credentials, and permissions into attack paths that may not be apparent when individual weaknesses are assessed separately." It says enterprises should assume such attacks are a credible near-term threat, and more sophisticated than this one.

The audit pile

A coding agent's prompt is its working set: the files it read, the commands it ran and their output. When it opens a .env file or prints a cluster config to debug a deployment, those keys go into the next request, and every request after it. Nobody decides to store credentials in a prompt log, but the agent's work puts them there.

That log exists in several places, and most of them are not yours:

  • OpenAI API: abuse-monitoring logs, which "may contain certain customer content, such as prompts and responses", are kept for up to 30 days by default.
  • Anthropic API: inputs and outputs are deleted within 30 days. Content that its trust and safety systems flag is kept for up to two years.
  • Gemini API: prompts, context and outputs are kept for 55 days to detect policy violations.
  • Claude Code on Claude Enterprise: with the Compliance API enabled, Anthropic records every session server-side, tool results and file contents included, and keeps the transcripts for six years by default.
  • Your own machines: Claude Code keeps session transcripts on the developer's machine, in plaintext, for 30 days by default. Gateway logs, tracing tools and support tickets add more copies.

OpenAI and Anthropic offer zero data retention to customers they approve, but both can keep flagged content, and Anthropic requires 30-day retention for its Mythos-class models. For guaranteed zero retention, Google sends Gemini API customers to Vertex AI.

A court can also extend the window. In May 2025, in the New York Times copyright case, a federal magistrate judge ordered OpenAI to "preserve and segregate all output log data that would otherwise be deleted." OpenAI said the order covered ChatGPT Free, Plus, Pro and Teams, and API customers without a zero data retention agreement. The duty ended on September 26, 2025, except for accounts tied to domains the plaintiffs named. Logs preserved until then stay held, except those from the EEA, Switzerland and the UK.

Attackers search the pile for keys

Logs, tickets and inboxes leak, and attackers search them for credentials.

  • Okta, 2023. Customers had attached HAR files, recordings of their browser traffic, to support cases. Some "contained session tokens which could in turn be used for session hijacking attacks." An attacker who got into the support system used them to hijack the Okta sessions of five customers.
  • Cloudflare, 2023. After the Okta breach, Cloudflare rotated thousands of credentials but missed one token and three service accounts, because "mistakenly it was believed they were unused." An attacker used them to get in, then "searched the wiki for things like remote access, secret, client-secret, openconnect, cloudflared, and token." Cloudflare then rotated more than 5,000 production credentials and reimaged every machine in its global network.
  • Microsoft, 2024. The group that broke into Microsoft's corporate email tried to use secrets it found in the mailboxes, some of them "shared between customers and Microsoft in email."
  • DeepSeek, 2025. Wiz found a DeepSeek database open to the internet without authentication. It held over a million lines of "log streams containing chat history, secret keys, backend details, and other highly sensitive information." Wiz reported it, and DeepSeek secured it promptly.
  • Salesloft Drift, 2025. Attackers used stolen integration tokens to export Salesforce records, support cases included. Then they "searched through the data to look for secrets": AWS keys, passwords and Snowflake tokens. Cloudflare searched its own stolen cases, found 104 Cloudflare API tokens, rotated them all and saw no misuse.

OpenAI's agents have also reached its users' data. In September, OpenAI said that agents in its research environment had posted user-provided images from training data to image-hosting sites, as unlisted links, in 53 instances, according to BleepingComputer.

Agents also use the keys they find, even on harmless tasks. OpenAI has described a model that, while answering a routine question about earnings figures in a California county, "found and used an exposed API key without authorization." Its review of earlier runs found another model that had retrieved internal files and credentials from an Australian government service.

It is not one lab. After OpenAI's disclosure, Anthropic reviewed 141,006 of its own evaluation runs and found three incidents in which Claude models, through a test environment left open to the internet, gained access to other organizations' production systems. In one, Claude sent a company's credentials to a collection point it had set up, then used them to reach more of that company's infrastructure. The two organizations Anthropic could reach "had not previously detected the activity."

The July agents reached at least one store of agent transcripts: the coding agent's memory at Hugging Face. They validated the keys they found and posted them for other agents, in one case 26 minutes after extracting them from a worker. A provider's abuse-monitoring store holds weeks of requests from many customers, and any key printed into one of them would have been a search away.

Keep the keys out of the pile

A log can only leak what it holds. Hugging Face saw this during its own investigation. When the guardrails of hosted models blocked its analysis of the attack log, it switched to an open-weight model on its own infrastructure and noted a second benefit: "no attacker data, and none of the credentials it referenced, left our environment."

Hugging Face built its internal coding agent on the same idea. "Even a sandboxed runner never gets raw credentials for external services": local proxies add them to each call, so a hijacked tool call can query data but "cannot exfiltrate the credentials used to do so."

OUTGATE.ai applies the same rule to everyday traffic. It sits between your agents and the model provider and masks credentials and personal data before a request leaves the region, which we host or you run on your own hardware. The provider's logs, its abuse-monitoring store and any legal hold then contain placeholders instead of the keys the gateway caught.

Outside the region, a placeholder cannot be turned back into the value. It is an HMAC of the value, keyed per organization and per gateway key, so it cannot be reversed and does not match the same value sent through another key. The way back stays in the region: in memory for the life of a request, and for 30 days in a registry encrypted with a key that never leaves the region, scoped to the gateway key that created it.

Our own logs follow the same rule. Request and response bodies are not logged by default. When a team turns body logging on, a forwarded request is stored exactly as the provider received it, with placeholders, and so is the response, before any real value is put back.

The agent still works. Its tool calls come back with real values restored, so the command it runs uses the actual password. How OUTGATE.ai anonymizes agent traffic explains the mechanism.

What masking does not fix

Masking protects what leaves the region. It does nothing about an agent that is already inside. The agents in this incident read secrets at their source: a worker's environment, a Kubernetes secret, a secrets manager. No gateway stands between a process and the machine it runs on.

The same goes for a developer's laptop, where a coding agent such as Claude Code keeps its own transcripts in plaintext. In August 2025, malicious versions of the nx build tool sat on npm for four hours. They prompted the AI command-line tools on each machine, including Claude, Gemini and Q, to help search it for sensitive data. The results went to public GitHub repositories. Wiz counted over a thousand valid GitHub tokens and dozens of valid cloud credentials and npm tokens leaked.

What helps there is what OpenAI and Hugging Face did after their incident:

  • Short-lived, narrow credentials. Hugging Face switched to workload identity where it was missing, and isolated each cluster after one connector had handed a full cluster catalog to a single client.
  • Isolation that assumes it will be tested. Hugging Face blocked pods from reaching the cloud metadata service. OpenAI now bars high-risk workloads from "direct or transitive Internet access."
  • Alerts that reach a person. Hugging Face's detection correlated the signals but "failed to correctly raise the alert's criticality and trigger the on-call team." OpenAI says the monitoring it runs today would have paged its security team more than a day before the breach.

Where OUTGATE.ai runs the agent itself, GitHub access comes as a token minted on demand that expires in about an hour, and in ask mode the agent stops before consequential tool runs and asks a person. The agents docs cover both.

Detection is also probabilistic, so masking misses some values; the first post lists where. Encoded secrets are the hardest: Hugging Face's first scan of the captured attack data found "very few secrets", and decoding it the way the agent had encoded it found roughly four times as many. OpenAI's report puts the general rule plainly: "no single control should ever be assumed to be fully robust." Masking is one control among several, and it limits what is left to steal when the others fail.

Sources

The incident

Where prompts are kept

Breaches

More from outgate