A coding agent sends its working set: the files it read, the shell output it saw, the .env it opened to debug a connection, the customer CSV it grepped. It resends all of it on every turn.
Masking sensitive values in that stream is easy to describe: swap them out before the request leaves, swap them back on the way in. The challenge is that the agent still has to write the real password into the config file, the response stream must not stall, and the provider's prompt cache has to keep hitting, or every turn gets slower and more expensive.
In June we introduced what the OUTGATE.ai gateway does. This post is about how its anonymization works, what broke on the way, and what it still cannot do.
Keep in mind that anonymization is not purely about trusting your service provider, but about keeping one more pile of audit logs, the one on your provider's side, free of your most sensitive data and credentials.
The idea: reversible placeholders
Every value the policy marks as sensitive is replaced with a placeholder such as OG_CREDENTIAL_4be1f07a93d2c58e or OG_PII_d81c3a6f0e297b45. The label says what kind of value it was. The 16 hex characters are derived from the value, so the same value always gets the same placeholder.
Say an agent debugging a failed migration runs cat .env. The tool output it sends contains:
DATABASE_PASSWORD=Tq7v-mR2x-Lp9w-Hc4n
ONCALL_EMAIL=maria.keller@example.comThe provider receives:
DATABASE_PASSWORD=OG_CREDENTIAL_4be1f07a93d2c58e
ONCALL_EMAIL=OG_PII_d81c3a6f0e297b45When the model tests the connection, it writes the placeholder into its tool call. The agent receives the real value and runs that:
model writes: PGPASSWORD=OG_CREDENTIAL_4be1f07a93d2c58e psql -h db -U billing -c 'select 1'
agent runs: PGPASSWORD=Tq7v-mR2x-Lp9w-Hc4n psql -h db -U billing -c 'select 1'The model learns the convention from a short note the gateway adds to the request:
Values like OG_CREDENTIAL_a3f291 and OG_PII_b8c042 are secure placeholders for real data. Use them exactly as-is in your response — they will be transparently swapped back to the original values before the user sees your reply. For tool calls, pass them as arguments normally; the actual credentials/data will be injected before execution.
In GDPR terms this is pseudonymization rather than anonymization, because a way back exists. That way back lives only in the region that carries the traffic, which we can host or you can run on your own hardware.
The round trip
The gateway masks the request before it leaves the region and restores the stream on its way back, so the map from placeholder to value never reaches the provider.
Finding what to mask
Masking starts with finding, and in agent traffic sensitive values hide in odd places. The gateway reads each request in the provider's own wire format and pulls out the content: system prompt, messages, tool results, tool-call arguments and reasoning summaries.
That list is longer than it looks. When we tested Codex, its tool calls lived in Responses API item types we did not read yet, and 37 percent of the characters in one captured request never reached detection. Unknown item types are now serialized and scanned instead of skipped.
Two detectors then run side by side:
- The vault holds values the organization has seen before, plus values a team always wants masked. It matches them token by token, so a key found in message 1 stays masked in message 50 even if the classifier misses it there. An auto-detected value keeps being masked as long as it keeps showing up, on a one-hour sliding window by default.
- A detector finds new values. Our hosted region runs the open OpenAI Privacy Filter classifier, merged with Microsoft Presidio; a region can also use a general-purpose LLM.
Classifiers are slow next to a proxy, so the detector only sees what is new. Each message block is hashed and its verdict cached for up to a week, so in a long conversation only the latest turn is scanned. On our GPU-backed test region the scan added 36 ms at the median and 408 ms at the 90th percentile, and it did not grow over 40-turn conversations.
Screenshots take the same path. With an OCR model configured, each image is transcribed and the text is scanned like typed text. An image whose text contains a masked value is withheld, and the model gets the masked transcript with a note instead. An image that cannot be read is withheld too; a clean image is forwarded unchanged.
Raw detector output is not trusted as is. A span like AWS_SECRET_ACCESS_KEY=<key> is trimmed to the key, so the model still sees which variable it is reading. Teams can mark values as never sensitive, and single everyday words that a classifier mistakes for names can be filtered out against a dictionary.
Getting the real values back
The model only ever sees placeholders, so its output can only contain placeholders. The gateway keeps the request's map in memory for the life of that request and rewrites the response on its way to the client.
Streaming makes that more than a find-and-replace. A placeholder can arrive split across two server-sent events, OG_PI in one and the rest in the next. The filter holds back any tail that could start a placeholder until the next event settles it, then releases it as a well-formed event of the same type.
An early version got both parts wrong. It flushed held text as malformed events, which the client dropped along with an end-of-block marker. It also rescanned the whole output on every event, so long tool calls stalled the stream.
Tool calls are where restoring pays off. Their arguments stream as fragments of JSON and are restored the same way, so the Write or Bash call the agent executes carries the real value. The tools run on the client or in its sandbox, so real values exist only there and in the region.
Models also mangle placeholders now and then. A second pass restores a placeholder whose letter case changed; any other rewrite is counted and logged.
Placeholders can also outlive the request that created them. The model may repeat one it saw many turns ago, after the value has left the current request's map, and the user would then get the placeholder instead of the value. Once that ended up in a file path the agent wrote to disk.
So every placeholder is also kept in a registry for 30 days, encrypted with a key that never leaves the region and scoped to the gateway key that created it. The registry only feeds the response. What the provider sees stays masked.
Same value, same placeholder, and why the cache cares
Placeholders are deterministic. The hex is an HMAC of the value, keyed per organization and per gateway key, so a key always gets the same placeholder for the same value, across turns and threads. A different key gets a different placeholder, so a colleague cannot confirm a guess by sending a value through their own key and comparing placeholders in a shared log.
Determinism lets the model follow one placeholder through a long task. It matters even more for the prompt cache. Providers cache the prefix of a conversation, and a cached prefix is faster and much cheaper to process. Change one byte in an early message and everything after it is processed again at full price.
We avoid that with what we call stable context. The first time the gateway sees a message block, it freezes the set of values masked inside it, and later requests reuse that set for that block. A value discovered later is masked in new messages only. In a controlled test, the turn after a new value entered the vault still read all 9,196 tokens from cache, and in a 40-turn soak each turn added only its own new tokens to the cache.
The trade-off is deliberate. A value first detected in turn 20 stays as it was in turns 1 to 19. Those bytes already went to the provider on earlier turns, so masking them now would cost the cache without taking anything back. Teams that want everything re-masked anyway can turn stable context off in the policy.
What it cannot do
- Detection is probabilistic. A detector can miss a value, which is what the vault is for, and it flags harmless strings: git hashes and checksums look like secrets, first names look like people. The ignore list, the dictionary filter and manual vault entries are how teams tune it.
- A placeholder carries no meaning. The model cannot check a masked IBAN's format or spot a typo in a masked key. Categories can be set to detect-only where the model needs the real value.
- Small models mangle placeholders. Case changes are restored; other rewrites are logged and stay unresolved.
- Images are only as safe as OCR. A transcript that misses a value lets its image through.
Turning it on
Anonymization is part of a guardrail policy on a provider in the Console. Personal information and credentials are masked by default. Other categories can be detect-only or blocking, and images are scanned when the region has an OCR model.
To see what a policy finds before real traffic does, og scan runs the same pipeline in dry-run mode: nothing reaches the model, and what it finds goes into the vault. The guardrails docs cover policies, the vault and the CLI.
If your agents touch real systems, secrets will reach the prompt. The question is whether they reach the provider.
