Try Bifrost Enterprise free for 14 days. Request access

Gateway-Level PII Redaction Before Provider Transmission

Gateway-Level PII Redaction Before Provider Transmission

TL;DR

  • Gateway-level PII redaction detects sensitive values in an LLM request and rewrites them before the request leaves your network, so no application has to sanitize its own prompts.
  • Bifrost implements this through guardrails: rules written in Common Expression Language decide when to evaluate, and profiles decide how, with input guardrails running before the provider call.
  • The built-in PII Detection template covers email addresses, US phone numbers, US Social Security numbers, credit-card-like numbers, and IPv4 addresses, evaluated in-process against Go's RE2 engine.
  • Three strategies control what replaces a match: replace keeps only the entity type, mask preserves approximate length, and hash substitutes a deterministic value you can still correlate on.
  • Three modes control where the rewrite lands, covering the live payload, Bifrost logs, and trace exports, which is what stops sensitive data from surviving in stored copies after the request itself is clean.

Sensitive information disclosure ranks second in the 2025 OWASP Top 10 for LLM Applications, covering the exposure of personally identifiable information (PII), credentials, and health records through the inputs and outputs of large language model applications. Every prompt an application sends to a third-party provider can carry names, email addresses, account numbers, and other regulated data, and once that text leaves your infrastructure you no longer control where it is logged or retained.

Gateway-level PII redaction addresses this by detecting and rewriting sensitive values in the request path before any data reaches an external provider. Bifrost, the open-source AI gateway built in Go by Maxim AI, applies redaction as a policy at the gateway, so every model call across every provider passes through the same controls. This post covers how PII exposure happens in LLM requests and how to configure redaction before provider transmission.

What Is Gateway-Level PII Redaction

Gateway-level PII redaction is the practice of detecting sensitive data in an LLM request at a central proxy and rewriting or removing that data before the request is forwarded to a model provider. Instead of relying on each application to sanitize its own prompts, redaction runs once at the gateway that sits between your applications and every provider. This makes the control uniform, auditable, and independent of the code that generated the prompt. The same placement argument, framed for audited environments, appears in PII redaction at the gateway layer for regulated industries.

The distinction that matters is placement. Application-level scrubbing depends on every team implementing the same logic correctly, and any service that skips it becomes a leak. A gateway enforces one policy on the request path that carries traffic to OpenAI, Anthropic, AWS Bedrock, Google Vertex AI, and other providers, which means a single configuration governs sensitive data handling everywhere. PII filtering and compliance at the AI gateway layer covers how that configuration maps onto compliance obligations.

Why PII Exposure in LLM Requests Matters for AI Teams

The prompt itself is a threat vector. Security analysis of LLM deployments increasingly treats the request, not just the response, as a source of sensitive-data disclosure, because prompts are logged, transmitted across services, and processed by third-party APIs. Truffle Security scanned the December 2024 Common Crawl archive, 400 terabytes drawn from 2.67 billion web pages, and found 11,908 live API keys and passwords that still authenticated successfully. Credential and PII leakage into model pipelines is a measured problem, not a hypothetical one.

Three properties of LLM request traffic make this difficult to contain:

  • Loss of control after transmission. Once a prompt reaches an external provider, retention, logging, and downstream use are governed by that provider's terms, not yours.
  • Distributed prompt assembly. Applications assemble prompts from system instructions, retrieved context, and conversation history, so PII can enter from sources the calling code never inspected. Academic work on indirect prompt injection for privacy extraction shows attackers can also engineer prompts to surface user data.
  • Regulatory scope. Names, emails, national identifiers, and health data are regulated under frameworks such as GDPR and HIPAA. Sending them to a provider without controls creates compliance exposure that is hard to remediate after the fact.

For teams running AI in production, this is where a gateway earns its place. Routing all traffic through Bifrost gives you one enforcement point where sensitive data is detected and rewritten before it crosses a network boundary, backed by the governance model that already controls access, budgets, and rate limits. Prompt injection and audit coverage sit on the same enforcement point, as LLM gateway security describes.

Where PII Leaks in the LLM Request Path

PII enters an LLM request from four places: what the user types, what a retrieval pipeline pulls in, what earlier turns of the conversation carry forward, and what the request leaves behind in logs and traces. Only the first is visible to the calling code, which is why sanitizing at the application layer misses most of the exposure.

Each source in turn:

  • Direct user input. End users paste emails, account numbers, and identifiers into chat interfaces.
  • Retrieved context. RAG pipelines pull documents that contain PII into the prompt at query time.
  • Conversation history. Multi-turn sessions accumulate sensitive values from earlier messages and replay them on every subsequent call.
  • Logs and traces. Even when the model call is acceptable, raw prompts written to logs or exported to observability backends can persist PII well beyond the request.

A redaction layer has to address both the live payload sent to the provider and the copies of that payload that land in logs and trace exports. Handling only one leaves the other exposed.

How Bifrost Redacts PII Before Provider Transmission

In the Bifrost AI gateway, redaction is implemented through guardrails that validate requests and responses in real time against policies you define. Guardrails are built around two concepts: rules, written in Common Expression Language (CEL), that determine when and what content is evaluated, and profiles, which define how content is checked. Input guardrails inspect a request before the gateway forwards it to the provider, which is exactly the enforcement point gateway-level redaction requires.

For PII specifically, the Custom Regex guardrail evaluates request and response text against patterns you define and ships with a built-in PII Detection template. The template pre-fills common patterns for:

  • Email addresses
  • US phone numbers
  • US Social Security numbers
  • Credit-card-like numbers
  • IPv4 addresses

Custom Regex runs in-process using Go's RE2-compatible engine, which means patterns cannot use lookaheads, lookbehinds, or backreferences, and each provider carries its own evaluation timeout. Because it is pattern-based rather than semantic, teams can extend it with organization-specific patterns for internal IDs, project names, or non-US national identifiers. When a pattern matches, the gateway applies the configured action: detect only, block, or redact.

Credential leakage is handled separately by the Secrets Detection guardrail, which uses embedded Gitleaks rules to catch API keys, tokens, and private keys in prompts and completions. Running PII redaction and secrets detection together on the same request covers both personal data and leaked credentials in one pass. For semantic PII classification beyond regex, Bifrost also supports Microsoft Presidio and Azure AI Language PII as guardrail profiles, and AWS Bedrock Guardrails in Bifrost walks through a third option end to end.

A typical Custom Regex profile for PII redaction is configured through the dashboard or the management API:

curl -X POST http://localhost:8080/api/guardrails/regex \
  -H "Content-Type: application/json" \
  -d '{
    "name": "pii-detection",
    "enabled": true,
    "config": {
      "timeout": 5,
      "patterns": [
        {
          "pattern": "\\b[A-Z0-9._%+-]+@[A-Z0-9.-]+\\.[A-Z]{2,}\\b",
          "description": "Email address",
          "entity_type": "EMAIL",
          "flags": "i",
          "action": "redact",
          "redaction_strategy": "replace",
          "redaction_mode": "runtime"
        }
      ]
    }
  }'

Redaction Strategies and Modes

When a guardrail's action is set to redact, the gateway rewrites matched text rather than blocking the request outright. This keeps applications functional while removing sensitive values from what the provider receives. Bifrost supports three redaction strategies that control the replacement value:

Strategy alex@example.com becomes When to use it
replace [EMAIL] The default. Keeps only the entity type, discarding everything else.
mask [EMAIL:****************] When approximate value length still matters to the reader or the model.
hash [EMAIL:8c7dd922ad47494f] When you need to correlate repeated occurrences of the same value without exposing it.

Strategies apply to the non-reversible runtime mode; the reversible modes use numbered placeholders such as [EMAIL-1] instead.

Redaction mode then decides where the rewrite applies, which is what separates a partial control from a complete one:

Mode Live request and response Bifrost logs and trace exports Original recoverable
runtime Rewritten, so sensitive text never reaches the provider Rewritten the same way No
logs_only Left raw, so the model call proceeds with the original text Rewritten with numbered placeholders In Bifrost logs only, with the Logs:Reveal permission
runtime_reversible Rewritten with numbered placeholders Rewritten with numbered placeholders In Bifrost logs only, with the Logs:Reveal permission

Only runtime and runtime_reversible keep sensitive values away from the provider. logs_only is the mode for when the model genuinely needs the original text and the stored copies are the exposure you are closing.

Because these modes cover the runtime payload, Bifrost logs, and trace-export connectors such as OpenTelemetry, Datadog, and Kafka together, redaction closes the gap where sensitive data survives in stored copies after the live request is clean. Redaction can also run on outputs. For streaming responses, runtime redaction checks buffered text segments and releases each one already rewritten, while logs-only redaction does not delay delivery at all. If the matched rule set can also block, Bifrost holds the complete stream until the final decision.

Building a Compliance-Ready Redaction Layer

Redaction is one control in a broader data-governance posture, and Bifrost pairs it with the mechanisms regulated teams need: scoped access, an accountable record of change, and a deployment boundary.

Virtual keys scope which customers and teams can send traffic, and data access control restricts what each identity can reach.

Audit logs record administrative activity as verifiable events, signed when an HMAC key is configured, with configurable retention and archiving to object storage. Guardrail interventions themselves are recorded in the request logs, which log exports can ship to a downstream system.

For organizations that cannot let request data traverse public networks at all, in-VPC deployment keeps the gateway and its logs inside private infrastructure, isolated within your VPC with no external network dependencies. Combined with redaction, this means sensitive prompts are both rewritten before provider transmission and confined to a controlled network perimeter. These capabilities are why Bifrost Enterprise is positioned for enterprises and regulated industries where data handling is audited rather than assumed.

Setting up a redaction layer follows a consistent path:

  1. Route application traffic through Bifrost as a drop-in replacement by changing the base URL in existing SDK code.
  2. Enable the Custom Regex PII Detection template and add organization-specific patterns.
  3. Add the Secrets Detection guardrail to catch leaked credentials in the same request.
  4. Set the action to redact and choose a strategy and mode that match your logging and reveal requirements.
  5. Attach the profiles to rules scoped to the inputs, outputs, or both that need protection.

The broader pattern of stopping an AI system from leaking PII or unsafe output follows the same sequence. Placing redaction at the gateway means these steps are done once rather than in every service, and the governance layer applies them uniformly across all connected applications and providers.

Frequently Asked Questions

Does redaction stop sensitive data from reaching the provider?

In runtime and runtime_reversible modes, yes: the rewrite happens on the request before the gateway forwards it, so the provider receives the placeholder rather than the value. In logs_only mode it does not, by design: that mode leaves the live payload intact and redacts only the stored copies.

Can the model still do its job on redacted text?

It depends on whether the value carries meaning the task needs. A support-ticket summary rarely needs the customer's real email, so replace costs nothing. A task that must reason about distinct people is better served by hash, which gives the same value the same placeholder, so the model can still tell two people apart without seeing either identity.

How does regex PII detection compare with semantic detection?

Regex is deterministic, runs in-process, and catches structured values such as card numbers and Social Security numbers with no external call. It does not detect names or unformatted values. Microsoft Presidio and Azure AI Language PII are available as guardrail profiles for semantic classification, and a single rule can chain both.

Can the original values be recovered after redaction?

Only in the reversible modes, and only by a user holding the Logs:Reveal permission, which reads the stored placeholder mapping in Bifrost logs. The runtime strategies are one-way: replace and mask discard the value, and hash is deterministic rather than reversible. Choose the mode from your audit and debugging requirements before enforcing.

Does redaction apply to retrieved context and conversation history?

Yes, because the guardrail inspects the assembled request rather than the code that built it. Whatever a retrieval pipeline or an earlier turn contributed is part of the payload the input guardrail evaluates, which is the reason gateway placement catches leaks that application-level scrubbing misses.

What happens to PII already sitting in old logs?

Redaction applies going forward, not retroactively. A value no guardrail detected at request time is not rewritten anywhere later, and stored copies from before a rule existed keep whatever they captured. Enable the rule in detect-only mode first to measure what your prompts contain, then decide on retention for the existing history.

Getting Started with Bifrost

Gateway-level PII redaction removes the reliance on every application to sanitize its own prompts and replaces it with one enforcement point on the request path. With Bifrost, you can detect PII with regex and semantic profiles, catch leaked secrets, and rewrite both the live payload and its stored copies before any data reaches an external provider, all governed by the same virtual keys, audit logs, and access controls that manage the rest of your AI traffic. Explore the Bifrost documentation and the governance resources to see how the pieces fit together.

To see how gateway-level PII redaction and governance work for your workloads, book a demo with the Bifrost team.