Not to be confused with semantic caching. Semantic caching is
Bifrost replaying a response it has already seen, so the provider is never called.
Prompt caching is the provider reusing the prefix of your request: the call still
happens and is still billed, but cached input is much cheaper than fresh input. The two
are independent and can both be on.
Overview
Providers such as Anthropic cache a prompt prefix only when the request marks where the cacheable region ends, using acache_control block on a message. Most SDKs let you add
that marker yourself, but agentic clients such as Codex send none at all. On Anthropic
models that means nothing is cached and every turn pays full price for a prompt that
barely changed. On providers that cache implicitly, the cached prefix slides onto the
newest message, so each turn writes a new cache entry and reads almost nothing back.
Bifrost can add the marker for the client. Turn on prompt_cache.auto_inject for a
provider and Bifrost marks the first cacheable content block of every request that
arrives without markers of its own. That block is the prefix an agent loop replays
verbatim each turn, so turn 1 writes the cache and turn 2 onward reads it.
Key properties:
- Off by default - a cache marker is a cost decision, and Bifrost never spends one the operator did not ask for.
- Caller markers always win - a request that already carries
cache_controlorprompt_cache_breakpointis forwarded unchanged. - Capability gated - a marker is only injected for models that can act on one. Implicit-caching providers are never sent a marker they would reject or ignore.
- Per-provider - configure it on each provider independently, from the provider config sheet, the management API, or
config.json. - Overridable per request - flip
auto_injectfor a single request with thex-bf-prompt-cache-auto-injectheader.
How it works
Injection runs on Chat Completions and Responses requests, including their streaming variants, and on every SDK integration route that Bifrost converts into one of those two shapes. A few rules govern what gets marked:- First cacheable block. With
auto_injectalone, Bifrost walks the messages in order and marks the first text, image, or file block it finds. A message whose content is a plain string is promoted to a single text block so the marker has somewhere to sit. The promotion is deterministic, so the cached prefix stays byte-identical across turns. - Caller markers win. If any message already carries a marker, the request is left completely alone. Injection is a default for clients that say nothing, never an override of a client that spoke.
- At most four markers. Anthropic rejects a request carrying more than four blocks with
cache_control, and every other dialect derives from that ceiling. Injection stops at four rather than relying on a downstream clamp that would silently discard the earliest marker. - Copy on write. The request Bifrost holds is never mutated. The marker is added to a copy handed to the provider, so plugins, retries, and fallbacks never see a marker the caller did not send.
- Per attempt. A fallback to a different provider re-evaluates injection against that provider’s own
prompt_cacheconfig and model capabilities. It does not inherit the previous provider’s decision.
Provider support
Bifrost injects one internal marker shape and each provider translates it to its own wire format. The capability gate is evaluated per model, not per provider, so a provider that serves both explicit-caching and implicit-caching models only injects on the former.The
Prompt Caching tab appears in the provider sheet for every provider, including
those where injection is a no-op. Saving auto_inject: true on an implicit-caching
provider is harmless: the capability gate answers false for every model, so no marker is
ever sent. The setting starts working the moment that provider gains a model that
accepts explicit markers.Configuration
Prompt caching is configured per provider. Three settings are available:- Web UI
- API
- config.json
- Go SDK
- Open Model Providers and select the provider you want to configure.
- Click Edit Provider Config to open the provider configuration sheet.
- Select the Prompt Caching tab.

- Turn on Auto-inject cache breakpoints.
- Optionally change Cache TTL from Provider default (5 minutes) to 1 hour.
- Optionally click Add injection point and set a Role, an Index, or both for each point. Adding any point replaces the default first-block strategy.

- Click Save Prompt Caching.
Injection points
cache_control_injection_points gives you precise control over which messages are
marked. It mirrors LiteLLM’s setting of the same name, so an existing LiteLLM
configuration carries over directly.
Each point has three fields:
The matching rules:
- Role and index combine with AND. A point with both matches only when the message at that index has that role.
- Role alone matches every message with that role. Four
usermessages produce four markers, which is the whole budget. - Index alone matches one message. An index past either end of the conversation matches nothing. A conversation shorter than the configured index is normal in the early turns of a session, so it is not treated as an error and no other message is marked in its place.
- A point with neither role nor index matches nothing. It is almost certainly a mistake, and marking every message would burn the whole budget.
- The last cacheable block of each match is marked, not the first. A point names a message you want cached through to its end, unlike the default strategy which names a prefix boundary.
- At most four markers are emitted, in the order the points are listed. Points beyond the budget are ignored.
- Any point replaces the default strategy.
auto_injectis still required to turn the feature on, but once at least one point is present the first-block rule no longer applies.
Per-request override
Thex-bf-prompt-cache-auto-inject header flips auto_inject for a single request.
Send true to inject on a request to a provider that has it off, or false to leave a
request alone when the provider has it on.
- Gateway (cURL)
- Go SDK
- It cannot manufacture opt-in. The header only takes effect on a provider that has a
prompt_cacheblock configured. A provider with no block is one whose operator has expressed no opinion, and a request header must not spend a cache marker or change the billing profile on their behalf. Configure{"auto_inject": false}on the provider if you want callers to opt in per request. - Only
auto_injectis overridable. The TTL and injection points stay a config-level decision. A caller that wants specific placement can send the markers itself, and caller markers always win.
Verifying that caching works
A200 response alone does not mean the cache was used. Read the usage block.
On Chat Completions, Bifrost surfaces the provider’s cache counters under
usage.prompt_tokens_details:
usage.input_tokens_details.
Expect the first request in a session to report cached_write_tokens and the following
requests to report cached_read_tokens. If every turn reports writes and no reads, the
cached prefix is changing between turns: check for a timestamp in the system prompt,
tools listed in a different order, or a client that rewrites earlier messages.
To see exactly where the marker landed, send x-bf-send-back-raw-request: true and
inspect extra_fields.raw_request. See
Request Options.
Cost considerations
Prompt caching is cheaper only when the cached prefix is read more often than it is written. Cache reads are billed well below the fresh-input rate, but the turn that writes the cache costs more than fresh input: on Anthropic, 1.25x for the default 5 minute TTL and 2x for the 1 hour TTL.- Agent loops replay the same prefix every turn, often dozens of times within a minute. This is the case the feature is built for and the savings are large.
- One-shot requests never read what they wrote. Injecting a marker there costs 25% to 100% more on the marked prefix for nothing in return. Leave
auto_injectoff on providers that only serve one-shot traffic, or use the header to opt those requests out. - A 1 hour TTL survives long pauses between turns but doubles the write cost. Use it when turns are minutes apart, such as a human-in-the-loop workflow, and stay on the default when turns are seconds apart.
Next steps
- Semantic Caching - Replay whole responses from Bifrost’s own cache instead of calling the provider.
- Request Options - Every per-request header and context key, including the prompt-cache override.
- Anthropic - Cache-control semantics for Claude, which injected markers follow.
- Bedrock - How markers become
cachePointblocks on Bedrock. - OpenAI - Explicit cache mode on the gpt-5.6 family.

