Skip to navigation

Prompt Caching

ASI:One runs prefix caching in front of every model. When a request repeats the beginning of an earlier one, the shared part is served from cache instead of being processed again, which cuts latency on workloads that reuse a system prompt, a tool list or a conversation history.

Caching is automatic. There is nothing to enable, and nothing to send.

Reading the counters

The response usage tells you what the cache did.

On /v1/chat/completions:

"prompt_tokens_details": {
"cached_tokens": 1920,
"created_cache_tokens": 0
}

On /v1/responses the same figure is input_tokens_details.cached_tokens.

The two fields answer different questions, which is what makes a cold cache easy to tell apart from one that is not working:

  • created_cache_tokens counts tokens written to the cache. It is high on the first call of a new prefix and drops to zero once that prefix is cached. This field is specific to /v1/chat/completions; /v1/responses reports the read counter only.
  • cached_tokens counts tokens read from the cache. It is zero on the first call and rises on the ones that follow.

So a first call reporting created_cache_tokens above zero is working correctly, even though nothing was read yet.

cached_tokens will not match prompt_tokens. The cache works in blocks, so a prefix is cached up to a block boundary and the remainder is processed normally. A ratio well below 1.0 on a stable prefix is the expected result, not a misconfiguration.

Very short prompts may cache nothing at all: a prefix has to be long enough to fill a block before there is anything to store.

What invalidates the cache

The cache matches on an exact prefix, so anything that changes the start of the request starts a new one:

  • Any edit to the prefix, down to a single character, a whitespace change or a timestamp embedded in the system prompt. Put stable content (system prompt, tool definitions, few-shot examples) first and variable content last.
  • Editing or reordering past turns. Replay conversation history verbatim. If you need to compress an old turn, do it once and keep that form stable.
  • Changes to the tool list. Tool schemas are part of the prompt, so adding, removing or renaming a tool invalidates the cache from that point on.
  • Dropping reasoning between turns on the models that return it. Replay reasoning_content alongside content, or set "use_reasoning_history": true so the model sees on turn N+1 what it produced on turn N. See Reasoning.

Sampling parameters never invalidate the cache. temperature, top_p and max_tokens act on the output distribution, not on the prompt, so you can vary them freely without losing hits.

Next steps

  1. Tuning - Which dial to reach for when output is not what you wanted
  2. Reasoning - Carrying reasoning across turns
  3. Chat Completions API - The default endpoint