Prompt Caching
ASI:One runs prefix caching in front of every model. When a request repeats the beginning of an earlier one, the shared part is served from cache instead of being processed again, which cuts latency on workloads that reuse a system prompt, a tool list or a conversation history.
Caching is automatic. There is nothing to enable, and nothing to send.
Reading the counters
The response usage tells you what the cache did.
On /v1/responses the same
figure is input_tokens_details.cached_tokens.
The two fields answer different questions, which is what makes a cold cache easy to tell apart from one that is not working:
created_cache_tokenscounts tokens written to the cache. It is high on the first call of a new prefix and drops to zero once that prefix is cached. This field is specific to/v1/chat/completions;/v1/responsesreports the read counter only.cached_tokenscounts tokens read from the cache. It is zero on the first call and rises on the ones that follow.
So a first call reporting created_cache_tokens above zero is working
correctly, even though nothing was read yet.
cached_tokens will not match prompt_tokens. The cache works in blocks, so
a prefix is cached up to a block boundary and the remainder is processed
normally. A ratio well below 1.0 on a stable prefix is the expected result,
not a misconfiguration.
Very short prompts may cache nothing at all: a prefix has to be long enough to fill a block before there is anything to store.
What invalidates the cache
The cache matches on an exact prefix, so anything that changes the start of the request starts a new one:
- Any edit to the prefix, down to a single character, a whitespace change or a timestamp embedded in the system prompt. Put stable content (system prompt, tool definitions, few-shot examples) first and variable content last.
- Editing or reordering past turns. Replay conversation history verbatim. If you need to compress an old turn, do it once and keep that form stable.
- Changes to the tool list. Tool schemas are part of the prompt, so adding, removing or renaming a tool invalidates the cache from that point on.
- Dropping reasoning between turns on the models that return it. Replay
reasoning_contentalongsidecontent, or set"use_reasoning_history": trueso the model sees on turn N+1 what it produced on turn N. See Reasoning.
Sampling parameters never invalidate the cache. temperature, top_p and
max_tokens act on the output distribution, not on the prompt, so you can vary
them freely without losing hits.
Next steps
- Tuning - Which dial to reach for when output is not what you wanted
- Reasoning - Carrying reasoning across turns
- Chat Completions API - The default endpoint