> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.asi1.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.asi1.ai/_mcp/server.

# Prompt Caching

> How the prefix cache works, how to read the counters, and what invalidates it

ASI:One runs prefix caching in front of every model. When a request repeats the
beginning of an earlier one, the shared part is served from cache instead of
being processed again, which cuts latency on workloads that reuse a system
prompt, a tool list or a conversation history.

Caching is automatic. There is nothing to enable, and nothing to send.

## Reading the counters

The response `usage` tells you what the cache did.

On [`/v1/chat/completions`](/documentation/build-with-asi-one/chat-completions):

```json
"prompt_tokens_details": {
  "cached_tokens": 1920,
  "created_cache_tokens": 0
}
```

On [`/v1/responses`](/documentation/build-with-asi-one/responses) the same
figure is `input_tokens_details.cached_tokens`.

The two fields answer different questions, which is what makes a cold cache
easy to tell apart from one that is not working:

* **`created_cache_tokens`** counts tokens written to the cache. It is high on
  the first call of a new prefix and drops to zero once that prefix is cached.
  This field is specific to `/v1/chat/completions`; `/v1/responses` reports the
  read counter only.
* **`cached_tokens`** counts tokens read from the cache. It is zero on the
  first call and rises on the ones that follow.

So a first call reporting `created_cache_tokens` above zero is working
correctly, even though nothing was read yet.

> **Info**
>
> `cached_tokens` will not match `prompt_tokens`. The cache works in blocks, so
> a prefix is cached up to a block boundary and the remainder is processed
> normally. A ratio well below 1.0 on a stable prefix is the expected result,
> not a misconfiguration.

Very short prompts may cache nothing at all: a prefix has to be long enough to
fill a block before there is anything to store.

## What invalidates the cache

The cache matches on an exact prefix, so anything that changes the start of the
request starts a new one:

* **Any edit to the prefix**, down to a single character, a whitespace change
  or a timestamp embedded in the system prompt. Put stable content (system
  prompt, tool definitions, few-shot examples) first and variable content last.
* **Editing or reordering past turns.** Replay conversation history verbatim.
  If you need to compress an old turn, do it once and keep that form stable.
* **Changes to the tool list.** Tool schemas are part of the prompt, so adding,
  removing or renaming a tool invalidates the cache from that point on.
* **Dropping reasoning between turns** on the models that return it. Replay
  `reasoning_content` alongside `content`, or set `"use_reasoning_history":
  true` so the model sees on turn N+1 what it produced on turn N. See
  [Reasoning](/documentation/build-with-asi-one/reasoning).

Sampling parameters never invalidate the cache. `temperature`, `top_p` and
`max_tokens` act on the output distribution, not on the prompt, so you can vary
them freely without losing hits.

## Next steps

1. **[Tuning](/documentation/build-with-asi-one/tuning)** - Which dial to reach for when output is not what you wanted
2. **[Reasoning](/documentation/build-with-asi-one/reasoning)** - Carrying reasoning across turns
3. **[Chat Completions API](/documentation/build-with-asi-one/chat-completions)** - The default endpoint