> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.asi1.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.asi1.ai/_mcp/server.

# Tuning

> Which dial to reach for when output is not what you wanted, and two worked request shapes

Most quality problems are solved by reasoning depth or by picking a different
model, not by sampling. This page is the order to work in, plus two request
shapes that show it applied.

If you are looking for the values themselves, they live with the thing they
describe: [context windows and sampling defaults](/documentation/models#sampling-defaults) on the models page,
[reasoning depth](/documentation/build-with-asi-one/reasoning) on the reasoning
page, and [rate limits](/documentation/build-with-asi-one/errors#rate-limits)
with the errors they produce.

## Where to start

1. **Match the model to the workload.** `asi1-mini` for high-volume chat,
   classification and autocomplete; `asi1` for general chat, tool calling and
   routine coding; `asi1-ultra` for deep research, code audits and long agentic
   runs. A workload on the wrong model cannot be fixed with sampling.
2. **Then set reasoning depth.** Depth moves quality by points where sampling
   moves it by impressions, so reach for
   [`reasoning_effort`](/documentation/build-with-asi-one/reasoning) first.
3. **Change sampling last, and one parameter at a time.** On `asi1` and
   `asi1-ultra` the defaults are the values the models were trained and
   evaluated at, so prefer steering through the prompt. `asi1-mini` tunes
   comfortably across a wider range.
4. **Sampling never costs you a cache hit.** `temperature`, `top_p` and
   `max_tokens` act on the output distribution, not the prompt prefix, so vary
   them freely. See [Prompt caching](/documentation/build-with-asi-one/prompt-caching).

## Symptom to fix

| Symptom                                                | Likely cause                               | Fix                                                                                                                           |
| ------------------------------------------------------ | ------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------- |
| Shallow answers, missed edge cases                     | Depth is balanced by default, not deep     | `"reasoning_effort": "high"` on `asi1` / `asi1-ultra`                                                                         |
| Reasoning too slow or too expensive for the task       | Deep reasoning where it is not needed      | `"reasoning_effort": "low"` for classification, extraction and autocomplete, or `asi1-mini` with thinking off                 |
| Coding assistant produces subtly wrong code            | Under-provisioned depth                    | `"reasoning_effort": "high"`, and `asi1-ultra` for audits and cross-file refactors                                            |
| Tool-calling agent forgets what it concluded steps ago | Prior-turn reasoning stripped from history | `"use_reasoning_history": true`, and replay `reasoning_content` verbatim                                                      |
| Multi-turn conversations slow down with cache misses   | History altered between turns              | Replay prior turns byte for byte, including reasoning. See [Prompt caching](/documentation/build-with-asi-one/prompt-caching) |

## Worked examples

### A coding assistant with tools

```json
{
  "model": "asi1",
  "messages": [],
  "reasoning_effort": "high",
  "use_reasoning_history": true,
  "stream": true
}
```

`asi1` is the right tier for routine-to-moderate coding; move to `asi1-ultra`
for audits and cross-file refactors. Deep reasoning earns its cost on code.
`use_reasoning_history` carries the agent's own conclusions across its
tool-call turns, which also keeps the prefix stable enough to cache. Sampling
stays at the defaults.

### High-volume extraction or classification

```json
{
  "model": "asi1-mini",
  "messages": [],
  "enable_thinking": false,
  "temperature": 0.2
}
```

`asi1-mini` is the cost and latency tier. Thinking is off by default here and
that is the right setting for short structured output; a low `temperature` is
comfortable on this model.

## Next steps

1. **[Reasoning](/documentation/build-with-asi-one/reasoning)** - The depth controls, per endpoint
2. **[Prompt caching](/documentation/build-with-asi-one/prompt-caching)** - How to keep a stable prefix and read the cache counters
3. **[ASI:One Models](/documentation/models)** - Which model to reach for, and the sampling defaults
4. **[Errors and Rate Limits](/documentation/build-with-asi-one/errors)** - What each status code means and what to retry