Skip to navigation

Tuning

Most quality problems are solved by reasoning depth or by picking a different model, not by sampling. This page is the order to work in, plus two request shapes that show it applied.

If you are looking for the values themselves, they live with the thing they describe: context windows and sampling defaults on the models page, reasoning depth on the reasoning page, and rate limits with the errors they produce.

Where to start

  1. Match the model to the workload. asi1-mini for high-volume chat, classification and autocomplete; asi1 for general chat, tool calling and routine coding; asi1-ultra for deep research, code audits and long agentic runs. A workload on the wrong model cannot be fixed with sampling.
  2. Then set reasoning depth. Depth moves quality by points where sampling moves it by impressions, so reach for reasoning_effort first.
  3. Change sampling last, and one parameter at a time. On asi1 and asi1-ultra the defaults are the values the models were trained and evaluated at, so prefer steering through the prompt. asi1-mini tunes comfortably across a wider range.
  4. Sampling never costs you a cache hit. temperature, top_p and max_tokens act on the output distribution, not the prompt prefix, so vary them freely. See Prompt caching.

Symptom to fix

SymptomLikely causeFix
Shallow answers, missed edge casesDepth is balanced by default, not deep"reasoning_effort": "high" on asi1 / asi1-ultra
Reasoning too slow or too expensive for the taskDeep reasoning where it is not needed"reasoning_effort": "low" for classification, extraction and autocomplete, or asi1-mini with thinking off
Coding assistant produces subtly wrong codeUnder-provisioned depth"reasoning_effort": "high", and asi1-ultra for audits and cross-file refactors
Tool-calling agent forgets what it concluded steps agoPrior-turn reasoning stripped from history"use_reasoning_history": true, and replay reasoning_content verbatim
Multi-turn conversations slow down with cache missesHistory altered between turnsReplay prior turns byte for byte, including reasoning. See Prompt caching

Worked examples

A coding assistant with tools

{
"model": "asi1",
"messages": [],
"reasoning_effort": "high",
"use_reasoning_history": true,
"stream": true
}

asi1 is the right tier for routine-to-moderate coding; move to asi1-ultra for audits and cross-file refactors. Deep reasoning earns its cost on code. use_reasoning_history carries the agent’s own conclusions across its tool-call turns, which also keeps the prefix stable enough to cache. Sampling stays at the defaults.

High-volume extraction or classification

{
"model": "asi1-mini",
"messages": [],
"enable_thinking": false,
"temperature": 0.2
}

asi1-mini is the cost and latency tier. Thinking is off by default here and that is the right setting for short structured output; a low temperature is comfortable on this model.

Next steps

  1. Reasoning - The depth controls, per endpoint
  2. Prompt caching - How to keep a stable prefix and read the cache counters
  3. ASI:One Models - Which model to reach for, and the sampling defaults
  4. Errors and Rate Limits - What each status code means and what to retry