Tuning
Most quality problems are solved by reasoning depth or by picking a different model, not by sampling. This page is the order to work in, plus two request shapes that show it applied.
If you are looking for the values themselves, they live with the thing they describe: context windows and sampling defaults on the models page, reasoning depth on the reasoning page, and rate limits with the errors they produce.
Where to start
- Match the model to the workload.
asi1-minifor high-volume chat, classification and autocomplete;asi1for general chat, tool calling and routine coding;asi1-ultrafor deep research, code audits and long agentic runs. A workload on the wrong model cannot be fixed with sampling. - Then set reasoning depth. Depth moves quality by points where sampling
moves it by impressions, so reach for
reasoning_effortfirst. - Change sampling last, and one parameter at a time. On
asi1andasi1-ultrathe defaults are the values the models were trained and evaluated at, so prefer steering through the prompt.asi1-minitunes comfortably across a wider range. - Sampling never costs you a cache hit.
temperature,top_pandmax_tokensact on the output distribution, not the prompt prefix, so vary them freely. See Prompt caching.
Symptom to fix
Worked examples
A coding assistant with tools
asi1 is the right tier for routine-to-moderate coding; move to asi1-ultra
for audits and cross-file refactors. Deep reasoning earns its cost on code.
use_reasoning_history carries the agent’s own conclusions across its
tool-call turns, which also keeps the prefix stable enough to cache. Sampling
stays at the defaults.
High-volume extraction or classification
asi1-mini is the cost and latency tier. Thinking is off by default here and
that is the right setting for short structured output; a low temperature is
comfortable on this model.
Next steps
- Reasoning - The depth controls, per endpoint
- Prompt caching - How to keep a stable prefix and read the cache counters
- ASI:One Models - Which model to reach for, and the sampling defaults
- Errors and Rate Limits - What each status code means and what to retry