> For clean Markdown of any page, append .md to the page URL. > For a complete documentation index, see https://docs.asi1.ai/documentation/models/llms.txt. > For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.asi1.ai/_mcp/server. # ASI:One Models > Three models that share one API - asi1 is the general default, asi1-ultra goes deeper, asi1-mini goes faster. ASI:One offers three models that share the same API surface, the same API key, and the same capabilities. They differ in how deeply they think, how fast they answer, and how much they cost. Switching between them is a one-line change. The `model` field is the only thing that changes between them: #### asi1 ```bash curl -X POST https://api.asi1.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $ASI_ONE_API_KEY" \ -d '{ "model": "asi1", "messages": [ {"role": "user", "content": "Help me plan a trip to Tokyo"} ] }' ``` #### asi1-ultra ```bash curl -X POST https://api.asi1.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $ASI_ONE_API_KEY" \ -d '{ "model": "asi1-ultra", "messages": [ {"role": "user", "content": "Audit this 800-line contract and flag risks"} ] }' ``` #### asi1-mini ```bash curl -X POST https://api.asi1.ai/v1/chat/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $ASI_ONE_API_KEY" \ -d '{ "model": "asi1-mini", "messages": [ {"role": "user", "content": "Classify this support ticket as billing, bug, or feature request"} ] }' ``` --- ## Choosing a model Start with `asi1`. Move off it only when you have a reason: 1. **Is latency or cost your binding constraint?** Use `asi1-mini`. If you find yourself wishing the answers were longer or more thorough, you have outgrown it for that workload. 2. **Is the task hard enough that quality beats speed?** Use `asi1-ultra`, and expect noticeably higher latency. 3. **Otherwise** use `asi1`. Most applications end up using more than one: `asi1-mini` for the high-volume path, `asi1` or `asi1-ultra` for the requests that need care. --- ## Comparison | | `asi1` | `asi1-ultra` | `asi1-mini` | | --------------------------------------------------------------------------------- | ----------------------- | ---------------- | ----------------- | | **Position** | General-purpose default | Most capable | Fastest, lightest | | **[Reasoning budget](/documentation/build-with-asi-one/reasoning)** | Configurable | Not configurable | Not configurable | | **[Image input](/documentation/build-with-asi-one/chat-completions#image-input)** | Yes | No | No | | **Context window** | 196,608 tokens | 1,000,000 tokens | 262,144 tokens | | **Latency** | Moderate | Highest | Lowest | | **Cost** | Moderate | Highest | Lowest | Everything not in that table is the same across all three: [streaming](/documentation/build-with-asi-one/chat-completions#streaming), [tool calling](/documentation/build-with-asi-one/tool-calling), [reasoning](/documentation/build-with-asi-one/reasoning) (the budget differs, support does not), [structured data](/documentation/build-with-asi-one/structured-data), and [OpenAI compatibility](/documentation/build-with-asi-one/openai-compatibility) on both [`/v1/chat/completions`](/documentation/build-with-asi-one/chat-completions) and [`/v1/responses`](/documentation/build-with-asi-one/responses). [Planner mode](/documentation/build-with-asi-one/planner) is the exception: it runs on `/v1/chat/completions`, on all three models. Latency and cost above are relative to each other rather than absolute. Your account's usage and balance are in your [ASI:One account](https://asi1.ai/account-and-settings/wallet). ### Pricing Prices are per million tokens, in US dollars. Every model bills input tokens, output tokens, and reasoning tokens separately; see [Reasoning](/documentation/build-with-asi-one/reasoning) for how reasoning billing works. | Model | Input | Cached input | Output | | ------------ | ------ | ------------ | ------ | | `asi1-mini` | \$0.18 | \$0.018 | \$0.45 | | `asi1` | \$0.50 | \$0.05 | \$1.90 | | `asi1-ultra` | \$1.26 | \$0.126 | \$3.96 | Cached input tokens are billed at 10% of the input rate, and reasoning tokens are billed at the output rate for the model. ### Context windows The context window is the **combined** budget for your input and the model's output. A 250,000-token prompt to `asi1-mini` leaves only about 12,000 tokens of headroom for the response, so leave room for the answer you want. > **Warning** > > Over-length requests are rejected rather than truncated. ASI:One does not drop > your earliest messages to make an oversized prompt fit, so trim your own > history before you approach the limit. There is no separate cap on output length. `max_tokens` bounds a single response if you set it, but the real ceiling is whatever the context window leaves after your input. --- ## What each model is for ### asi1 General chat and assistants, tool calling, and mixed workloads where you cannot tell in advance how hard any individual request will be. ### asi1-ultra Work where quality matters more than speed: deep research and synthesis, code and contract review, long multi-step runs, and strategic planning. ### asi1-mini Real-time chat and voice, classification and routing, autocomplete, and high-volume paths where each request is straightforward on its own. What you trade for the speed is depth and response length. --- ## Switching models Change one field. Everything else about the request stays the same. ```diff { - "model": "asi1", + "model": "asi1-mini", "messages": [...] } ``` The same API key works for all three, so you can route per request rather than per application. Two patterns worth knowing: * **Cascade.** Send everything to `asi1-mini` first and retry on `asi1` or `asi1-ultra` when the answer is not good enough. * **Route by request.** Classify the incoming request and pick the model up front. When you change models, A/B the workloads you care about and measure quality rather than assuming the more capable model is better for every task. On simple requests it is mostly slower. --- ## Next steps 1. **[Chat Completions API](/documentation/build-with-asi-one/chat-completions)** - The default endpoint in full, including streaming 2. **[Reasoning](/documentation/build-with-asi-one/reasoning)** - Turn reasoning on and control how much of it you pay for 3. **[Tool Calling](/documentation/build-with-asi-one/tool-calling)** - Let the model call your own functions 4. **[Planner Mode](/documentation/build-with-asi-one/planner)** - Multi-step work against Agentverse agents > Three models that share one API - asi1 is the general default, asi1-ultra goes deeper, asi1-mini goes faster.