> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.asi1.ai/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.asi1.ai/_mcp/server.

# ASI:One Models

> Three models that share one API - asi1 is the general default, asi1-ultra goes deeper, asi1-mini goes faster.

ASI:One offers three models that share the same API surface, the same API key, and
the same capabilities. They differ in how deeply they think, how fast they answer,
and how much they cost. Switching between them is a one-line change.

The `model` field is the only thing that changes between them:

#### asi1

```bash
curl -X POST https://api.asi1.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ASI_ONE_API_KEY" \
  -d '{
    "model": "asi1",
    "messages": [
      {"role": "user", "content": "Help me plan a trip to Tokyo"}
    ]
  }'
```

#### asi1-ultra

```bash
curl -X POST https://api.asi1.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ASI_ONE_API_KEY" \
  -d '{
    "model": "asi1-ultra",
    "messages": [
      {"role": "user", "content": "Audit this 800-line contract and flag risks"}
    ]
  }'
```

#### asi1-mini

```bash
curl -X POST https://api.asi1.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $ASI_ONE_API_KEY" \
  -d '{
    "model": "asi1-mini",
    "messages": [
      {"role": "user", "content": "Classify this support ticket as billing, bug, or feature request"}
    ]
  }'
```

---

## Choosing a model

Start with `asi1`. Move off it only when you have a reason:

1. **Is latency or cost your binding constraint?** Use `asi1-mini`. If you find
   yourself wishing the answers were longer or more thorough, you have outgrown it
   for that workload.
2. **Is the task hard enough that quality beats speed?** Use `asi1-ultra`, and
   expect noticeably higher latency.
3. **Otherwise** use `asi1`.

Most applications end up using more than one: `asi1-mini` for the high-volume path,
`asi1` or `asi1-ultra` for the requests that need care.

---

## Comparison

|                                                                                   | `asi1`                  | `asi1-ultra`     | `asi1-mini`       |
| --------------------------------------------------------------------------------- | ----------------------- | ---------------- | ----------------- |
| **Position**                                                                      | General-purpose default | Most capable     | Fastest, lightest |
| **[Reasoning budget](/documentation/build-with-asi-one/reasoning)**               | Configurable            | Not configurable | Not configurable  |
| **[Image input](/documentation/build-with-asi-one/chat-completions#image-input)** | Yes                     | No               | No                |
| **Context window**                                                                | 196,608 tokens          | 1,000,000 tokens | 262,144 tokens    |
| **Latency**                                                                       | Moderate                | Highest          | Lowest            |
| **Cost**                                                                          | Moderate                | Highest          | Lowest            |

Everything not in that table is the same across all three:
[streaming](/documentation/build-with-asi-one/chat-completions#streaming), [tool calling](/documentation/build-with-asi-one/tool-calling), [reasoning](/documentation/build-with-asi-one/reasoning)
(the budget differs, support does not), [structured data](/documentation/build-with-asi-one/structured-data), and [OpenAI compatibility](/documentation/build-with-asi-one/openai-compatibility) on both
[`/v1/chat/completions`](/documentation/build-with-asi-one/chat-completions) and
[`/v1/responses`](/documentation/build-with-asi-one/responses).

[Planner mode](/documentation/build-with-asi-one/planner) is the exception: it
runs on `/v1/chat/completions`, on all three models.

Latency and cost above are relative to each other rather than absolute. Your
account's usage and balance are in your [ASI:One account](https://asi1.ai/account-and-settings/wallet).

### Pricing

Prices are per million tokens, in US dollars. Every model bills input tokens,
output tokens, and reasoning tokens separately; see [Reasoning](/documentation/build-with-asi-one/reasoning)
for how reasoning billing works.

| Model        | Input  | Cached input | Output |
| ------------ | ------ | ------------ | ------ |
| `asi1-mini`  | \$0.18 | \$0.018      | \$0.45 |
| `asi1`       | \$0.50 | \$0.05       | \$1.90 |
| `asi1-ultra` | \$1.26 | \$0.126      | \$3.96 |

Cached input tokens are billed at 10% of the input rate, and reasoning tokens are
billed at the output rate for the model.

### Context windows

The context window is the **combined** budget for your input and the model's
output. A 250,000-token prompt to `asi1-mini` leaves only about 12,000 tokens of
headroom for the response, so leave room for the answer you want.

> **Warning**
>
> Over-length requests are rejected rather than truncated. ASI:One does not drop
> your earliest messages to make an oversized prompt fit, so trim your own
> history before you approach the limit.

There is no separate cap on output length. `max_tokens` bounds a single response
if you set it, but the real ceiling is whatever the context window leaves after
your input.

---

## What each model is for

### asi1

General chat and assistants, tool calling, and mixed workloads where you cannot
tell in advance how hard any individual request will be.

### asi1-ultra

Work where quality matters more than speed: deep research and synthesis, code and
contract review, long multi-step runs, and strategic planning.

### asi1-mini

Real-time chat and voice, classification and routing, autocomplete, and
high-volume paths where each request is straightforward on its own. What you
trade for the speed is depth and response length.

---

## Switching models

Change one field. Everything else about the request stays the same.

```diff
  {
-   "model": "asi1",
+   "model": "asi1-mini",
    "messages": [...]
  }
```

The same API key works for all three, so you can route per request rather than per
application. Two patterns worth knowing:

* **Cascade.** Send everything to `asi1-mini` first and retry on `asi1` or
  `asi1-ultra` when the answer is not good enough.
* **Route by request.** Classify the incoming request and pick the model up front.

When you change models, A/B the workloads you care about and measure quality rather
than assuming the more capable model is better for every task. On simple requests
it is mostly slower.

---

## Next steps

1. **[Chat Completions API](/documentation/build-with-asi-one/chat-completions)** - The default endpoint in full, including streaming
2. **[Reasoning](/documentation/build-with-asi-one/reasoning)** - Turn reasoning on and control how much of it you pay for
3. **[Tool Calling](/documentation/build-with-asi-one/tool-calling)** - Let the model call your own functions
4. **[Planner Mode](/documentation/build-with-asi-one/planner)** - Multi-step work against Agentverse agents