> ## Documentation Index
> Fetch the complete documentation index at: https://docs.macroscope.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Models & Token Pricing

> The models Macroscope runs on, how reasoning and effort behave, and the raw per-token rate for each one

This page is the single source of truth for the models Macroscope runs on and what their tokens cost. The rates below are raw model cost, so a model's price is the same wherever you use it.

Each product bills these rates under its own pricing — see [Model Selection & Costs](/check-run-agents/costs#how-runs-are-billed) for Check Run Agents.

## Available Models

| Model                 | Reasoning                                                | Effort                            | Notes                                                                                                                                |
| --------------------- | -------------------------------------------------------- | --------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| **Anthropic**         |                                                          |                                   |                                                                                                                                      |
| `claude-opus-4-5`     | `low` (default), `medium`, `high`                        | `low` (default), `medium`, `high` | Earlier Opus generation. Smaller context window — best for short, targeted checks.                                                   |
| `claude-opus-4-6`     | auto                                                     | `low` (default), `medium`, `high` | Strongest general-purpose option. Large context window.                                                                              |
| `claude-opus-4-7`     | auto                                                     | `low` (default), `medium`, `high` | Large context window.                                                                                                                |
| `claude-opus-4-8`     | auto                                                     | `low` (default), `medium`, `high` | Large context window.                                                                                                                |
| `claude-opus-5`       | auto                                                     | `low` (default), `medium`, `high` | Large context window.                                                                                                                |
| `claude-opus-5-5`     | auto                                                     | `low` (default), `medium`, `high` | Latest Opus. Large context window.                                                                                                   |
| `claude-sonnet-4-5`   | `low` (default), `medium`, `high`                        | `low` (default), `medium`, `high` | Faster than Opus. Smaller context window.                                                                                            |
| `claude-sonnet-4-6`   | auto                                                     | `low` (default), `medium`, `high` | Faster than Opus, with a large context window.                                                                                       |
| `claude-sonnet-5`     | auto                                                     | `low` (default), `medium`, `high` | Latest Sonnet. Faster than Opus, with a large context window.                                                                        |
| `claude-fable-5`      | auto                                                     | `low` (default), `medium`, `high` | Highest intelligence. Best for hard tasks. **<u>Note: Fable 5 does not support Zero Data Retention.</u>**                            |
| `claude-fable-5-1`    | auto                                                     | `low` (default), `medium`, `high` | Latest Fable. Highest intelligence. Best for hard tasks. **<u>Note: Fable 5.1 does not support Zero Data Retention.</u>**            |
| **OpenAI**            |                                                          |                                   |                                                                                                                                      |
| `gpt-5-2`             | `low` (default), `medium`, `high`                        | —                                 | Earlier GPT generation.                                                                                                              |
| `gpt-5-4`             | `low` (default), `medium`, `high`                        | —                                 | More capable than 5.2.                                                                                                               |
| `gpt-5-5`             | `low` (default), `medium`, `high`                        | —                                 | Previous flagship GPT.                                                                                                               |
| `gpt-5-6-sol`         | `low` (default), `medium`, `high`, `xhigh`, `max`        | —                                 | Latest GPT. Frontier tier — highest quality.                                                                                         |
| `gpt-5-6-terra`       | `low` (default), `medium`, `high`, `xhigh`, `max`        | —                                 | Balanced price and quality.                                                                                                          |
| `gpt-5-6-luna`        | `low` (default), `medium`, `high`, `xhigh`, `max`        | —                                 | Cost-optimized.                                                                                                                      |
| `gpt-6-astra`         | `low` (default), `medium`, `high`, `xhigh`, `max`        | —                                 | Latest flagship. Large context window. Always on; `off` maps to low.                                                                 |
| `gpt-6-sol`           | `off`, `low` (default), `medium`, `high`, `xhigh`, `max` | —                                 | Balances intelligence and cost for coding and agentic work. Large context window. Every thinking level is distinct, including `off`. |
| `gpt-6-luna`          | `off`, `low` (default), `medium`, `high`, `xhigh`, `max` | —                                 | High-volume, cost-optimized tier. Large context window. Every thinking level is distinct, including `off`.                           |
| **xAI**               |                                                          |                                   |                                                                                                                                      |
| `grok-4-6`            | `low` (default), `medium`, `high`, `xhigh`               | —                                 | Large context window. Always on. Each level is distinct; `off` maps to low.                                                          |
| **Open source**       |                                                          |                                   |                                                                                                                                      |
| `deepseek-v4-1-flash` | `low` (default), `high`, `xhigh`                         | —                                 | Large context window. Always on. `off`/`low` map to low effort, `medium`/`high` to high, `xhigh` to max.                             |
| `deepseek-v4-flash`   | `low` (default), `high`, `xhigh`                         | —                                 | Large context window. Always on. `off`/`low` map to low effort, `medium`/`high` to high, `xhigh` to max.                             |
| `deepseek-v4-pro`     | `off` or on (default)                                    | —                                 | Large context window. Any thinking level maps to on.                                                                                 |
| `kimi-k3`             | `low` (default), `high`, `xhigh`                         | —                                 | Large context window. Always on. `off`/`low` map to low effort, `medium`/`high` to high, `xhigh` to max.                             |
| `kimi-k2-7-code`      | always on                                                | —                                 | Reasoning cannot be turned off.                                                                                                      |
| `glm-5-3`             | `low` (default), `high`, `xhigh`                         | —                                 | Large context window. Always on. `off`/`low` map to low effort, `medium`/`high` to high, `xhigh` to max.                             |
| `glm-5-3-flash`       | `low` (default), `high`, `xhigh`                         | —                                 | Large context window. Always on. `off`/`low` map to low effort, `medium`/`high` to high, `xhigh` to max.                             |
| `glm-5-2`             | `off`, `high` (default), `xhigh`                         | —                                 | Large context window. `low`/`medium` map to `high`.                                                                                  |
| `minimax-m3`          | `off` or on (default)                                    | —                                 | Large context window. Any thinking level maps to on.                                                                                 |
| `qwen-3-8-max`        | `off`, `low` (default), `medium`, `high`, `xhigh`        | —                                 | Large context window. Every thinking level is distinct, including `off`.                                                             |

* **Default** = the value used when the field is omitted.
* **Effort** applies to Anthropic models only — GPT, Grok, and open-source models ignore it (shown as `—`). On those models every `effort` level collapses to the same variant, so it has no effect.
* Newer Anthropic models (Opus 4.6 / 4.7 / 4.8 / 5 / 5.5, Sonnet 4.6 / 5, Fable 5 / 5.1) set thinking automatically, so their `reasoning` field is ignored (shown as auto).
* Writing a level a model does not support, such as `xhigh` on `gpt-5-2` or `max` on `kimi-k3`, falls back to `low` and surfaces a warning.

<Note>
  Open-source models are opt-in and typically cost a fraction of the frontier options — a good fit for specialized, high-volume, or narrowly-scoped work, while Opus remains the default where the strongest model matters.
</Note>

<Note>
  Open-source models are served through Fireworks AI's hosted inference — Macroscope does not run local or self-hosted models, and you cannot point at your own endpoints. As we expand the set of inference providers we support, capacity fluctuations from these providers (during heavy load) mean latency may occasionally be higher due to request queuing.
</Note>

## Reasoning and Effort

* `reasoning` accepts `off | low | medium | high | xhigh | max`
* `effort` accepts `low | medium | high`

Which field applies depends on the model:

* **Anthropic** — `effort` always applies. `claude-opus-4-5` and `claude-sonnet-4-5` additionally honor `reasoning`, which maps to [extended thinking](https://platform.claude.com/docs/en/build-with-claude/extended-thinking#supported-models); newer Anthropic models set thinking automatically and ignore `reasoning`.
* **OpenAI, xAI, and open source** — `reasoning` selects the thinking level; `effort` is ignored.

Writing a `reasoning` level a model does not support falls back to `low` and surfaces a warning.

For how to set these fields on a Check Run Agent, see [Model Selection & Costs](/check-run-agents/costs#reasoning-and-effort).

## Pricing

Raw model cost, USD per 1M tokens. The four columns are the four token counts a request is metered on, and they do not overlap — each token is counted once:

* **Input** — prompt tokens read fresh, with no cache involved.
* **Output** — everything the model generated, reasoning tokens included.
* **Cache write** — prompt tokens written into the provider's prompt cache. **—** means the provider charges nothing to create a cache entry.
* **Cached input** — prompt tokens read back out of the cache instead of re-read.

Model ids are written here the way you set them in configuration. A provider's own usage view may spell the same model differently — `gpt-5-6-sol` is `gpt-5.6-sol` there, `grok-4-6` is `grok-4.6`, and the open-source models appear under vendor paths such as `moonshotai/Kimi-K3`.

| Model                                                                                                                                                                      | Input   | Output  | Cache write | Cached input |
| -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------- | ------- | ----------- | ------------ |
| **Anthropic**                                                                                                                                                              |         |         |             |              |
| `claude-opus-4-5`                                                                                                                                                          | \$5.00  | \$25.00 | \$6.25      | \$0.50       |
| `claude-opus-4-6`                                                                                                                                                          | \$5.00  | \$25.00 | \$6.25      | \$0.50       |
| `claude-opus-4-7`                                                                                                                                                          | \$5.00  | \$25.00 | \$6.25      | \$0.50       |
| `claude-opus-4-8`                                                                                                                                                          | \$5.00  | \$25.00 | \$6.25      | \$0.50       |
| `claude-opus-5`                                                                                                                                                            | \$5.00  | \$25.00 | \$6.25      | \$0.50       |
| `claude-opus-5-5`                                                                                                                                                          | \$4.00  | \$20.00 | \$5.00      | \$0.20       |
| `claude-sonnet-4-5`                                                                                                                                                        | \$3.00  | \$15.00 | \$3.75      | \$0.30       |
| `claude-sonnet-4-6`                                                                                                                                                        | \$3.00  | \$15.00 | \$3.75      | \$0.30       |
| `claude-sonnet-5`                                                                                                                                                          | \$2.00  | \$10.00 | \$2.50      | \$0.20       |
| `claude-fable-5`                                                                                                                                                           | \$10.00 | \$50.00 | \$12.50     | \$1.00       |
| `claude-fable-5-1`                                                                                                                                                         | \$10.00 | \$50.00 | \$12.50     | \$0.25       |
| **OpenAI**                                                                                                                                                                 |         |         |             |              |
| `gpt-5-2`                                                                                                                                                                  | \$1.75  | \$14.00 | —           | \$0.18       |
| `gpt-5-4`                                                                                                                                                                  | \$2.50  | \$15.00 | —           | \$0.25       |
| `gpt-5-5`                                                                                                                                                                  | \$5.00  | \$30.00 | —           | \$0.50       |
| `gpt-5-6-sol`                                                                                                                                                              | \$4.00  | \$20.00 | \$5.00      | \$0.40       |
| `gpt-5-6-terra`                                                                                                                                                            | \$2.00  | \$12.00 | \$2.50      | \$0.20       |
| `gpt-5-6-luna`                                                                                                                                                             | \$0.20  | \$1.20  | \$0.25      | \$0.02       |
| `gpt-6-astra`                                                                                                                                                              | \$10.00 | \$50.00 | \$12.50     | \$1.00       |
| `gpt-6-sol`                                                                                                                                                                | \$2.00  | \$10.00 | \$2.50      | \$0.20       |
| `gpt-6-luna`                                                                                                                                                               | \$0.10  | \$0.50  | \$0.13      | \$0.01       |
| **xAI**                                                                                                                                                                    |         |         |             |              |
| `grok-4-6`                                                                                                                                                                 | \$2.00  | \$6.00  | —           | \$0.50       |
| **Open source**                                                                                                                                                            |         |         |             |              |
| <Tooltip headline="New rates from October 1, 2026" tip="Input (uncached) $0.30, output $1.20, cached input $0.01 — all USD per 1M tokens.">`deepseek-v4-1-flash`</Tooltip> | \$0.22  | \$0.66  | —           | \$0.01       |
| `deepseek-v4-flash`                                                                                                                                                        | \$0.44  | \$1.32  | —           | \$0.02       |
| `deepseek-v4-pro`                                                                                                                                                          | \$1.32  | \$3.96  | —           | \$0.05       |
| `kimi-k3`                                                                                                                                                                  | \$3.00  | \$15.00 | —           | \$0.30       |
| `kimi-k2-7-code`                                                                                                                                                           | \$0.95  | \$4.00  | —           | \$0.19       |
| `glm-5-3`                                                                                                                                                                  | \$1.40  | \$4.40  | —           | \$0.26       |
| `glm-5-3-flash`                                                                                                                                                            | \$0.15  | \$0.50  | —           | \$0.03       |
| `glm-5-2`                                                                                                                                                                  | \$1.40  | \$4.40  | —           | \$0.14       |
| `minimax-m3`                                                                                                                                                               | \$0.30  | \$1.20  | —           | \$0.06       |
| `qwen-3-8-max`                                                                                                                                                             | \$2.00  | \$6.00  | —           | \$0.25       |

<Note>
  **Large prompts bill at a higher rate on some models.** Crossing the threshold reprices the *whole* request, not just the tokens past it. "Prompt tokens" means uncached input plus cached input plus cache writes.

  * `gpt-5-4`, `gpt-5-5`, `gpt-5-6-sol`, `gpt-5-6-terra`, `gpt-5-6-luna`, `gpt-6-astra`, `gpt-6-sol`, `gpt-6-luna` — above 272,000 prompt tokens: 2x input, 1.5x output, and 2x on both cache columns where they are charged.
  * `claude-sonnet-4-5` — above 200,000 prompt tokens: 2x input, 1.5x output, 2x cache write, 2x cached input.
  * `grok-4-6` — at 200,000 prompt tokens or more: 2x on every token, output included.

  Every other model above bills the same rate at any prompt size.
</Note>

<Warning>
  `deepseek-v4-1-flash` pricing changes on **October 1, 2026**. New rates, USD per 1M tokens:

  **Input:** \$0.30, **Output:** \$1.20, **Cached input:** \$0.01

  The table above shows the rates in effect until then.
</Warning>
