> ## Documentation Index
> Fetch the complete documentation index at: https://docs.macroscope.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Benchmarking API

> Run Macroscope's standardized code-review benchmark against your own model endpoint (bring-your-own-key) and get back recall, precision, F1, and cost metrics.

Run Macroscope's standardized code-review benchmark against **your own model
endpoint** (bring-your-own-key) and get back recall / precision / F1 and related
metrics. The dataset, review harness, prompts, and scoring are fixed and identical
for every customer, so results are directly comparable across models — you supply
only the model endpoint.

<Note>
  A run is asynchronous: you `POST` a run, receive a `benchmarkId`, then poll a
  status URL until it completes and returns the metrics.
</Note>

## Authentication

Every request requires a bearer token issued to you by Macroscope:

```
Authorization: Bearer mapi_{key_id}_{secret}
```

Keys are provisioned by Macroscope (they are not self-service). Treat the token as
a secret; it identifies and rate-limits your account.

The **base URL** is provided with your key (per environment). All paths below are
relative to it.

## Start a run

```
POST /api/v1/benchmark-runs
Content-Type: application/json
```

### Request body

```jsonc theme={null}
{
  "modelBackend": {
    "provider":  "openai",                 // required — WIRE PROTOCOL, see below
    "baseUrl":   "https://api.your-host/v1",// required — HTTPS endpoint base URL
    "modelName": "your-model-id",           // required
    "apiKey":    "sk-...",                  // required — key for YOUR endpoint
    "effort":    "high",                    // required — reasoning effort ("off" if none)
    "contextWindow":   200000,              // required — your model's context window
    "maxOutputTokens": 64000,               // required — your model's max output tokens
    "costRate": {                           // optional — price your model (see below)
      "inputCost": 300,  "outputCost": 1500,
      "cacheReadCost": 30, "cacheWriteCost": 375
    }
  },
  "k": 1,                                    // optional — draws per bug (1–5, default 3)
  "systemPrompt": "You are a reviewer..."    // optional — custom detection system prompt
}                                            //            (makes the run non-comparable)
```

Inside `modelBackend`, every field is required except `costRate`. The top-level `k` and
`systemPrompt` are optional. There are otherwise no defaults: you declare effort and
capacities explicitly, so a run can't be silently made non-comparable by a value we guessed
on your behalf.

| Field | Required | Notes |
| - | - | - |
| `provider` | yes | The request/response **wire protocol**, not a vendor. `openai` = OpenAI-compatible Chat Completions — which covers **most open-source / self-hosted models** (e.g. Kimi, Llama, DeepSeek) served via Fireworks, vLLM, or your own endpoint; `anthropic` = Anthropic Messages. What matters is the wire format your endpoint speaks, not who made the model. |
| `baseUrl` | yes | HTTPS only. No user-info, query string, or fragment (put the key in `apiKey`, not the URL). |
| `modelName` | yes | The model identifier your endpoint expects. It is used **only** to call your endpoint — it is never echoed back in the status response, and the run is reported under its opaque `benchmarkId`. |
| `apiKey` | yes | Bearer key for **your** endpoint. It is KMS-encrypted at rest, never logged, and never written to workflow history; it is used only to call your endpoint and blanked when the run finishes. |
| `effort` | yes | One of `off`, `low`, `medium`, `high`, `xhigh` (alias `extrahigh`), `max`. Sets the reasoning effort on your model's calls. **Use `off` for a model with no reasoning/effort setting** — it sends no effort parameter. |
| `contextWindow` | yes | Your model's context window in tokens. The harness fits the packed input to it (`input = context_window − output`), so a smaller window reviews a proportionally smaller pack. **Minimum eligible context is 128,000.** A smaller window packs less context than a large one, so scores across very different context tiers aren't directly comparable. |
| `maxOutputTokens` | yes | Your model's max output tokens. |
| `costRate` | no | Per-token price so cost metrics are populated for a model we don't price by default. See [Pricing your model](#pricing-your-model). Omit it to leave cost unpriced. |
| `k` | no | Number of independent review draws per bug (1–5). Default **3**. |
| `systemPrompt` | no | A custom detection **system** prompt (bring-your-own-prompt). When set it replaces **only** the issue-detection system text — the harness still injects the PR code, provides the `report_issue`/`review_complete` tools, enforces the output schema, and scores with the same frozen validation judge. When present it must be non-empty and ≤ 50,000 bytes. **Supplying it makes the run non-comparable** to the public board — see [Custom detection prompt](#custom-detection-prompt-non-comparable). Omit it to run the fixed control prompt. |

Example request:

```bash theme={null}
curl -s -X POST "$BASE_URL/api/v1/benchmark-runs" \
  -H "Authorization: Bearer $MACROSCOPE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "modelBackend": {
      "provider": "openai",
      "baseUrl": "https://api.your-host/v1",
      "modelName": "your-model-id",
      "apiKey": "sk-...",
      "effort": "high",
      "contextWindow": 200000,
      "maxOutputTokens": 64000
    },
    "k": 3
  }'
```

### How the model name is handled

You send the real `modelName` (and `baseUrl`) because we need them to call your endpoint.
They are handled like your `apiKey`: used only for the inference call, encrypted at rest, and
**never written to the results, exports, logs, or API responses**. The run is reported under
its opaque `benchmarkId` — the status response does **not** echo the model name back (you poll
by `benchmarkId`, so you already know which model it is). Because cost is attributed under that
id rather than a known model name, supply `costRate` to price a model we don't price by default
(see below); otherwise the run is unpriced.

### Pricing your model

If your model isn't one we price by default, supply `costRate` so `totalCost` /
`costPerReview` are populated instead of left unpriced. Each field is the price in
**cents per million tokens** — so a model that costs \$3.00 per million input tokens has
`inputCost: 300`:

| Field | Meaning | Example: \$3.00 / 1M input tokens |
| - | - | - |
| `inputCost` | prompt tokens | `300` |
| `outputCost` | completion tokens | — |
| `cacheReadCost` | cache-read tokens | — |
| `cacheWriteCost` | cache-write tokens | — |

Each value must be in `[0, 10000000]`; a bucket set to `0` is priced free. Omit the
whole `costRate` object to run **unpriced** (cost reported as 0).

### Custom detection prompt (non-comparable)

By default every run uses the **fixed control detection prompt**, which is what makes
scores comparable across models and customers. If you supply `systemPrompt`, you replace
**only** the detection system prompt — your instructions for *what to look for and how to
think*. Everything else stays fixed:

* the harness still injects the PR code and the reference-graph context,
* it still provides the `report_issue` / `review_complete` tools and enforces the output
  schema (your `systemPrompt` cannot change these),
* the instruction prompt and the frozen validation **judge** that scores findings are
  unchanged, so a finding is scored exactly as it would be on a control run.

**Keep it reasonably short.** Your `systemPrompt` shares the model's context window with the
PR code and the reference-graph context the harness injects. When your prompt plus the PR
code does not fit the declared `contextWindow`, that review **errors** — its result is not
scored, rather than being silently reviewed on a truncated pack. A review that errors counts
toward the run's error rate like any other errored review; if enough reviews error, the whole
run is **invalidated with a context-window reason** (status `errored`, message naming the
context window — see the error table below). The 50,000-byte ceiling is a hard limit, not a
target: a focused prompt leaves room for the code under review, so more reviews fit the
window and are scored.

**A run that supplies `systemPrompt` is not comparable to the public benchmark board.**
Its score reflects your own detection instructions, not the control harness, so it is
segregated automatically:

* it runs under a **distinct `benchmarkHarnessVersion`** (the custom prompt is part of the
  hashed harness surface), so it never shares a bucket with control runs in any export or
  dashboard; and
* the status response marks it explicitly with `comparable: false` / `customPrompt: true`
  and a `comparabilityNote` (see below).

Two runs that supply the **same** `systemPrompt` share a `benchmarkHarnessVersion`, so they
remain comparable **to each other** — just not to the board. Omit `systemPrompt` to run the
comparable control prompt.

### Response — `202 Accepted`

```json theme={null}
{
  "benchmarkId": "6dbdd89d-171f-481c-96fc-d1fdfbda5dec",
  "statusUrl": "https://<base-url>/api/v1/benchmark-runs/6dbdd89d-171f-481c-96fc-d1fdfbda5dec",
  "k": 1,
  "benchmarkApiVersion": "2026-09-11",
  "benchmarkHarnessVersion": "4d954bee2fb72fe444118c86d5add75af41936170ba8db1ce797b2981ca161d4",
  "benchmarkDatasetId": "code_review_dataset.json"
}
```

* `benchmarkId` — opaque run id; use it to poll status.
* `benchmarkHarnessVersion` — content hash of the fixed harness (prompts, tools,
  scoring, standard layer). Two runs with the same value were scored identically;
  it changes only when the harness changes.
* `benchmarkDatasetId` — at submit time this is the dataset name; the **status**
  response reports the dataset content SHA once the run completes.

## Get run status / results

```
GET /api/v1/benchmark-runs/{benchmarkId}
```

Poll until `status` is `completed` (or `errored`). While running, `results` is
`null` and `percentComplete` reports progress.

```bash theme={null}
curl -s "$BASE_URL/api/v1/benchmark-runs/6dbdd89d-171f-481c-96fc-d1fdfbda5dec" \
  -H "Authorization: Bearer $MACROSCOPE_API_KEY"
```

### Response

```jsonc theme={null}
{
  "benchmarkId": "6dbdd89d-171f-481c-96fc-d1fdfbda5dec",
  "status": "completed",              // running | completed | errored
  // "percentComplete": 60,          // present ONLY while status is "running"
  "message": "",                     // errored runs carry a specific reason (see "When a run errors")
  "benchmarkApiVersion": "2026-09-11",
  "benchmarkHarnessVersion": "4d954bee2f...",
  "benchmarkDatasetId": "cd4560f98b...", // dataset content SHA on completion
  // comparable / customPrompt / comparabilityNote appear once there is a scored
  // result (omitted while running). A control run is comparable:true; a run that
  // supplied a custom systemPrompt is comparable:false / customPrompt:true with a note.
  "comparable": true,
  "customPrompt": false,
  "results": {
    "k": 1,
    "f1Score": 0.646,
    "recallAvgK": 0.6,
    "recallMaxK": 0.6,
    "precisionAvgK": 0.7,
    "signalToNoiseRatio": 1.222,
    "bySeverity": {
      "critical": { "recall": 0.38, "precision": 0.97, "f1Score": 0.55, "precisionScored": 40 },
      "high":     { "recall": 0.56, "precision": 0.87, "f1Score": 0.68, "precisionScored": 907 },
      "medium":   { "recall": 0.65, "precision": 0.88, "f1Score": 0.75, "precisionScored": 660 },
      "low":      { "recall": 0.48, "precision": 0.82, "f1Score": 0.61, "precisionScored": 188 },
      "none":     { "recall": null, "precision": 0.63, "f1Score": null, "precisionScored": 38 }
    },
    "totalCost": 17.91,
    "costPerReview": 1.79,
    "latencyMsP50": 394832,
    "latencyMsP90": 632559,
    "latencyMsMax": 649254,
    "totalDraws": 100,                  // review attempts (records x k)
    "okDraws": 98,
    "erroredDraws": 2,
    "errorRate": 0.02,                  // erroredDraws / totalDraws
    "errors": {                         // present only when erroredDraws > 0
      "erroredDraws": 2,
      "contextWindowExceededDraws": 0,
      "byReason": [
        { "reason": "endpoint_unreachable", "count": 2 }
      ],
      "byLanguage": [
        { "language": "kotlin", "count": 2 }
      ]
    },
    "effectiveModel": {
      "provider": "anthropic",
      "effort": "",
      "contextWindow": 1050000
    }
  }
}
```

### Fields

| Field | Meaning |
| - | - |
| `k` | Draws per bug used for this run. |
| `f1Score` | Harmonic mean of `precisionAvgK` and per-sample recall. |
| `recallAvgK` / `recallMaxK` | Fraction of known bugs detected — averaged over the k draws / best of the k draws. |
| `precisionAvgK` | Fraction of surfaced comments that are valid, averaged over k. |
| `signalToNoiseRatio` | Ratio of signal (valid, catalogued/high-severity findings) to noise (everything else). |
| `bySeverity.{critical,high,medium,low,none}` | Per-severity breakdown. Each has `recall`, `precision`, `f1Score`, and `precisionScored` (the number of scored comments at that severity — the precision denominator). `none` is precision-only (`recall`/`f1Score` are `null`); a severity with no catalogued bugs may report `null` recall. |
| `totalCost` / `costPerReview` | Cost in USD — total for the run, and per review. |
| `latencyMsP50/P90/Max` | Per-review latency (review generation) percentiles, in ms. |
| `totalDraws` / `okDraws` / `erroredDraws` | Review-attempt counts — one draw per dataset record × `k`. `okDraws` completed and were scored; `erroredDraws` failed (e.g. your endpoint erroring or returning an undecodable response) and were excluded from the metrics. |
| `errorRate` | `erroredDraws / totalDraws`. **A `completed` run can still have a non-zero `errorRate`** (up to the run-level ceiling) — the scores above are computed only over `okDraws`, so a high `errorRate` means the metrics cover a smaller, potentially skewed sample. Check this before trusting a score. |
| `errors` | **Present only when `erroredDraws > 0`** (omitted on a clean run). Characterizes the errored draws: `erroredDraws` (total), `contextWindowExceededDraws` (how many overflowed the context window), `byReason` (counts per failure-reason slug — one of `endpoint_unreachable`, `endpoint_auth_failed`, `endpoint_bad_output`, `context_window_exceeded`, `internal_error`), and `byLanguage` (counts per language). The **by-language** split is the most useful signal: a failure rate that concentrates in one language points to those *records* (large repos, slow clones), while one spread evenly points to a transient *load* problem. Both lists are sorted by count, descending. |
| `effectiveModel` | Echo of the resolved backend settings actually used — `provider`, `effort`, `contextWindow`. It never includes the key, the base URL, or the **model name** (you polled by `benchmarkId`, so the name is not echoed back). |

The top-level comparability fields appear on the status response (not inside `results`), once a run has a scored result:

| Field | Meaning |
| - | - |
| `comparable` | `true` for a control run; `false` for a run that supplied a custom `systemPrompt`. Omitted while the run is still `running`. |
| `customPrompt` | `true` iff the run supplied a custom detection `systemPrompt`. The cause of a `comparable: false`. |
| `comparabilityNote` | Present only on a custom-prompt run: a short, static explanation that the score reflects your own instructions and is off the public board. |

## Run lifecycle

<Steps>
  <Step title="Submit">
    `POST` → `202` with `benchmarkId` (malformed requests are rejected here with a
    `400` before any run starts — see Errors).
  </Step>

  <Step title="Eligibility probe">
    An eligibility probe validates the endpoint before a full run is spent on it: it
    dials your endpoint with your declared `effort` and checks auth, tool-calling,
    and capacity. If the endpoint is unreachable, unauthenticated, or can't serve the
    request (e.g. it rejects the effort you asked for, or the context window is below
    the minimum), the run terminates as `errored` with an explanatory `message` —
    this surfaces on a **poll**, not on the initial `POST`.
  </Step>

  <Step title="Running">
    Poll `GET .../{benchmarkId}`: `status: running` with `percentComplete` climbing.
  </Step>

  <Step title="Terminal">
    `status: completed` (with `results`) or `status: errored` (with a `message`).
  </Step>
</Steps>

## When a run errors

An `errored` run always carries a `message` naming the cause. The messages are a fixed
set — we never echo raw errors from your endpoint, so the text is safe to log and never
contains your `apiKey` or internal URLs. The ones you can act on:

| message (abridged) | what it means / what to do |
| - | - |
| "Authentication to the model endpoint failed" | Your `apiKey` was rejected (401/403). Check the key. |
| "The model endpoint was unreachable" | We couldn't reach `baseUrl`. Check the host and that it's publicly reachable over HTTPS. |
| "The declared context window is below the 128,000-token minimum" | Raise `contextWindow` to ≥ 128,000. |
| "…did not honor tool-calling on the readiness probe…" | Your endpoint ignored the forced tool call — confirm it supports tool use and the `effort` you requested. |
| "…responses could not be used (unparseable output…)" | Your endpoint returned output the harness couldn't parse as a review. |
| "Too many reviews errored… (see errorRate)" | Enough reviews failed that the score would be untrustworthy; the `results` still include `errorRate` and the partial metrics. |
| "Too many reviews exceeded the context window…" | Your `systemPrompt` plus the PR code did not fit the declared `contextWindow` on enough reviews to invalidate the run. Shorten the `systemPrompt` or declare a model with a larger context window. The `results` still include `errorRate` and the partial metrics. |
| "An internal error occurred…" | A fault on our side — retry, and contact support if it persists. |

A run rejected before it starts fails at `POST` with a `4xx` instead (see Errors).

## Rate limits

Limits are **per API key**:

| Limit | Default |
| - | - |
| Concurrent in-flight runs | 5 |
| Run starts | 500 / 24h (burst 10) |
| Status polls | 3600 / hour (burst 120) |

Exceeding a limit returns **`429 Too Many Requests`** with a `Retry-After` header
(seconds). Back off and retry.

## Errors

| Status | Meaning |
| - | - |
| `400 Bad Request` | Malformed body, missing required field, unsupported `provider`, a `baseUrl` that isn't a plain HTTPS URL / points at a disallowed host, a `costRate` outside `[0, 10000000]`, or a `systemPrompt` that is present-but-empty or over 50,000 bytes. |
| `401 Unauthorized` | Missing or malformed `Authorization: Bearer` header. |
| `403 Forbidden` | Invalid, revoked, or unauthorized key (valid header, but the key is not accepted for this operation). |
| `404 Not Found` | Unknown `benchmarkId`, a run that isn't yours, or the API is not enabled for your account. |
| `429 Too Many Requests` | Rate limit hit; see `Retry-After`. |

## Notes

* **`provider` is a wire protocol, not a vendor.** Any OpenAI-compatible endpoint
  (Fireworks, vLLM, etc.) uses `openai`; the Anthropic Messages API uses
  `anthropic`.
* **The harness is fixed.** You choose only the model backend and `k`; the
  dataset, prompts, tools, and scoring are identical for every run so scores are
  comparable across models. `benchmarkHarnessVersion` pins exactly which harness
  produced a result. The one exception is the optional `systemPrompt`: supplying it
  overrides the detection system prompt and makes the run **non-comparable** (see
  [Custom detection prompt](#custom-detection-prompt-non-comparable)) — it changes
  `benchmarkHarnessVersion` so those runs segregate automatically.
* **Key handling.** Your `apiKey` is encrypted at rest, used only to call your
  declared `baseUrl`, never logged or persisted in plaintext, and blanked when the
  run reaches a terminal state.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.