Skip to main content
Run Macroscope’s standardized code-review benchmark against your own model endpoint (bring-your-own-key) and get back recall / precision / F1 and related metrics. The dataset, review harness, prompts, and scoring are fixed and identical for every customer, so results are directly comparable across models — you supply only the model endpoint.
A run is asynchronous: you POST a run, receive a benchmarkId, then poll a status URL until it completes and returns the metrics.

Authentication

Every request requires a bearer token issued to you by Macroscope:
Keys are provisioned by Macroscope (they are not self-service). Treat the token as a secret; it identifies and rate-limits your account. The base URL is provided with your key (per environment). All paths below are relative to it.

Start a run

Request body

Inside modelBackend, every field is required except costRate. The top-level k and systemPrompt are optional. There are otherwise no defaults: you declare effort and capacities explicitly, so a run can’t be silently made non-comparable by a value we guessed on your behalf. Example request:

How the model name is handled

You send the real modelName (and baseUrl) because we need them to call your endpoint. They are handled like your apiKey: used only for the inference call, encrypted at rest, and never written to the results, exports, logs, or API responses. The run is reported under its opaque benchmarkId — the status response does not echo the model name back (you poll by benchmarkId, so you already know which model it is). Because cost is attributed under that id rather than a known model name, supply costRate to price a model we don’t price by default (see below); otherwise the run is unpriced.

Pricing your model

If your model isn’t one we price by default, supply costRate so totalCost / costPerReview are populated instead of left unpriced. Each field is the price in cents per million tokens — so a model that costs $3.00 per million input tokens has inputCost: 300: Each value must be in [0, 10000000]; a bucket set to 0 is priced free. Omit the whole costRate object to run unpriced (cost reported as 0).

Custom detection prompt (non-comparable)

By default every run uses the fixed control detection prompt, which is what makes scores comparable across models and customers. If you supply systemPrompt, you replace only the detection system prompt — your instructions for what to look for and how to think. Everything else stays fixed:
  • the harness still injects the PR code and the reference-graph context,
  • it still provides the report_issue / review_complete tools and enforces the output schema (your systemPrompt cannot change these),
  • the instruction prompt and the frozen validation judge that scores findings are unchanged, so a finding is scored exactly as it would be on a control run.
Keep it reasonably short. Your systemPrompt shares the model’s context window with the PR code and the reference-graph context the harness injects. When your prompt plus the PR code does not fit the declared contextWindow, that review errors — its result is not scored, rather than being silently reviewed on a truncated pack. A review that errors counts toward the run’s error rate like any other errored review; if enough reviews error, the whole run is invalidated with a context-window reason (status errored, message naming the context window — see the error table below). The 50,000-byte ceiling is a hard limit, not a target: a focused prompt leaves room for the code under review, so more reviews fit the window and are scored. A run that supplies systemPrompt is not comparable to the public benchmark board. Its score reflects your own detection instructions, not the control harness, so it is segregated automatically:
  • it runs under a distinct benchmarkHarnessVersion (the custom prompt is part of the hashed harness surface), so it never shares a bucket with control runs in any export or dashboard; and
  • the status response marks it explicitly with comparable: false / customPrompt: true and a comparabilityNote (see below).
Two runs that supply the same systemPrompt share a benchmarkHarnessVersion, so they remain comparable to each other — just not to the board. Omit systemPrompt to run the comparable control prompt.

Response — 202 Accepted

  • benchmarkId — opaque run id; use it to poll status.
  • benchmarkHarnessVersion — content hash of the fixed harness (prompts, tools, scoring, standard layer). Two runs with the same value were scored identically; it changes only when the harness changes.
  • benchmarkDatasetId — at submit time this is the dataset name; the status response reports the dataset content SHA once the run completes.

Get run status / results

Poll until status is completed (or errored). While running, results is null and percentComplete reports progress.

Response

Fields

The top-level comparability fields appear on the status response (not inside results), once a run has a scored result:

Run lifecycle

1

Submit

POST → 202 with benchmarkId (malformed requests are rejected here with a 400 before any run starts — see Errors).
2

Eligibility probe

An eligibility probe validates the endpoint before a full run is spent on it: it dials your endpoint with your declared effort and checks auth, tool-calling, and capacity. If the endpoint is unreachable, unauthenticated, or can’t serve the request (e.g. it rejects the effort you asked for, or the context window is below the minimum), the run terminates as errored with an explanatory message — this surfaces on a poll, not on the initial POST.
3

Running

Poll GET .../{benchmarkId}: status: running with percentComplete climbing.
4

Terminal

status: completed (with results) or status: errored (with a message).

When a run errors

An errored run always carries a message naming the cause. The messages are a fixed set — we never echo raw errors from your endpoint, so the text is safe to log and never contains your apiKey or internal URLs. The ones you can act on: A run rejected before it starts fails at POST with a 4xx instead (see Errors).

Rate limits

Limits are per API key: Exceeding a limit returns 429 Too Many Requests with a Retry-After header (seconds). Back off and retry.

Errors

Notes

  • provider is a wire protocol, not a vendor. Any OpenAI-compatible endpoint (Fireworks, vLLM, etc.) uses openai; the Anthropic Messages API uses anthropic.
  • The harness is fixed. You choose only the model backend and k; the dataset, prompts, tools, and scoring are identical for every run so scores are comparable across models. benchmarkHarnessVersion pins exactly which harness produced a result. The one exception is the optional systemPrompt: supplying it overrides the detection system prompt and makes the run non-comparable (see Custom detection prompt) — it changes benchmarkHarnessVersion so those runs segregate automatically.
  • Key handling. Your apiKey is encrypted at rest, used only to call your declared baseUrl, never logged or persisted in plaintext, and blanked when the run reaches a terminal state.