A run is asynchronous: you
POST a run, receive a benchmarkId, then poll a
status URL until it completes and returns the metrics.Authentication
Every request requires a bearer token issued to you by Macroscope:Start a run
Request body
modelBackend, every field is required except costRate. The top-level k and
systemPrompt are optional. There are otherwise no defaults: you declare effort and
capacities explicitly, so a run can’t be silently made non-comparable by a value we guessed
on your behalf.
Example request:
How the model name is handled
You send the realmodelName (and baseUrl) because we need them to call your endpoint.
They are handled like your apiKey: used only for the inference call, encrypted at rest, and
never written to the results, exports, logs, or API responses. The run is reported under
its opaque benchmarkId — the status response does not echo the model name back (you poll
by benchmarkId, so you already know which model it is). Because cost is attributed under that
id rather than a known model name, supply costRate to price a model we don’t price by default
(see below); otherwise the run is unpriced.
Pricing your model
If your model isn’t one we price by default, supplycostRate so totalCost /
costPerReview are populated instead of left unpriced. Each field is the price in
cents per million tokens — so a model that costs $3.00 per million input tokens has
inputCost: 300:
Each value must be in
[0, 10000000]; a bucket set to 0 is priced free. Omit the
whole costRate object to run unpriced (cost reported as 0).
Custom detection prompt (non-comparable)
By default every run uses the fixed control detection prompt, which is what makes scores comparable across models and customers. If you supplysystemPrompt, you replace
only the detection system prompt — your instructions for what to look for and how to
think. Everything else stays fixed:
- the harness still injects the PR code and the reference-graph context,
- it still provides the
report_issue/review_completetools and enforces the output schema (yoursystemPromptcannot change these), - the instruction prompt and the frozen validation judge that scores findings are unchanged, so a finding is scored exactly as it would be on a control run.
systemPrompt shares the model’s context window with the
PR code and the reference-graph context the harness injects. When your prompt plus the PR
code does not fit the declared contextWindow, that review errors — its result is not
scored, rather than being silently reviewed on a truncated pack. A review that errors counts
toward the run’s error rate like any other errored review; if enough reviews error, the whole
run is invalidated with a context-window reason (status errored, message naming the
context window — see the error table below). The 50,000-byte ceiling is a hard limit, not a
target: a focused prompt leaves room for the code under review, so more reviews fit the
window and are scored.
A run that supplies systemPrompt is not comparable to the public benchmark board.
Its score reflects your own detection instructions, not the control harness, so it is
segregated automatically:
- it runs under a distinct
benchmarkHarnessVersion(the custom prompt is part of the hashed harness surface), so it never shares a bucket with control runs in any export or dashboard; and - the status response marks it explicitly with
comparable: false/customPrompt: trueand acomparabilityNote(see below).
systemPrompt share a benchmarkHarnessVersion, so they
remain comparable to each other — just not to the board. Omit systemPrompt to run the
comparable control prompt.
Response — 202 Accepted
benchmarkId— opaque run id; use it to poll status.benchmarkHarnessVersion— content hash of the fixed harness (prompts, tools, scoring, standard layer). Two runs with the same value were scored identically; it changes only when the harness changes.benchmarkDatasetId— at submit time this is the dataset name; the status response reports the dataset content SHA once the run completes.
Get run status / results
status is completed (or errored). While running, results is
null and percentComplete reports progress.
Response
Fields
The top-level comparability fields appear on the status response (not inside
results), once a run has a scored result:
Run lifecycle
1
Submit
POST → 202 with benchmarkId (malformed requests are rejected here with a
400 before any run starts — see Errors).2
Eligibility probe
An eligibility probe validates the endpoint before a full run is spent on it: it
dials your endpoint with your declared
effort and checks auth, tool-calling,
and capacity. If the endpoint is unreachable, unauthenticated, or can’t serve the
request (e.g. it rejects the effort you asked for, or the context window is below
the minimum), the run terminates as errored with an explanatory message —
this surfaces on a poll, not on the initial POST.3
Running
Poll
GET .../{benchmarkId}: status: running with percentComplete climbing.4
Terminal
status: completed (with results) or status: errored (with a message).When a run errors
Anerrored run always carries a message naming the cause. The messages are a fixed
set — we never echo raw errors from your endpoint, so the text is safe to log and never
contains your apiKey or internal URLs. The ones you can act on:
A run rejected before it starts fails at
POST with a 4xx instead (see Errors).
Rate limits
Limits are per API key:
Exceeding a limit returns
429 Too Many Requests with a Retry-After header
(seconds). Back off and retry.
Errors
Notes
provideris a wire protocol, not a vendor. Any OpenAI-compatible endpoint (Fireworks, vLLM, etc.) usesopenai; the Anthropic Messages API usesanthropic.- The harness is fixed. You choose only the model backend and
k; the dataset, prompts, tools, and scoring are identical for every run so scores are comparable across models.benchmarkHarnessVersionpins exactly which harness produced a result. The one exception is the optionalsystemPrompt: supplying it overrides the detection system prompt and makes the run non-comparable (see Custom detection prompt) — it changesbenchmarkHarnessVersionso those runs segregate automatically. - Key handling. Your
apiKeyis encrypted at rest, used only to call your declaredbaseUrl, never logged or persisted in plaintext, and blanked when the run reaches a terminal state.