Grade an eval run with LLM as Judge
Spends. Runs the goal-completion judge over the finished run, scoring each case’s final answer against its expected output.
202: scheduled, not done. Read the grades from the run detail’s judges.goalCompletion rather than re-requesting — a second POST only spends again.
A run’s grading config is pinned when the run is created, so turning the judge on for the suite does not reach an already-recorded run: enable: true is what grades one, and it changes nothing beyond that run. Omitting model and threshold clears any override a previous request left on the run.
Authorizations
MCPJam API key (sk_…). Create one at Settings → API keys. Guest sessions cannot use the API, and API keys cannot manage other API keys.
Headers
Which vocabulary this request and its response speak. Absent means 1, which is byte-for-byte today's contract: the same request fields, the same refusals, the same response projection. 2 is the canonical vocabulary. Any other value is a 400 with code: "VALIDATION_ERROR".
Today it decides one thing: the spelling of an evaluator's policy role. Vocabulary 1 accepts and returns gating; vocabulary 2 accepts both spellings and returns the canonical required. Sending required without the header is a 400, deliberately — vocabulary 1 is not widened to meet vocabulary 2 half way, because a boundary that accepts a spelling it does not announce is one two implementations can disagree about.
A response that varies by vocabulary sends Vary: x-mcpjam-eval-vocabulary.
1, 2 Path Parameters
ID of the hosted project that contains the server.
Eval run ID, as returned by POST /eval-runs.
Body
Optional. A bodyless POST grades with the suite's own config.
Re-grade a run that already has a result. SPENDS AGAIN. Must be a real boolean — "false" is rejected rather than read as consent.
Grade this run even though the judge was off when it ran. A per-RUN answer, not a suite edit: grading reads the config pinned when the run was created, so enabling the judge on the suite does not reach an already-recorded run.
Judge model for this run only.
Pass threshold for this run only.
0 <= x <= 1
