> ## Documentation Index
> Fetch the complete documentation index at: https://docs.2501.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Benchmarks from the API

> Create scenarios, test them, run benchmarks, and poll results over /api/v1

Everything the Benchmarks pages do to scenarios is available under `/api/v1`, so you can keep scenarios as code and run benchmarks from a pipeline. This page walks through the usual flow; every endpoint, field, and response is in the **API Reference** under **Scenarios** and **Benchmarks**.

Authenticate with an API key as a bearer token - see [API Overview](/0.16/api/overview#authentication).

```bash theme={null}
export CC=https://<your-command-center-host>
export API_KEY=2501_ak_...
export ORG=org_460b541c-...
```

## 1. Create a scenario

Send the whole [scenario document](/0.16/benchmark/scenario), plus `org_id`. Hosts and agents are referenced by name; an unknown name is a `400`, a duplicate title a `409`. A new scenario is a `draft` unless you say otherwise.

```bash theme={null}
curl -X POST "$CC/api/v1/scenarios" \
  -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
  -d @- <<EOF
$(jq --arg org "$ORG" '. + {org_id: $org}' disk-large-file.json)
EOF
```

The response is the saved scenario with its `id` (`scn_...`).

To find host and agent names, use `GET /api/v1/hosts?org_id=$ORG` and `GET /api/v1/agents?org_id=$ORG`.

## 2. Test setup and restore

`POST /api/v1/scenarios/{id}/run` without an agent runs only the phases you ask for, on **every** host of the scenario. Nothing is sent to an agent and nothing is graded.

```bash theme={null}
# Break the machine and leave it broken, to inspect it
curl -X POST "$CC/api/v1/scenarios/$SCENARIO/run" \
  -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
  -d '{"phases": ["prepare"]}'
```

```json theme={null}
{ "run_id": "wfl_7d3b9c2e-...", "workflow_id": "wfl_7d3b9c2e-...", "phases": ["prepare"], "with_agent": false }
```

| Body | Effect |
| - | - |
| `{"phases": ["prepare"]}` | Target Setup only. |
| `{"phases": ["restore"]}` | Target Restore only - also how you clean a dirty host. |
| `{}` | Both, setup then restore. |
| `"host": "db-01"` | Limit the phase to the steps of one host. |

Run setup then restore twice in a row before trusting a scenario - see [Before you trust a scenario](/0.16/benchmark/designing#before-you-trust-a-scenario).

## 3. Poll a run

```bash theme={null}
curl "$CC/api/v1/scenarios/$SCENARIO/runs/$RUN_ID" -H "Authorization: Bearer $API_KEY"
```

```json theme={null}
{
  "status": "completed",
  "execution_status": null,
  "evaluation_status": "not_evaluated",
  "passing": null,
  "restore_status": null,
  "errors": null,
  "job_id": null,
  "workflow_id": "wfl_7d3b9c2e-...",
  "workflow_status": "completed",
  "steps": [
    {
      "name": "Confirming disk pressure",
      "status": "completed",
      "duration_ms": 412,
      "record": { "label": "Confirming disk pressure", "host": "app-01", "status": "ok", "command": "test ...", "rc": 0, "stdout": "", "stderr": "" }
    }
  ]
}
```

`steps` grows as the run progresses: each record carries the host, the command (secrets redacted), the exit code, the output, and an `error` when the step failed. A phase run has no verdict: `passing` stays `null`, and the step records are what tell you whether your commands did what you meant.

Poll every few seconds until `status` is `completed`, `failed`, or `cancelled`.

## 4. Run it graded

Add `"with_agent": true` to hand the ticket to the agent and grade the run. It always runs both phases, on every host.

```bash theme={null}
curl -X POST "$CC/api/v1/scenarios/$SCENARIO/run" \
  -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
  -d '{"with_agent": true, "gateway": "runner", "main_engine": "<model-key>"}'
```

```json theme={null}
{ "run_id": "scr_ca850e32-...", "benchmark_id": "bch_0f4a6d1b-...", "report_id": "scr_ca850e32-...", "workflow_id": "wfl_...", "phases": ["prepare", "restore"], "with_agent": true }
```

Poll `run_id` as above. Once the run completes, `passing` holds the verdict, `evaluation_status` is `passed` or `failed`, and `job_id` points to the job the ticket created. `gateway`, `main_engine`, `secondary_engine`, and `gateway_engine` are optional and work as in the [run dialog](/0.16/benchmark/command-center#running-a-benchmark).

## 5. Run a benchmark

```bash theme={null}
curl -X POST "$CC/api/v1/benchmarks" \
  -H "Authorization: Bearer $API_KEY" -H "Content-Type: application/json" \
  -d '{
    "org_id": "'"$ORG"'",
    "scenario_ids": ["scn_...", "scn_..."],
    "gateway": "runner",
    "main_engine": "<model-key>"
  }'
```

```json theme={null}
{ "benchmark_id": "bch_...", "workflow_id": "wfl_...", "report_ids": ["scr_...", "scr_..."] }
```

`report_ids` are in the order of `scenario_ids`. Poll each one at `/api/v1/scenarios/{scenario_id}/runs/{report_id}`. Up to 50 scenarios per benchmark; add `"sequential": true` to run them one at a time.

There is no scheduling: to run a benchmark nightly, call this endpoint from your scheduler or CI.

## 6. Stop a run

```bash theme={null}
curl -X POST "$CC/api/v1/scenarios/$SCENARIO/runs/$RUN_ID/stop" -H "Authorization: Bearer $API_KEY"
```

```json theme={null}
{ "run_id": "scr_...", "cancelled": true, "reason": "cancelling" }
```

`reason` is `cancelling`, `never_started`, or `already_ending`. A stopped run still runs its restore.

## Update and delete

* `POST /api/v1/scenarios/{id}` **replaces the whole document** - there is no partial update. `GET` the scenario, change the JSON, and send it back; the server-owned fields the `GET` returns are accepted and ignored.
* `DELETE /api/v1/scenarios/{id}` removes the scenario. Its past runs stay in the results.

An API key acts as an administrator within its scope, so it can do all of the above. In Command Center, the same actions follow your [role](/0.16/benchmark/overview#who-can-do-what).

## Keep scenarios as code

Scenarios are plain JSON, so you can keep a folder of them in a repository and sync it to an organization:

1. `GET /api/v1/scenarios?org_id=$ORG` to list what exists, by title.
2. For each file, `POST /api/v1/scenarios/{id}` if a scenario with that title exists, otherwise `POST /api/v1/scenarios`.
3. Never delete from the sync: retire scenarios by setting them `disabled`.

Matching by title means a rename in the files creates a new scenario with a fresh history. Rename in Command Center instead, or update by `id`.

## What the API does not cover yet

* **Per-rule results.** The run response gives the verdict and the step records, not each validation check. Read those on the run's page in Command Center.
* **Compliance and remediation** are shown in Command Center, not in the run response.
* **Benchmarks as a whole.** You can launch a benchmark, but not list, read, stop, or delete one: poll and stop its runs individually, and delete it from Command Center.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.