> ## Documentation Index
> Fetch the complete documentation index at: https://docs.2501.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Reading results

> Read a scenario run, tell a broken scenario from a failing agent, and decide what to change

A red run is a question, not an answer. Before changing anything - the scenario, an operational rule, a specialty - work out **who failed**: the scenario, the platform, or the agent.

## What a run tells you

Open a scenario run from a benchmark or from the scenario's history.

| Field | Meaning |
| - | - |
| **Status** | **Completed**, **Error**, or **Cancelled**: did the run reach a verdict. |
| **Verdict** | **Pass** or **Fail**, when the run completed. |
| **Compliance** | The share of rules that passed, from 0 to 100. |
| **Remediation** | Whether the job resolved as `success`. |
| **Checks** | Each rule with its result and what it inspected: the matching command, the exit code and output of a `host_state` command, the resolution status. |
| **Restore failed** | Shown when restore did not put the host back. |

The run also links to the job the ticket created, where you can follow every task, command, and message, as with any job.

### Error is not Fail

* **Fail** is a verdict: the agent ran and missed the bar. That is a result.
* **Error** means no verdict was produced: setup failed, the ticket never produced a job, the job ended in a technical failure, or a host stopped answering. Somebody has to fix something before the run says anything about the agent.

## Who failed

Classify every red run before touching anything:

| What you see | Who failed | What to do |
| - | - | - |
| **Error**, and the errors name a setup step | **The scenario's setup** | Fix Target Setup. The agent never ran. |
| **Restore failed** | **The scenario's restore** | Fix Target Restore now: every later run on that host scores against a broken machine. Then clean the host with **Target Restore**. |
| A check failed although the agent's work was right - a valid alternative fix, a read-only command that matched a negated rule | **The scenario's grading** | Fix the rule. Prefer a `host_state` check on the outcome. |
| The checks are right and the agent did the wrong thing, or gave up | **The agent** | A real signal - see [below](#when-the-agent-failed). |
| **Error** with a job timeout, a technical job failure, *"not part of this benchmark run"*, *"produced no job"*, or transport errors | **The platform or the setup around it** | Re-run before changing the scenario. If it repeats, see the table below. |

Most red runs on a **new** scenario fall in the first three rows. Never "fix" the agent for a broken scenario.

### Common errors

| Error | Cause |
| - | - |
| *Agent ... is not part of this benchmark run* | The ticket went to an agent the scenario does not list, often a second agent on the same host. Add it to **Evaluated Agents**. |
| *This organization has no runner gateway* | Create a Runner gateway in **Gateways**. |
| *Host ... has no remote execution agent* | The host has no agent that can run commands. Add one. |
| *Ticket ... produced no job* | The gateway did not open a job for the ticket - check the gateway's configuration and inbound prompt. |
| *No job appeared for ticket ... within 600s* | With ServiceNow: the gateway never ingested the incident. Check the [ServiceNow metadata](/0.16/benchmark/command-center#running-through-servicenow) and that the gateway is active. |
| *Job ... did not finish within 2700s* | The job ran over 45 minutes. Look at the job to see where it stalled. |

## When the agent failed

Only change your configuration for a genuine agent failure, and only when the change would be right in production too. The benchmark measures the configuration your real tickets run on: a rule added only to pass a scenario makes the score a lie.

* **The agent broke a policy** - restarted a service without approval, flushed a firewall, deleted data instead of archiving it. If your organization really has that policy, add it as an [operational rule](/0.16/configure/operational-rules), and add a negated rule to the scenario that checks it.
* **The agent could not do the job** - it lacked the domain knowledge or tools, or the work went to an agent whose specialty does not fit. Adjust the specialty or the agents using it. That changes every job those agents handle, not just this scenario.
* **Neither** - the model is not capable enough, or the ticket is ambiguous. Record it, try another model with a **Main engine** override, or make the ticket as clear as a real one would be.

## Look at the pattern, not one run

Agents are not deterministic. One failure out of ten runs and ten out of ten are different findings. Run a scenario several times before drawing conclusions, and use the trend on the Benchmarks page to see whether a change made things better or worse.

To compare models, run the same set of scenarios once per **Main engine** override and filter the Benchmarks page by engine.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.