What a run tells you
Open a scenario run from a benchmark or from the scenario’s history.
The run also links to the job the ticket created, where you can follow every task, command, and message, as with any job.
Error is not Fail
- Fail is a verdict: the agent ran and missed the bar. That is a result.
- Error means no verdict was produced: setup failed, the ticket never produced a job, the job ended in a technical failure, or a host stopped answering. Somebody has to fix something before the run says anything about the agent.
Who failed
Classify every red run before touching anything:
Most red runs on a new scenario fall in the first three rows. Never “fix” the agent for a broken scenario.
Common errors
When the agent failed
Only change your configuration for a genuine agent failure, and only when the change would be right in production too. The benchmark measures the configuration your real tickets run on: a rule added only to pass a scenario makes the score a lie.- The agent broke a policy - restarted a service without approval, flushed a firewall, deleted data instead of archiving it. If your organization really has that policy, add it as an operational rule, and add a negated rule to the scenario that checks it.
- The agent could not do the job - it lacked the domain knowledge or tools, or the work went to an agent whose specialty does not fit. Adjust the specialty or the agents using it. That changes every job those agents handle, not just this scenario.
- Neither - the model is not capable enough, or the ticket is ambiguous. Record it, try another model with a Main engine override, or make the ticket as clear as a real one would be.

