Skip to main content
A benchmark is only as trustworthy as its scenarios. A scenario that breaks the machine differently each time, leaves debris behind, or fails an agent for a valid fix produces numbers that look precise and mean nothing. The rules below come from real scenarios that went wrong. Follow them and most of your red runs will be about the agent, not the scenario.

Start from the incident

Before writing a command, write down three things:
  1. The root cause you will inject.
  2. What a correct fix looks like - and which fixes are also acceptable.
  3. What the agent must never do while fixing it.
If you have a runbook or a past incident report, take them from it: its root cause is your fault, its fix becomes your host_state check, its “never do” list becomes your negated rules. Put all three in the scenario’s description so the next person knows what the scenario is meant to test.

Target Setup

One deterministic fault

One root cause, injected the same way every time. A scenario that sometimes fills the disk to 91% and sometimes to 99%, or picks a random service to stop, gives you noise instead of a trend.

Break something the scenario owns

Inject the fault into something the scenario creates and can safely destroy: a loop-mounted filesystem, a throwaway container, a unit file it installs, a firewall rule with a recognizable comment. Use names, paths, and ports unique to the scenario. Before using a port, path, or container name, check that nothing real on the host uses it. A scenario that binds a port a production service already holds fails on every run, for reasons unrelated to the agent.

Clean up first

Start setup by removing whatever an earlier, interrupted run might have left: stop the process, unmount, delete the files, remove the rule. Then setup works on a dirty host too. Cleanup must not fail when there is nothing to clean:
On Windows, guard the command or end it with ; exit 0 - see Windows.

Prove the fault

End setup with a step that checks the incident is really there, and fails otherwise:
Without it you can grade an agent against a machine that was never broken - and it “passes” by doing nothing.

Each step is its own shell

Nothing carries over between steps: no variables, no cd. Recompute what you need in each step, or use scenario variables for values that never change.

Shared hosts mean shared paths

Two scenarios that mount the same directory, bind the same port, or install the same unit collide sooner or later - one’s leftovers break the other’s setup. Give each scenario its own resources. If they must share one, release it explicitly at the start of setup (stop whatever holds it, then unmount).

Target Restore

  • Revert everything setup did, and everything the agent might plausibly have added on top: a drop-in directory, an extra rule, a backup file.
  • Never fail on “already gone”. Restore stops at its first failing step, so one strict cleanup command leaves the rest of the mess behind. Use || true, rm -f, -ErrorAction with a guard.
  • Restore runs on every path - after a pass, a fail, an Error, or a stop - so it must work whatever state the machine is in.
A restore that exits 0 is not the same as a clean machine. See Keeping scenarios healthy for how to check.

The ticket

  • Describe symptoms, the way a user or a monitoring alert would. Never name the cause or the fix: “disk full on /opt/appdata”, not “delete the old dump in /opt/appdata/backups”.
  • Give the facts a real ticket would have: host name, IP, error message, since when.
  • For a diagnose-only scenario, put @2501:investigate in the ticket. The agent must diagnose without changing anything. Grade it on what it found and on the fault being still there afterwards.

Grading

Grade the outcome first, then policy, then method.

Outcome: did the machine get fixed

Use a host_state rule: a command run on the host after the job, which passes on exit code 0. Test what the user cares about, not how the agent got there. “Port 5432 is reachable from the app server” is a good check. “The DROP rule was deleted” is not: an agent that added a scoped ACCEPT rule above the DROP fixed the problem too, and a too-narrow check fails it for a valid fix.
For an investigate scenario, invert it: check that the fault is still there.

Resolution: how the job ended

A job_resolution_status rule, usually ["success"] - also for investigate scenarios, where a completed diagnosis is a success. Add partial only if a partial answer is acceptable.

Policy: what the agent must never do

Negated pattern_match rules on executed_commands: flush the firewall, reboot, rm -rf a data directory, SSH from one host to another. Make them required.

Method: how the agent worked

Positive pattern_match rules for evidence of a sound diagnosis, for example “checked inode usage, not only block usage”. Make them optional unless the method is the point: there are many valid ways to fix most incidents, and a required method rule fails the agent that found another one.

Efficiency

A loose task_count maximum catches an agent that thrashes.

Regular expression traps

pattern_match rules are where most unfair failures come from. Know how a pattern is matched - each value on its own, any match for a positive rule, no match at all for a negated one - then avoid these traps:
  • Anchor command words. A pattern matches anywhere inside a command, including inside an echo or a comment: \bssh\b.*10\.0\.0\.5 matches echo "=== ssh check ==="; curl 10.0.0.5. Anchor the command at a command boundary instead:
  • Match the destructive form in negated rules. A negated rule fails on any match, even a read-only command: rp_pool alone matches grep rp_pool haproxy.cfg. Write disable\s+server\s+rp_pool.
  • Don’t use task-level rules to prove nothing happened. With no task, they fail. To check that no task was created, use task_count with max: 0.
  • Escape for JSON. \s is written \\s inside a JSON string. The form does this for you; raw JSON does not.
Test every pattern before saving it, against at least three commands that must match and three that must not, for example with your browser console: new RegExp(pattern, 'i').test(command).

Before you trust a scenario

  1. Run Target Setup and Target Restore on their own, twice in a row. The second setup proves the scenario recovers from the first.
  2. Check the machine yourself after restore, not just the exit codes.
  3. Run one graded run and read every check: did each rule do what you meant?
  4. Run it several times. One run of an agent proves nothing.
  5. Then publish it.
When a run fails, sort out whether the scenario or the agent is at fault before changing anything - see Reading results.