Start from the incident
Before writing a command, write down three things:- The root cause you will inject.
- What a correct fix looks like - and which fixes are also acceptable.
- What the agent must never do while fixing it.
host_state check, its “never do” list becomes your negated rules. Put all three in the scenario’s description so the next person knows what the scenario is meant to test.
Target Setup
One deterministic fault
One root cause, injected the same way every time. A scenario that sometimes fills the disk to 91% and sometimes to 99%, or picks a random service to stop, gives you noise instead of a trend.Break something the scenario owns
Inject the fault into something the scenario creates and can safely destroy: a loop-mounted filesystem, a throwaway container, a unit file it installs, a firewall rule with a recognizable comment. Use names, paths, and ports unique to the scenario. Before using a port, path, or container name, check that nothing real on the host uses it. A scenario that binds a port a production service already holds fails on every run, for reasons unrelated to the agent.Clean up first
Start setup by removing whatever an earlier, interrupted run might have left: stop the process, unmount, delete the files, remove the rule. Then setup works on a dirty host too. Cleanup must not fail when there is nothing to clean:; exit 0 - see Windows.
Prove the fault
End setup with a step that checks the incident is really there, and fails otherwise:Each step is its own shell
Nothing carries over between steps: no variables, nocd. Recompute what you need in each step, or use scenario variables for values that never change.
Shared hosts mean shared paths
Two scenarios that mount the same directory, bind the same port, or install the same unit collide sooner or later - one’s leftovers break the other’s setup. Give each scenario its own resources. If they must share one, release it explicitly at the start of setup (stop whatever holds it, then unmount).Target Restore
- Revert everything setup did, and everything the agent might plausibly have added on top: a drop-in directory, an extra rule, a backup file.
- Never fail on “already gone”. Restore stops at its first failing step, so one strict cleanup command leaves the rest of the mess behind. Use
|| true,rm -f,-ErrorActionwith a guard. - Restore runs on every path - after a pass, a fail, an Error, or a stop - so it must work whatever state the machine is in.
The ticket
- Describe symptoms, the way a user or a monitoring alert would. Never name the cause or the fix: “disk full on /opt/appdata”, not “delete the old dump in /opt/appdata/backups”.
- Give the facts a real ticket would have: host name, IP, error message, since when.
- For a diagnose-only scenario, put
@2501:investigatein the ticket. The agent must diagnose without changing anything. Grade it on what it found and on the fault being still there afterwards.
Grading
Grade the outcome first, then policy, then method.Outcome: did the machine get fixed
Use ahost_state rule: a command run on the host after the job, which passes on exit code 0.
Test what the user cares about, not how the agent got there. “Port 5432 is reachable from the app server” is a good check. “The DROP rule was deleted” is not: an agent that added a scoped ACCEPT rule above the DROP fixed the problem too, and a too-narrow check fails it for a valid fix.
Resolution: how the job ended
Ajob_resolution_status rule, usually ["success"] - also for investigate scenarios, where a completed diagnosis is a success. Add partial only if a partial answer is acceptable.
Policy: what the agent must never do
Negatedpattern_match rules on executed_commands: flush the firewall, reboot, rm -rf a data directory, SSH from one host to another. Make them required.
Method: how the agent worked
Positivepattern_match rules for evidence of a sound diagnosis, for example “checked inode usage, not only block usage”. Make them optional unless the method is the point: there are many valid ways to fix most incidents, and a required method rule fails the agent that found another one.
Efficiency
A loosetask_count maximum catches an agent that thrashes.
Regular expression traps
pattern_match rules are where most unfair failures come from. Know how a pattern is matched - each value on its own, any match for a positive rule, no match at all for a negated one - then avoid these traps:
- Anchor command words. A pattern matches anywhere inside a command, including inside an
echoor a comment:\bssh\b.*10\.0\.0\.5matchesecho "=== ssh check ==="; curl 10.0.0.5. Anchor the command at a command boundary instead: - Match the destructive form in negated rules. A negated rule fails on any match, even a read-only command:
rp_poolalone matchesgrep rp_pool haproxy.cfg. Writedisable\s+server\s+rp_pool. - Don’t use task-level rules to prove nothing happened. With no task, they fail. To check that no task was created, use
task_countwithmax: 0. - Escape for JSON.
\sis written\\sinside a JSON string. The form does this for you; raw JSON does not.
new RegExp(pattern, 'i').test(command).
Before you trust a scenario
- Run Target Setup and Target Restore on their own, twice in a row. The second setup proves the scenario recovers from the first.
- Check the machine yourself after restore, not just the exit codes.
- Run one graded run and read every check: did each rule do what you meant?
- Run it several times. One run of an agent proves nothing.
- Then publish it.

