Drift: the silent failure
Drift is anything left on a host that should not be there: a file a restore forgot, a mount stacked on another mount, a unit an agent installed, a config block from an earlier run. It is dangerous because nothing turns red. Restore exits 0, the next setup exits 0, and the scenario now grades the agent against a different machine than the one you designed. Real examples:- A dirty baseline. A scenario saved a copy of a config file at setup, broke the file, and put the copy back at restore. One run’s restore was skipped after an agent had “fixed” the file. The next setup saved the already-modified file as its clean copy. From then on, every restore faithfully put back the wrong config - with exit code 0.
- Stacked mounts. A scenario’s setup unmounted a stale mount once, then mounted its own on top. A foreign mount from another scenario stayed underneath and resurfaced after every restore.
- Orphans from another scenario. Someone ran a scenario’s setup to test it and never ran its restore. Its inactive services sat on a host used by other scenarios - a perfect red herring for the agent.
- What the agent adds. Restore removed the unit file its setup installed, but not the
override.confdrop-in directory the agent created next to it, nor the editor swap file the agent left behind.
Prevent it in the scenario
Most drift is designed in. The design rules prevent it:- Don’t snapshot and restore whole files. Saving a file at setup and copying it back at restore captures whatever was there - including leftovers. Instead, change something the scenario owns and remove exactly that: a marked config block, a tagged firewall rule, a dedicated loop filesystem. See the firewall example.
- Clean up at the start of setup, not only in restore. If restore was skipped, setup still starts from a known state.
- Release shared resources completely. Unmount until nothing is mounted, stop whatever holds a directory before unmounting it (
fuser -km), detach every loop device on your backing file. - Restore what the agent may have added, not only what setup did: drop-in directories, extra rules, backup copies.
- End setup with a check that the fault is there, and check values your scenario controls. A value the scenario never writes - a timeout of
10swhen setup writes2s- is proof the host was dirty.
Verify restores against the machine, not the exit code
A restore that exits 0 only means its last command succeeded. After changing a scenario, and periodically after that:- Run Target Setup, then Target Restore, from the scenario’s page.
- Check the host yourself: compare the files, services, mounts, and rules the scenario touches against how the host should look.
- Run both again. The second setup proves the scenario recovers from the first.
When a run says Restore failed
Fix it right away: every later run on that host scores against a broken machine.- Open the run and read which restore step failed and its output.
- Fix the step. Usually it is a cleanup that fails on “already gone” - add
|| true,rm -f, or a guard. - Run Target Restore from the scenario’s page until the host is clean.
- Treat runs on that host since the failure with suspicion.
Keep results comparable
- A new scenario starts a new history. Run history follows the scenario, not its name, so renaming it is safe. Deleting a scenario and recreating it, or a sync script that matches scenarios by title after a rename, starts from zero.
- Changing a scenario changes what it measures. Every save is a new version, and every run is graded against the version it launched with. When you change the fault or the grading substantially, note it - the pass rate before and after are not the same measurement.
- Change one thing at a time. To measure the effect of a model, a specialty, or an operational rule change, run the same scenario set before and after, with nothing else changed.
- Watch the trend, not one benchmark. Agents are not deterministic. A dip over one benchmark may be noise; the same dip over several is a regression.
Housekeeping
- Retire, don’t delete. Set a scenario you no longer run to
disabled: it stays visible and editable, and cannot be launched by accident. Deleting a scenario keeps its past runs but loses its definition. - Delete test benchmarks. Benchmarks you launched only to test a scenario add noise to your trends. Delete benchmark removes the benchmark with its runs and the jobs, tasks, and tickets it created.
- Keep scenarios as code. Scenarios are plain JSON. Keep them in a repository and sync them with the API, so changes are reviewed and every environment runs the same definitions.
- Watch the hosts’ agents. A scenario’s hosts and agents cannot be deleted while a scenario references them. When you replace a host, update the scenarios that target it.

