> ## Documentation Index
> Fetch the complete documentation index at: https://docs.2501.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Validation

> Validation rules, validators, and scoring model

Validation rules determine whether a scenario passed. They are evaluated after the agent finishes executing.

## Command Center scenarios

Scenarios stored in Command Center use a flat `validation` array. The engine always checks that the job completed, even when the array is empty.

```json theme={null}
"validation": [
  {
    "label": "Agent restarted nginx",
    "validator": "pattern_match",
    "pattern": "systemctl.*restart.*nginx",
    "where": "executed_commands"
  },
  {
    "label": "Job resolved acceptably",
    "validator": "job_resolution_status",
    "statuses": ["success", "partial"]
  },
  {
    "label": "Nginx is active",
    "validator": "host_state",
    "run": "systemctl is-active nginx"
  }
]
```

The available validators are:

| Validator | Fields | What it checks |
| - | - | - |
| `pattern_match` | `pattern`, `where`, optional `negate` | A case-insensitive regular expression against gateway, job, or task evidence. |
| `task_count` | optional `min`, optional `max` | The number of tasks created for the job. At least one bound is required. |
| `job_resolution_status` | `statuses` | The final resolution status. More than one status can be accepted. |
| `host_state` | `run`, optional `on`, optional `exit_code` | A shell command on a scenario host. It passes when the command returns the expected exit code, which defaults to `0`. |

Every rule needs a human-readable `label`. `required` defaults to `true`; an optional rule still contributes to the compliance score but does not fail the run.

`pattern_match` can inspect these targets:

| Scope | Targets |
| - | - |
| Gateway | `gateway_messages`, `gateway_summary` |
| Job | `job_resolution`, `job_plan` |
| Task | `task_summary`, `task_description`, `executed_commands`, `write_file`, `update_file`, `use_mcp_tool`, `tool_calls`, `agent_messages`, `operational_rules`, `task_plan` |

Setup and restoration are also plain shell. Give each step a `name` so reports describe the action rather than displaying the command as its label:

```json theme={null}
"prepare": [
  {
    "name": "Stopping nginx service",
    "run": "sudo systemctl stop nginx"
  },
  {
    "name": "Replacing nginx configuration",
    "write": "/etc/nginx/site.conf",
    "content": "server_name example.test;\n",
    "sudo": true
  }
]
```

A `write` step creates or replaces the complete file. It does not append or patch existing content.

## CLI runner scenarios

The classic file-based runner declares validation in the grouped `validation` object of `scenario.json`.

***

## Structure

```json theme={null}
"validation": {
  "gateway": [...],
  "job": [...],
  "tasks": [...]
}
```

Rules are organized into three scopes based on what they check:

| Scope | When used | What it checks |
| - | - | - |
| `gateway` | `--from gateway` only | The ServiceNow ticket |
| `job` | All entry points | The job record (status, plan, task count) |
| `tasks` | All entry points | Per-task data (commands, summaries, plans) |

For `tasks` rules, each rule is checked against every task. A rule passes if it matches on at least one task.

***

## Rule Structure

Every rule shares a common set of fields:

```json theme={null}
{
  "label": "Agent restarted nginx",
  "validator": "pattern_match",
  "pattern": "systemctl.*(restart|reload|start).*nginx",
  "where": "executed_commands",
  "required": true,
  "negate": false
}
```

| Field | Default | Description |
| - | - | - |
| `label` | - | Required. Shown in the validation report. Make it descriptive. |
| `validator` | - | Required. The type of check to run. See validators below. |
| `required` | `true` | When `false`, the rule is informational: it contributes to the compliance score but does not block a pass. |
| `negate` | `false` | Invert the result. The rule passes when the condition is NOT met. |

***

## Validators

### `pattern_match`

Checks whether a regex pattern matches in a specific field of the job or task data.

```json theme={null}
{
  "label": "Agent edited the nginx config",
  "validator": "pattern_match",
  "pattern": "/etc/nginx/",
  "where": "executed_commands"
}
```

| Field | Description |
| - | - |
| `pattern` | Regular expression. Matching is case-insensitive. |
| `where` | The field to search. |

**`where` targets:**

| Target | Content |
| - | - |
| `executed_commands` | All commands run by the agent, one per line |
| `tool_calls` | Every tool call as `<name> <args-json>` - the only target that sees the tools carrying neither a command nor a path, such as the terminal session and background-job families |
| `task_summary` | The task's own account of how it ended - why it stopped, or what it got done before it was terminated |
| `task_description` | The task description as created |
| `task_plan` | Always empty: the agent now plans inside its first turn and no longer writes a separate plan. Kept so existing scenarios still validate |
| `agent_messages` | Full agent reasoning history |
| `job_resolution` | The job's resolution summary |
| `job_plan` | The job-level plan |
| `gateway_messages` | Messages posted by the gateway bot on the ticket |
| `gateway_summary` | The gateway's summary of ticket resolution |
| `operational_rules` | Operational constraints from the agent's context |

***

### `job_resolution_status`

Checks the job's final resolution status.

```json theme={null}
{
  "label": "Job resolved successfully",
  "validator": "job_resolution_status",
  "pattern": "success"
}
```

`pattern` fixes one status. `statuses` accepts any of several, so a read-only scenario the agent ends as either `success` or `partial` does not fail on the difference:

```json theme={null}
{
  "label": "Job resolved",
  "validator": "job_resolution_status",
  "statuses": ["success", "partial"]
}
```

Allowed values: `success`, `agentic_failure`, `hard_failure`, `partial`, `no_action`.

***

### `ticket_status`

Checks the status of the ServiceNow ticket. Only applicable with `--from gateway`.

```json theme={null}
{
  "label": "Ticket was resolved",
  "validator": "ticket_status",
  "pattern": "resolved"
}
```

***

### `task_count`

Verifies that the number of tasks created under the job falls within a range.

```json theme={null}
{
  "label": "Resolved efficiently",
  "validator": "task_count",
  "min": 1,
  "max": 3
}
```

***

### `ansible`

Runs an Ansible playbook and treats its exit code as pass/fail. The most reliable way to assert actual machine state.

```json theme={null}
{
  "label": "Nginx is running and serving traffic",
  "validator": "ansible",
  "ansiblePath": "validate.yml"
}
```

A non-zero exit code fails the resolution gate and marks the scenario as failed. See [Playbooks](/0.16/benchmark/playbooks) for how to write `validate.yml`.

***

## Scoring Model

**Compliance score**: percentage of all non-Ansible rules that passed (required + optional combined). Informational.

Two gates determine the actual pass/fail result:

**Compliance gate**: passes when every `required` non-Ansible rule passes.

**Resolution gate**: if a `validate.yml` Ansible rule exists, passes when the playbook exits 0. If no `validate.yml` rule is defined, it passes when the compliance gate passes and the compliance score is at least 80%.

**A scenario passes only when both gates pass.**


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.