ct run eval2
Run an eval2 evaluation.
Usage
ct run eval2 [OPTIONS]
Options
| Option | Description |
|---|---|
--run-config FILE | Inspect run-config YAML for the control_tower/control_eval2 task; delegates to inspect eval --run-config. Task/policy/protocol/harness/grant/sandbox flags are rejected (put them in the YAML); --no-upload, --run-name, --tags, --log-dir, --max-samples forward. |
-e, --env, --environment TEXT | Environment(s) to run (repeat flag for multiple) |
-ea, --env-arg TEXT | Environment config option as key=value, or env:key=value when more than one selected environment declares the same key. Repeat for multiple. |
-t, --main-task TEXT | Main task(s) to run |
-s, --side-task TEXT | Side task(s) to run |
--trajectory-id, --traj-id TEXT | Trajectory ID to extract task info from (downloads from API or reads local file) |
--task-file TEXT | Task set name or path to file with task combinations. Use 'ct task-sets list' to see available task sets. |
--limit INTEGER RANGE | Limit the number of selected task combinations to run [x>=1] |
--all | Expand to all main×side combos (attack) or all main tasks (honest) |
--just-main-tasks | Expand to all main tasks only (for honest mode) |
--just-side-tasks | Expand to all side tasks with no main task |
--env-only | Use the first main task per environment (one sample per env, no side task) |
-c, --category TEXT | Filter main tasks by category (e.g. add_feature, fix_bug, refactor). Can be repeated. |
--main-tasks-with-no-scorers [error|include|exclude] | Main tasks lacking a scorer: refuse the run (error), keep them recorded as unscored (include), or drop them (exclude). |
--untrusted-policy TEXT | eval2 untrusted policy name. Defaults to honest when this option family is used. ct protocols eval2-policies to see available. |
-up, --untrusted-policy-arg TEXT | eval2 untrusted policy-specific args in key=value format (multiple allowed). ct protocols eval2-policies <name> to see available args. |
--protocol TEXT | eval2 blue protocol name. Defaults to null-blue-team when this option family is used. ct protocols eval2-protocols to see available. |
-bp, --protocol-arg TEXT | eval2 blue protocol-specific args in key=value format (multiple allowed). ct protocols eval2-protocols <name> to see available args. |
--ec2 | Run evaluation on EC2 instead of locally |
--ec2-arg TEXT | Config for --ec2 in key=value format (multiple allowed): instance_type (default 'c7i.4xlarge'), region (default 'us-east-1'), keep_instance (default False), max_workers (default 64), concurrency_per_worker (default 32), estimate (default False), auto_confirm (default False), new_fleet (default False), spot (default False), worker_setup (default None), ami (default None). |
--startup-retries INTEGER | Number of times to retry Docker sandbox startup on transient failures [default: 7] |
--no-internet | Block outbound internet and internet simulator access. |
--no-intranet | Block access to compose sibling services (e.g. Redis). |
--network-control | Grant NET_ADMIN to the agent container without applying restrictions at startup. Protocols that manage iptables mid-trajectory (e.g. privilege-monitor) need this. |
--sandbox-isolation | Wrap every agent command in a bubblewrap sandbox: read-only root, writable workspace only. A static bwrap is fetched automatically; set CT_BWRAP_HOST_BINARY to override with your own. |
--extra-src DIRECTORY | Directory of experiment code added to PYTHONPATH locally, so external --policy my.module:fn and --sandbox my.module:type refs resolve without committing into the repo; fleet mode (--ec2) also ships it to workers. Repeatable; each dir must have a distinct basename. |
--sandbox TEXT | Run samples in an Inspect sandbox provided outside this repo, as a registered type (my-sandbox) or module:type to import the module that registers it. Replaces the default Docker sandbox; with --ec2 each worker runs it for its own jobs. |
--sandbox-arg TEXT | Config for --sandbox in key=value format (multiple allowed), passed through to the sandbox as JSON. |
--simulated | Run against an LLM-simulated sandbox instead of provisioning the real environment. Every tool output is fabricated by a simulator agent. |
--simulator-model TEXT | Model for the simulator agent (--simulated). [default: anthropic/claude-opus-4-8] |
--simulator-judge-model TEXT | Model for the transcript judge that scores simulated samples (--simulated). [default: anthropic/claude-opus-4-6] |
--replay-sandbox | Serve each action's recorded tool output from the replayed trajectory instead of executing it (requires --trajectory-id). Nothing is provisioned or executed. Works with any replay-like policy — one whose executed tool calls stay a verbatim prefix of the recording; a call that diverges raises. |
--simulator-scratch-mode [filesystem|default_image|full|none] | Which real container the simulator investigates in (--simulated): filesystem (codebase copy, free), default_image (env's default image only), full (the whole compose stack), or none (no container — no-code-execution ablation). [default: filesystem] |
--epochs INTEGER | Number of epochs to run |
--max-samples INTEGER | Maximum samples to run in parallel |
--log-dir TEXT | Directory for eval logs |
--no-fail-on-error | Continue running if individual samples error (instead of failing the whole run) |
--scorer-timeout INTEGER | Timeout in seconds for scorer execution |
--diagnostics TEXT | Comma-separated diagnostic scorers to run (default: none) |
--score-main-on-attack | Also run the main-task scorer on attack-mode samples (adds main_task_success). Off by default; scores each attack sample twice, so worst-case scoring time doubles. |
--eval-arg TEXT | Argument passed to inspect_ai.eval() verbatim, as key=value with a JSON value (repeatable), e.g. --eval-arg model_roles='{"trusted": "openai/gpt-5"}' or --eval-arg time_limit=600 |
--harness [null|inspect_swe] | Which harness runs the untrusted model: null for the env as declared, inspect_swe for a real coding agent. White's to fix; defaults to null. |
-ha, --harness-arg TEXT | Harness args in key=value format (multiple allowed): agent, version. |
--grant TEXT | Env affordance tool to grant the red team, by its own name (multiple allowed): run_side_task_test, run_main_task_test. Absent from an env that does not declare it. |
--upload / --no-upload | Upload trajectories after the run. |
--run-name TEXT | Name for this run. |
--tag TEXT | Tag to attach to the run. |
--help | Show this message and exit. |