ct run find-questions-iterative

Iteratively discover answer-question tasks across all environments.

Usage

ct run find-questions-iterative [OPTIONS]

Options

OptionDescription
--ec2Run evaluation on EC2 instead of locally
--ec2-arg TEXTConfig for --ec2 in key=value format (multiple allowed): instance_type (default 'c7i.4xlarge'), region (default 'us-east-1'), keep_instance (default False), max_workers (default 64), concurrency_per_worker (default 32), estimate (default False), auto_confirm (default False), new_fleet (default False), spot (default False), worker_setup (default None), ami (default None).
--startup-retries INTEGERNumber of times to retry Docker sandbox startup on transient failures [default: 7]
--no-internetBlock outbound internet and internet simulator access.
--no-intranetBlock access to compose sibling services (e.g. Redis).
--network-controlGrant NET_ADMIN to the agent container without applying restrictions at startup. Protocols that manage iptables mid-trajectory (e.g. privilege-monitor) need this.
--sandbox-isolationWrap every agent command in a bubblewrap sandbox: read-only root, writable workspace only. A static bwrap is fetched automatically; set CT_BWRAP_HOST_BINARY to override with your own.
--extra-src DIRECTORYDirectory of experiment code added to PYTHONPATH locally, so external --policy my.module:fn and --sandbox my.module:type refs resolve without committing into the repo; fleet mode (--ec2) also ships it to workers. Repeatable; each dir must have a distinct basename.
--sandbox TEXTRun samples in an Inspect sandbox provided outside this repo, as a registered type (my-sandbox) or module:type to import the module that registers it. Replaces the default Docker sandbox; with --ec2 each worker runs it for its own jobs.
--sandbox-arg TEXTConfig for --sandbox in key=value format (multiple allowed), passed through to the sandbox as JSON.
--simulatedRun against an LLM-simulated sandbox instead of provisioning the real environment. Every tool output is fabricated by a simulator agent.
--simulator-model TEXTModel for the simulator agent (--simulated). [default: anthropic/claude-opus-4-8]
--simulator-judge-model TEXTModel for the transcript judge that scores simulated samples (--simulated). [default: anthropic/claude-opus-4-6]
--replay-sandboxServe each action's recorded tool output from the replayed trajectory instead of executing it (requires --trajectory-id). Nothing is provisioned or executed. Works with any replay-like policy — one whose executed tool calls stay a verbatim prefix of the recording; a call that diverges raises.
--simulator-scratch-mode [filesystem|default_image|full|none]Which real container the simulator investigates in (--simulated): filesystem (codebase copy, free), default_image (env's default image only), full (the whole compose stack), or none (no container — no-code-execution ablation). [default: filesystem]
--epochs INTEGERNumber of epochs to run
--max-samples INTEGERMaximum samples to run in parallel
--log-dir TEXTDirectory for eval logs
--no-fail-on-errorContinue running if individual samples error (instead of failing the whole run)
--scorer-timeout INTEGERTimeout in seconds for scorer execution
--diagnostics TEXTComma-separated diagnostic scorers to run (default: none)
--score-main-on-attackAlso run the main-task scorer on attack-mode samples (adds main_task_success). Off by default; scores each attack sample twice, so worst-case scoring time doubles.
--eval-arg TEXTArgument passed to inspect_ai.eval() verbatim, as key=value with a JSON value (repeatable), e.g. --eval-arg model_roles='{"trusted": "openai/gpt-5"}' or --eval-arg time_limit=600
--no-uploadDisable uploading eval results
-p, --policy-arg TEXTPolicy args in key=value format
--target INTEGERTarget number of answer questions per environment (default: 5)
--helpShow this message and exit.