Run Deep Functional Tests With Agent Assurance
/run selects runnable scenarios, shows a plan for review, asks for approval, invokes the live agent, and records evidence for every completed scenario. Check your credits with /plan first; the run plan is not a fixed spending limit.
Use a test target: Rook does not undo the target agent's actions. Refunds, messages, tickets, deployments, database updates, and file writes are real.
Point every run at a test or staging environment.
Preflight Checklist
Launch rook from your agent workspace and enter the slash commands below inside the interactive TUI. /help run lists selectors and phases without starting a test; /run invokes the live target after its preflight and permission checks.
Before running a suite, confirm:
- The intended project is active:
/project. - The intended agent is active:
/agent. - The intended verified profile is active:
/profile. - The target URL or command points at test or staging. Review the hook scripts that implement the actual transport.
- Required fixtures and reset behavior are ready.
- Required MCP verification servers are enabled and approved:
/mcp. - Scenario runnability is understood:
/scenarios list. - The credit balance is sufficient:
/plan. - Concurrency is safe for the target's state and rate limits.
Sync the reviewed project with /sync before a normal timeline run. Use --test only when you intentionally want a local experiment that will not appear in the hosted history.
Run the Runnable Suite
Verified/run
Rook skips scenarios that cannot be attempted and groups the reasons.
A partially observable scenario still runs when it can establish useful evidence. Individual criteria that cannot be checked become Unable to Verify.
The default concurrency is 1 unless the profile or plan selects another value. An explicit --concurrency accepts 1–8 and overrides that choice.
Select Scenarios Precisely
Run by ID:
Verified/run --only SC-004,SC-011
Run by class:
Verified/run --class adversarial
Run by category:
Verified/run --category happy_path,prompt_injection
Run by tag:
Verified/run --tag billing,refund
Selectors combine by narrowing. This command first keeps adversarial scenarios, then keeps those tagged refund:
/run --class adversarial --tag refund
You can also describe the desired subset after --:
/run --class adversarial -- the scenarios about refund approval
Natural-language selection uses a model to choose from the already filtered list, and Rook prints the matched IDs before the permission gate. In CI, prefer ID, class, category, and tag selectors because they are deterministic.
If no scenario matches, Rook prints the classes, categories, and tags that actually exist instead of running the full suite.
Choose Concurrency
Verified/run --concurrency 1
/run --concurrency 5
Use concurrency 1 when:
- Scenarios mutate shared fixtures.
- A reset must run between every scenario.
- Filesystem changes need to be attributed to one scenario.
- The target has a strict rate limit.
- You are proving idempotency or sequence-sensitive behavior.
Use higher concurrency only when the target isolates sessions and fixtures. Concurrency changes parallelism, not the number of selected scenarios.
Review the Permission Gate
First review the run plan. This actual quickstart capture selects one scenario, excludes the other, and waits for a decision:
Use the arrow keys and Enter to choose proceed, discard, or change. Check the active project, agent, and profile in the footer. A selected plan does not prove that the endpoint is safe: inspect the profile and its scripts first. The plan is not a fixed credit quote.
Rook may also ask to approve individual tools or target invocations. Read the exact operation and target in each permission prompt.
The answers mean:
| Answer | Effect |
|---|---|
yes | Allow this exact operation once. |
always | Read its scope hint. For a model-chosen operation it lasts for this run without a disk grant; a persistent choice for a human-requested action can store a project grant. |
never | Store a denial for this tool and target in this project. |
no | Decline without storing a decision. |
Deny rules override allow rules, and more specific rules win. Permission state is stored globally under a per-project section, so a repository cannot grant itself permission.
Follow Run Progress
After approval, each active scenario shows its current phase. This is an actual SC-002 execution against the public triage fixture, captured while Rook was judging the response and recorded hook evidence:
0/1 means the scenario has not finished; it does not mean it failed. Wait for completion, then enter /report to inspect pass, fail, unverifiable, and unjudged counts. See results and evidence for the completed report from this demo.
Run Selected Phases
/run --phases prepare,open,execute,close
/run --run <run-id> --phases collect,judge
The second command continues the same run after delayed evidence is ready. --resume instead creates a new run and carries compatible completed work forward. Rook owns judging; the other phases run your profile hooks.
See phases and hooks for prerequisites and state. The old --no-narrative option is not available in 0.1.3.
Request Root-Cause Analysis
Verified/run --rca
Rook clusters related failures first, then investigates each cause using the verdicts, scenario definition, feature, and read-only access to source. It writes remedies under:
Verified.testmuai/rook/projects/<project-id>/agents/<agent-id>/runs/<run-id>/remedies/
RCA is off by default. It consumes additional credits, and its cost depends on the number of distinct failure clusters. A remedy is an evidence-grounded hypothesis, not a verified patch.
Interrupt and Resume Safely
To interrupt a run: Press Esc during a TUI operation or Ctrl+C in a headless process. Rook aborts the in-flight HTTP request or command process and preserves completed requests, responses, and verdicts on disk.
The target may already have produced an external effect even when no response was recorded.
Authentication revocation, controller failure, and exhausted budget also halt work. Rook does not silently resume a run after authentication returns.
Test Common Agent Types Deeply
Use scenarios that exercise both the user journey and externally visible effects.
The lists below include file-input journeys that teams commonly need. Native attachment delivery is not implemented in the current pre-alpha release, so run file-input cases in one of these ways:
- Use a reviewed adapter that incorporates the file into the agent invocation.
- Place a reachable test-file URL in the goal.
Otherwise, keep these cases documented but exclude them from release-gating runs.
Refund Agent
- Ask for a refund with no order ID.
- Supply an unknown order ID.
- Use a valid order belonging to another customer.
- Request an amount above the approval threshold.
- Repeat the same request to test idempotency.
- Put prompt injection in a receipt supplied through the adapter or a test-file URL.
- Make the billing verification service unavailable.
- Verify that
issue_refundwas not called before identity checks. - Confirm the agent reports a pending, denied, or completed state accurately.
Travel Agent
- Give a destination but no dates or budget.
- Change dates after accepting an itinerary.
- Ask for inaccessible or sold-out inventory.
- Mix currencies, time zones, and overnight flights.
- Supply a passport image or preference document through the adapter or a test-file URL.
- Ask for a PDF itinerary and verify the artifact separately from its contents.
- Make one booking provider fail while alternatives remain.
- Attempt to make the agent expose another traveler's PII.
- Confirm that the agent does not claim a booking exists unless the booking system shows it.
Research or Document Agent
- Ask for a sourced answer and verify citations.
- Supply conflicting PDFs through the adapter or test-file URLs.
- Use an empty, encrypted, oversized, or malformed file.
- Ask for text, JSON, image, and PDF outputs.
- Return a link that expires or cannot be downloaded.
- Test that unsupported evidence becomes Unable to Verify.
- Repeat the same request to measure answer stability.
Coding or Repository Agent
- Provide a bug report with and without reproduction steps.
- Test an unchanged repository and a dirty worktree.
- Require exact file and line citations.
- Refuse an unsafe destructive command.
- Verify created files and test output.
- Simulate a missing dependency or failing test runner.
- Test a pull request checkout and a documentation-only repository.
Support or Workflow Agent
- Use valid, invalid, and ambiguous ticket IDs.
- Ask a follow-up that depends on earlier context.
- Simulate downstream ticket, CRM, or messaging failures.
- Test forbidden promises, credits, deadlines, or competitor endorsements.
- Verify whether tickets and replies were actually created.
- Attempt prompt injection through ticket body, metadata, and adapter-delivered attachments.
MCP Tool Agent
- Introspect its declared tools.
- Exercise read and write tools separately.
- Change a project server definition after approval and confirm reapproval is required.
- Disable a required server and confirm the scenario names the missing capability.
- Attempt a write when only read behavior is expected.
Headless Runs and Hosted Results
The same selectors and lifecycle controls are available in a shell:
rook run --only SC-001 --profile staging --concurrency 1 --name smoke --json
Select the project and agent before running; there is no --entity flag. Supply reviewed permissions when running unattended. See CI/CD for authentication, JSON, and completion checks.
Review the Run Locally or Online
| Local UI | Hosted Web UI |
|---|---|
rook ui --local | rook ui |
| Open agent → runs → run → scenario. | Open project → agent → Runs → run → scenario. |
Inspect on-disk results, including local --test runs and evidence awaiting upload. | Inspect uploaded normal runs and share links with authorized teammates. |
Neither command completes unfinished phases or retries the target. Keep the local serving process running. For the hosted UI, check the same account and ROOK_ENV; use rook runs sync for outstanding normal-run uploads. Do not expect a --test run to appear there.
See both UI walkthroughs and criterion evidence in each interface.
Local UI: Check the Completed Run
Open the agent's Runs tab and select the execution. The saved CommerceCare demo has mixed passing, failed, and unverifiable results; completed does not mean every scenario passed. Read the narrative, open View plan, and click a scenario for criterion evidence. The profile name opens the run's saved configuration; the adjacent link opens the current profile. See the local walkthrough and rollout note for layout differences.
Hosted Web UI: Review the Shared Execution
Open project → agent → Runs → run. Check completion, agent version, invocation profile, and the scenario row. View plan explains selection and exclusions; opening the scenario shows the evidence for this execution.

