For AI agents and LLMs: a machine-readable index is available at llms.txt. A plain-Markdown version of any documentation page is available by appending .md to its URL.
Skip to main content

Run Deep Functional Tests With Agent Assurance

/run selects runnable scenarios, shows a plan for review, asks for approval, invokes the live agent, and records evidence for every completed scenario. Check your credits with /plan first; the run plan is not a fixed spending limit.

Use a test target: Rook does not undo the target agent's actions. Refunds, messages, tickets, deployments, database updates, and file writes are real.

Point every run at a test or staging environment.

Preflight Checklist​

Launch rook from your agent workspace and enter the slash commands below inside the interactive TUI. /help run lists selectors and phases without starting a test; /run invokes the live target after its preflight and permission checks.

Before running a suite, confirm:

  1. The intended project is active: /project.
  2. The intended agent is active: /agent.
  3. The intended verified profile is active: /profile.
  4. The target URL or command points at test or staging. Review the hook scripts that implement the actual transport.
  5. Required fixtures and reset behavior are ready.
  6. Required MCP verification servers are enabled and approved: /mcp.
  7. Scenario runnability is understood: /scenarios list.
  8. The credit balance is sufficient: /plan.
  9. Concurrency is safe for the target's state and rate limits.

Sync the reviewed project with /sync before a normal timeline run. Use --test only when you intentionally want a local experiment that will not appear in the hosted history.

Run the Runnable Suite​

Verified
/run

Rook skips scenarios that cannot be attempted and groups the reasons.

A partially observable scenario still runs when it can establish useful evidence. Individual criteria that cannot be checked become Unable to Verify.

The default concurrency is 1 unless the profile or plan selects another value. An explicit --concurrency accepts 1–8 and overrides that choice.

Select Scenarios Precisely​

Run by ID:

Verified
/run --only SC-004,SC-011

Run by class:

Verified
/run --class adversarial

Run by category:

Verified
/run --category happy_path,prompt_injection

Run by tag:

Verified
/run --tag billing,refund

Selectors combine by narrowing. This command first keeps adversarial scenarios, then keeps those tagged refund:

Verified
/run --class adversarial --tag refund

You can also describe the desired subset after --:

Verified
/run --class adversarial -- the scenarios about refund approval

Natural-language selection uses a model to choose from the already filtered list, and Rook prints the matched IDs before the permission gate. In CI, prefer ID, class, category, and tag selectors because they are deterministic.

If no scenario matches, Rook prints the classes, categories, and tags that actually exist instead of running the full suite.

Choose Concurrency​

Verified
/run --concurrency 1
/run --concurrency 5

Use concurrency 1 when:

  • Scenarios mutate shared fixtures.
  • A reset must run between every scenario.
  • Filesystem changes need to be attributed to one scenario.
  • The target has a strict rate limit.
  • You are proving idempotency or sequence-sensitive behavior.

Use higher concurrency only when the target isolates sessions and fixtures. Concurrency changes parallelism, not the number of selected scenarios.

Review the Permission Gate​

First review the run plan. This actual quickstart capture selects one scenario, excludes the other, and waits for a decision:

Rook waiting for approval of a one-scenario run plan, with the excluded scenario and proceed, discard, and change choices

Use the arrow keys and Enter to choose proceed, discard, or change. Check the active project, agent, and profile in the footer. A selected plan does not prove that the endpoint is safe: inspect the profile and its scripts first. The plan is not a fixed credit quote.

Rook may also ask to approve individual tools or target invocations. Read the exact operation and target in each permission prompt.

The answers mean:

AnswerEffect
yesAllow this exact operation once.
alwaysRead its scope hint. For a model-chosen operation it lasts for this run without a disk grant; a persistent choice for a human-requested action can store a project grant.
neverStore a denial for this tool and target in this project.
noDecline without storing a decision.

Deny rules override allow rules, and more specific rules win. Permission state is stored globally under a per-project section, so a repository cannot grant itself permission.

Follow Run Progress​

After approval, each active scenario shows its current phase. This is an actual SC-002 execution against the public triage fixture, captured while Rook was judging the response and recorded hook evidence:

Rook's interactive run progress with SC-002 in judging, evidence reads, a completion counter, and active project and profile

0/1 means the scenario has not finished; it does not mean it failed. Wait for completion, then enter /report to inspect pass, fail, unverifiable, and unjudged counts. See results and evidence for the completed report from this demo.

Run Selected Phases​

/run --phases prepare,open,execute,close
/run --run <run-id> --phases collect,judge

The second command continues the same run after delayed evidence is ready. --resume instead creates a new run and carries compatible completed work forward. Rook owns judging; the other phases run your profile hooks.

See phases and hooks for prerequisites and state. The old --no-narrative option is not available in 0.1.3.

Request Root-Cause Analysis​

Verified
/run --rca

Rook clusters related failures first, then investigates each cause using the verdicts, scenario definition, feature, and read-only access to source. It writes remedies under:

Verified
.testmuai/rook/projects/<project-id>/agents/<agent-id>/runs/<run-id>/remedies/

RCA is off by default. It consumes additional credits, and its cost depends on the number of distinct failure clusters. A remedy is an evidence-grounded hypothesis, not a verified patch.

Interrupt and Resume Safely​

To interrupt a run: Press Esc during a TUI operation or Ctrl+C in a headless process. Rook aborts the in-flight HTTP request or command process and preserves completed requests, responses, and verdicts on disk.

The target may already have produced an external effect even when no response was recorded.

Authentication revocation, controller failure, and exhausted budget also halt work. Rook does not silently resume a run after authentication returns.

Test Common Agent Types Deeply​

Use scenarios that exercise both the user journey and externally visible effects.

The lists below include file-input journeys that teams commonly need. Native attachment delivery is not implemented in the current pre-alpha release, so run file-input cases in one of these ways:

  • Use a reviewed adapter that incorporates the file into the agent invocation.
  • Place a reachable test-file URL in the goal.

Otherwise, keep these cases documented but exclude them from release-gating runs.

Refund Agent​

  • Ask for a refund with no order ID.
  • Supply an unknown order ID.
  • Use a valid order belonging to another customer.
  • Request an amount above the approval threshold.
  • Repeat the same request to test idempotency.
  • Put prompt injection in a receipt supplied through the adapter or a test-file URL.
  • Make the billing verification service unavailable.
  • Verify that issue_refund was not called before identity checks.
  • Confirm the agent reports a pending, denied, or completed state accurately.

Travel Agent​

  • Give a destination but no dates or budget.
  • Change dates after accepting an itinerary.
  • Ask for inaccessible or sold-out inventory.
  • Mix currencies, time zones, and overnight flights.
  • Supply a passport image or preference document through the adapter or a test-file URL.
  • Ask for a PDF itinerary and verify the artifact separately from its contents.
  • Make one booking provider fail while alternatives remain.
  • Attempt to make the agent expose another traveler's PII.
  • Confirm that the agent does not claim a booking exists unless the booking system shows it.

Research or Document Agent​

  • Ask for a sourced answer and verify citations.
  • Supply conflicting PDFs through the adapter or test-file URLs.
  • Use an empty, encrypted, oversized, or malformed file.
  • Ask for text, JSON, image, and PDF outputs.
  • Return a link that expires or cannot be downloaded.
  • Test that unsupported evidence becomes Unable to Verify.
  • Repeat the same request to measure answer stability.

Coding or Repository Agent​

  • Provide a bug report with and without reproduction steps.
  • Test an unchanged repository and a dirty worktree.
  • Require exact file and line citations.
  • Refuse an unsafe destructive command.
  • Verify created files and test output.
  • Simulate a missing dependency or failing test runner.
  • Test a pull request checkout and a documentation-only repository.

Support or Workflow Agent​

  • Use valid, invalid, and ambiguous ticket IDs.
  • Ask a follow-up that depends on earlier context.
  • Simulate downstream ticket, CRM, or messaging failures.
  • Test forbidden promises, credits, deadlines, or competitor endorsements.
  • Verify whether tickets and replies were actually created.
  • Attempt prompt injection through ticket body, metadata, and adapter-delivered attachments.

MCP Tool Agent​

  • Introspect its declared tools.
  • Exercise read and write tools separately.
  • Change a project server definition after approval and confirm reapproval is required.
  • Disable a required server and confirm the scenario names the missing capability.
  • Attempt a write when only read behavior is expected.

Headless Runs and Hosted Results​

The same selectors and lifecycle controls are available in a shell:

rook run --only SC-001 --profile staging --concurrency 1 --name smoke --json

Select the project and agent before running; there is no --entity flag. Supply reviewed permissions when running unattended. See CI/CD for authentication, JSON, and completion checks.

Review the Run Locally or Online​

Local UIHosted Web UI
rook ui --localrook ui
Open agent → runs → run → scenario.Open project → agent → Runs → run → scenario.
Inspect on-disk results, including local --test runs and evidence awaiting upload.Inspect uploaded normal runs and share links with authorized teammates.

Neither command completes unfinished phases or retries the target. Keep the local serving process running. For the hosted UI, check the same account and ROOK_ENV; use rook runs sync for outstanding normal-run uploads. Do not expect a --test run to appear there.

See both UI walkthroughs and criterion evidence in each interface.

Local UI: Check the Completed Run​

Open the agent's Runs tab and select the execution. The saved CommerceCare demo has mixed passing, failed, and unverifiable results; completed does not mean every scenario passed. Read the narrative, open View plan, and click a scenario for criterion evidence. The profile name opens the run's saved configuration; the adjacent link opens the current profile. See the local walkthrough and rollout note for layout differences.

Local CommerceCare run showing mixed scenario outcomes, narrative, View plan, and profile metadata

Hosted Web UI: Review the Shared Execution​

Open project → agent → Runs → run. Check completion, agent version, invocation profile, and the scenario row. View plan explains selection and exclusions; opening the scenario shows the evidence for this execution.

Hosted completed run with one passed scenario, profile information, and a link to the scenario result

Terminal First Testing With Kane CLI

Natural language browser & mobile app tests right from terminal.

×
Schedule Your Personal Demo
Kane CLI terminal

Help and Support

Related Articles