For AI agents and LLMs: a machine-readable index is available at llms.txt. A plain-Markdown version of any documentation page is available by appending .md to its URL.
Skip to main content

How to Get Started With Agent Assurance

Run one small test before connecting a business-critical agent. This walkthrough uses Rook's public support-triage sample: a local HTTP service with in-memory tickets, no model key, and no external customer actions.

This is an interactive TUI walkthrough, not a CI pipeline. Use your normal shell only to install Rook and start the sample. Then launch rook without a subcommand and enter the slash commands below inside its prompt, one step at a time.

Install → launch Rook → sign in → select a project
→ explore → select the agent → create and verify a profile
→ generate and review scenarios → sync → approve one run → inspect evidence

The interactive workflow and terminal screenshots were tested with Rook 0.1.5 on September 28, 2026. Discovery, profile generation, scenario generation, and judging use TestMu AI credits; even a small suite can involve several model calls. Review the proposed work and credit balance before approving it. The browser examples below are separately identified saved runs.

Already have a live agent? Follow the same sequence with your own source or requirements and invocation profile.

Install Rook​

Choose one public install method. On Windows, first follow native PowerShell setup, then use the PowerShell alternatives below. Use rook.cmd in PowerShell and the same slash commands inside Rook's terminal.

# Homebrew
brew install lambdatest/rook/rook
# Shell installer
curl -fsSL https://raw.githubusercontent.com/LambdaTest/rook/main/install.sh | bash
# npm (use Node.js 22 or newer)
npm install -g @testmuai/rook

PowerShell on Windows x64:

npm.cmd install -g @testmuai/rook@0.1.5

Then check your installation:

rook --version

See Install Rook for PATH, upgrade, and checksum help, including migration from the old Homebrew tap. The packaged CLI includes its runtime. The sample below also needs a separate Node.js installation and Git available in the same terminal.

Use the Public Service​

Public packages default to ROOK_ENV=prod, the service behind Rook Projects. The launch commands in step 2 set it explicitly. Sign in inside the TUI with /login; there is no separate headless login step. This environment selects Rook's service, not your target agent's endpoint. Keep the target on a disposable or non-production environment while testing.

Existing credentials

If LT_USERNAME and LT_ACCESS_KEY are exported, they take precedence over stored browser login. Use credentials for the selected environment, or unset both in this terminal before using browser login. In PowerShell, use Remove-Item Env:LT_USERNAME, Env:LT_ACCESS_KEY -ErrorAction SilentlyContinue to remove them from this session only. Never paste credentials into documentation, prompts, or screenshots.

Test Your First Agent​

1. Start the public sample​

The triage-service sample is the target being tested—not Rook's backend.

In a separate terminal:

git clone --depth 1 --filter=blob:none --sparse https://github.com/LambdaTest/rook.git rook-samples
cd rook-samples
git sparse-checkout set samples/triage-service
cd samples/triage-service

Start the sample in that terminal. On macOS/Linux:

PORT=19110 node src/server.mjs

On Windows PowerShell:

$env:PORT = '19110'
node src/server.mjs

Leave it running. It keeps fixture changes in memory; restarting it resets them. From another terminal, verify the target:

curl -fsS http://127.0.0.1:19110/healthz
curl -fsS http://127.0.0.1:19110/v1/triage \
-H 'content-type: application/json' \
-d '{"input":"please look at T-1043"}'

In PowerShell, use Invoke-RestMethod instead of the Bash cURL example:

Invoke-RestMethod -Uri 'http://127.0.0.1:19110/healthz'
$rookProbe = @{ input = 'please look at T-1043' } | ConvertTo-Json
Invoke-RestMethod -Uri 'http://127.0.0.1:19110/v1/triage' -Method Post -ContentType 'application/json' -Body $rookProbe

The response should say T-1043 triaged as S1 and assigned to platform. and include the recorded tool steps.

2. Open the TUI, Sign In, and Explore​

Open the sample folder in your Rook terminal:

cd rook-samples/samples/triage-service
export ROOK_ENV=prod
rook

Or, in PowerShell:

Set-Location 'rook-samples/samples/triage-service'
$env:ROOK_ENV = 'prod'
rook.cmd

The Rook banner, command prompt, and footer confirm that you are in interactive mode. If you are already signed in, Rook can open the project chooser automatically:

Fresh Rook interactive session showing the Explore, Generate, Run, and Report journey and project chooser before discovery

If not signed in, enter /login, complete the browser flow, and return to this terminal. Use /doctor to check readiness and /plan to check credits. At a chooser, use arrow keys and Enter; Esc goes back. At the prompt, Tab completes commands. The remaining commands belong inside this TUI, not in your shell.

/project

Select an existing test project, or create one:

/project create "Rook quickstart"

Run exploration only after the intended test project is selected:

/explore .

Rook reads the sample and shows discovery progress. When it proposes an agent, review the name, source path, and purpose before choosing yes to register it. This is an actual exploration of the public sample:

Rook exploration finding the support triage agent and asking whether to register it, with yes and no choices and a live progress footer

Wait for exploration to finish, then open the agent picker:

/agent

Select the discovered triage agent. Generated IDs can differ: use the actual ID/name shown by your session. Confirm that the discovered features describe the ticket service, not unrelated files.

3. Generate and verify the profile from a prompt​

Enter /profile add local-triage in the TUI. When Rook asks how to reach the agent, paste this description:

Reach the running service at http://127.0.0.1:19110/v1/triage.
For execute, POST JSON {"input": <the goal read from stdin>}.
The response's output field is the agent's answer: return it as agent_reply.
Map each response step's tool and args to calls[].name and calls[].arguments.
Use "please look at T-1043" as the harmless verification goal.
No credentials are needed. Do not start another server or install dependencies.
This fixture is single-turn. Its echoed session_id is not conversational state.
Do not report the fixture's zero usage values as measured model usage.
Use concurrency 1.

Rook generates the hook and verifies it against the running sample. Approve only the intended local HTTP call and script work; do not approve server startup, dependency installation, or unrelated commands. Wait for the result before continuing.

Actual TUI profile-generation conversation with the triage HTTP request description and an approval prompt for the generated hook script

This checkpoint shows authoring in progress. Read the exact requested file or command, then approve once with yes if it matches the sample setup; the prompt is not evidence that verification has passed.

To use a saved description instead, put the same text in triage-profile.txt in the sample folder and use /profile add local-triage --from triage-profile.txt. These are alternatives—do not create the same profile twice.

Inspect and test the generated result inside the TUI:

/profile show local-triage
/profile test local-triage --goal "please look at T-1043"
/profile use local-triage

Review the generated profiles/local-triage.yaml and scripts/ files below the active agent directory. A successful probe should return the real answer and four tool calls. Stop here if the probe fails; generation and testing will not repair a broken connection automatically.

Rook completing profile authoring after a successful triage probe, reporting four mapped tool calls and selecting local-triage as the active profile

This is the successful connection checkpoint: the hook answered, the generated profile was saved, and the footer now includes local-triage. Profile verification is not a scenario verdict.

The sample needs only an execute hook. Use additional lifecycle hooks for login, session setup, teardown, or delayed trace collection.

4. Generate a small suite and review it​

/generate --total 2 --class functional --category happy_path -- Create two single-turn cases for existing tickets only: T-1043 must be S1/platform and T-1041 must be S2/billing. Check the answer and recorded calls. Do not require external verification or repeated samples.

Before writing scenarios, Rook shows the plan and offers proceed, discard, or change. Check both the included cases and the features it intentionally leaves uncovered:

Rook's actual generation plan for two triage scenarios, with included and excluded features and proceed, discard, or change choices

Choose proceed only if the scope matches the two requested tickets. Wait for writing to finish, then inspect the saved suite:

/scenarios list
Rook listing the two generated triage scenarios, their functional happy-path classification, four criteria each, and runnability against local-triage

The count is a generation target, not a guarantee. In this capture, two scenarios were written and both are runnable against local-triage; neither has passed yet. Review the resulting files before running them. Check that every criterion can be evaluated from the answer or the calls your hook actually returns. For example, a JSON-path check against $.steps cannot inspect that field if your hook returned only an answer string.

Choose the scenario for T-1043. Do not assume it will always be SC-002.

5. Sync, then run one scenario​

/sync

Wait for the agent to be recorded upstream before starting the run:

Rook synchronizing the triage agent and confirming one agent recorded upstream
/run --only <your-T-1043-scenario-id> --profile local-triage --concurrency 1 --name first-triage-run

Replace the placeholder with the generated scenario ID. Review the selected scenario and active profile before proceeding. This plan selects only SC-002 and explicitly skips SC-001:

Interactive run plan selecting one triage scenario, skipping the other, and waiting for proceed, discard, or change

Choose proceed to execute this reviewed scope, discard to run nothing, or change to revise it. Read any additional permission prompts before approving. The plan is not a fixed credit quote; check /plan before starting.

While Rook runs, the scenario row shows its current phase. This capture is in judging, with 0/1 completed; it is not a final verdict:

Live Rook test run showing SC-002 in the judging phase with evidence reads and zero of one scenarios completed

Wait for completion before opening the report. Press Esc only if you intend to interrupt; stopping Rook does not undo actions already performed by the target.

A normal run needs a synchronized agent version. --test is for an intentionally local, unsynchronized experiment; it does not add a run to the shared timeline.

6. Open the results​

First read the report, then choose either UI:

/report
Completed interactive report for first-triage-run showing one executed scenario, one pass, zero failures, and a warning about narrow feature coverage

The report names the saved run and its evidence directory. Check the counts, not just the percentage: this run executed one scenario and left four discovered feature areas untested. Reading /report does not rerun the target; do not add --rca unless you intend to request additional paid analysis.

Review on this machineReview with your team
Run /ui --local.Run /ui.
Open triage-service → runs → your run → scenario.Open project → triage agent → Runs → first-triage-run → scenario in the Web UI.
In the redesigned viewer, read Acceptance criteria, then choose Request, Response, Verdict, or Artefacts from Evidence. For older CLI layouts, see the note below.Read the criteria, then open Request, Response, Verdict, or Artefacts from Evidence.
Works with on-disk evidence, including --test runs; keep the TUI open while reviewing.Requires browser sign-in and uploaded results; teammates need project access.

Both routes inspect the recorded evidence without running the agent again. The local and hosted UI walkthrough shows the different screens and explains missing results.

In the verified smoke test, the selected scenario passed with four observed tool calls and no unverifiable criteria. That proves this one fixture path worked—not that the whole agent is reliable. Review the four other discovered features before expanding the suite.

Local UI: Review a Result​

The redesigned local result has Acceptance criteria filters and an Evidence drawer for request, response, verdict, and artifacts. This screenshot uses a separate saved CommerceCare demo (SC-006), not the triage execution above. It illustrates one failed requirement and two unverifiable criteria; your scenario IDs and outcomes will differ. Public CLI 0.1.5 still uses the earlier scrolling layout—see the rollout and navigation note.

Redesigned local result for the separate CommerceCare SC-006 demo, with failed and unverifiable criteria

Hosted Web UI: The Uploaded Result​

The browser capture below shows the earlier September 11 triage smoke test, not the September 28 TUI run above. Open Evidence → Verdict to read a saved evaluation in the drawer, then close it to return to the criterion cards. Use Expand all to read passing criteria, which start collapsed. Opening either UI reviews existing evidence; it does not execute the agent again.

Hosted quickstart result with verdict.yaml open in the evidence drawer

7. Stop the sample when finished​

Enter /exit to leave Rook. Stop the sample server with Ctrl+C in its terminal. Rook results stay below:

.testmuai/rook/projects/<project-id>/agents/<agent-id>/runs/<run-id>/

Continue After Your First Test​

To repeat this workflow through your coding assistant, choose a client-specific Rook skill guide. To automate the reviewed suite, use GitHub Actions, Jenkins, or Argo CD.

Use /status inside the TUI to check the selected project, active agent, and local/upstream state. Bare /project, /agent, and /profile open pickers; select with the arrow keys and Enter. Shell equivalents are covered in the command reference; you do not need to leave the TUI to continue this walkthrough.

Ask in Plain Language​

Use /ask in the TUI when you know the outcome but not the command:

Verified
/ask generate adversarial tests for refund-policy bypasses

Rook resolves the request to the appropriate operation. Any operation that spends credits or needs permission still shows its plan and asks first.

Local Changes and Sync​

Exploration, generation, profile authoring, and curation write plain files under .testmuai/rook/. They do not silently publish workspace state.

/sync records the current project tree upstream. Profile files contain environment-variable references, never their secret values. Run results are saved locally as they happen and can be reconciled upstream after connectivity returns.

When to Repeat a Step​

ChangeRepeat
Agent source, prompt, tools, or policy changedexplore, then regenerate affected scenarios
Test intent changed without an implementation changegenerate with an instruction, then curate
Endpoint, authentication, or response shape changedprofile test, then profile fix if needed
Only the deployed target changedrun against the intended profile
Evidence arrives asynchronouslyContinue the same run with --run <id> --phases collect,judge
Local project metadata needs publishingsync

Connect Your Own Agent Next​

Use staging credentials and disposable data. Tell the profile author the real request, answer field, authentication, session semantics, and evidence sources. Do not claim multi-turn state, observable calls, or measured usage unless the target actually supplies them.

Prompt-based profiles · Phases and hooks · Review results · CI/CD

Terminal First Testing With Kane CLI

Natural language browser & mobile app tests right from terminal.

×
Schedule Your Personal Demo
Kane CLI terminal

Help and Support

Related Articles