For AI agents and LLMs: a machine-readable index is available at llms.txt. A plain-Markdown version of any documentation page is available by appending .md to its URL.
Skip to main content

Inbound Phone Agent Testing With TestMu AI


An inbound phone agent answers calls. To test one, the Agent Testing Platform places a real call to the agent's phone number, and a simulated caller drives the scenario. Typical use cases are IVR menus, inbound support, appointment scheduling, and billing.

Inbound testing runs in two modes: pre-evaluation with live simulated calls, and post-evaluation on recordings from real production calls.

Pre-Evaluation: Live Test Calls


In pre-evaluation, the platform simulates customers calling your voice agent, then evaluates the resulting conversations.

Phone Number Management. Register the numbers your agent answers on, with country code selection (20+ countries), a default number, masked display, and edit or delete.

Scenario Management. Generate up to 20 inbound scenarios with configurable personas, languages, and special instructions, or create them manually. Choose from available personas or create custom ones.

Voice Configuration (per scenario). Select a voice from the library with audio preview, enable one of 15 background-noise presets, set the response timing (0.5 to 5.0 seconds) and a maximum call duration (60 to 1800 seconds), and choose who speaks first.

Agent Profiles. Create reusable caller personas with name, phone number, voice, and background noise, stored in an organization-level library with an active or inactive toggle.

Test Suites. Group scenarios with per-scenario voice and phone configuration, associate test and agent profiles, and run the whole suite with one action.

Call Execution and Monitoring. Initiate live test calls, track status in real time, watch a live duration counter, and terminate a call in progress.

Post-Evaluation: Recording Analysis


In post-evaluation, you upload recordings from real production calls and score them with the same metrics, without placing new calls.

Voice Analytics. Upload production recordings (MP3, WAV) and transcripts, analyze them in parallel batches, select which metric categories or individual metrics to run, and bookmark, tag, search, and filter recordings.

Recording Playback. Play any call with play, pause, and duration controls, follow a speaker-identified transcript, see DTMF keypad inputs (0 to 9, star, pound) captured in the transcript, and download the audio and transcript.

Shared Across Both Modes


Go-Live Assessment. Get a Green (score at least 80), Yellow (65 to 79), or Red (below 65) verdict, with confidence based on call volume, dimension scores, scenario coverage, failure pattern analysis, validation-criteria compliance, and prioritized action items.

Metric Configuration. Select which metric categories or individual metrics to run per project.

Scheduled Runs. Automate runs with cron-based scheduling, IANA timezones, pause and resume, and run history.

Metrics


Phone agents are evaluated across 8 metric categories with 30+ individual metrics.

A. Conversation Flow and Interaction Dynamics

MetricUnitWhat it measures
Average LatencymsTime to respond after the user stops speaking
Words Per MinutewpmAgent speaking speed
AI Talk Ratio%Share of call time the agent is speaking
User Talk Ratio%Share of call time the user is speaking
AI Interrupting User%How often the agent interrupts the user
User Interrupting AI%How often the user interrupts the agent

B. Accuracy and Effectiveness

MetricUnitWhat it measures
First Call Resolution%Whether the issue was resolved in a single call
Intent Recognition Accuracy%How accurately the agent understood intent
Task Completion Success Rate%Share of assigned tasks completed
Instruction Following%Adherence to configured instructions
Response Consistency%Consistency of responses to similar inputs

C. User Experience and Satisfaction

MetricUnitWhat it measures
CSAT%Overall customer satisfaction score
CSAT ReasonTextExplanation for the satisfaction score
User SentimentTextDetected emotional sentiment from user speech
Early Termination%Share of calls not terminated prematurely

D. Business Operational Metrics

MetricUnitWhat it measures
Containment Rate%Issues resolved without human escalation
AI to Human Handoff Rate%Frequency of escalation to a human agent

E. Audio Voice Quality

MetricUnitWhat it measures
Average PitchHzVoice pitch (normal: 85 to 300 Hz)
Voice Quality Index0 to 5Composite voice quality score
Signal-to-Noise Ratio%Audio clarity versus background noise

F. Speech-to-Text Evaluation

MetricUnitWhat it measures
STT Accuracy%Transcription accuracy
STT VerdictPass/FailOverall transcription quality judgment
STT SummaryTextDetailed transcription quality notes
Mismatch ExamplesListInstances where transcription differed from speech

G. Validation Results

MetricUnitWhat it measures
Compliance%Compliance rate against custom validation criteria
Pass/Fail/Unable to VerifyCountPer-criterion validation breakdown

H. Detected Issue Tags (automated). Every recording is auto-scanned for: latency issues, hallucination in call flow, transcript issues, patchy audio, running in a loop, incorrect STT, interruption handling, number issues, background noise, no response, and blank or empty STT.

Threshold reference

MetricExcellentGoodPoor
Average Latencyat most 1000 msat most 2500 msover 2500 ms
Words Per Minuteat least 160 (fast)131 to 160under 110 (slow)
Voice Quality Indexat least 2.5 / 5n/aunder 2.5 / 5
Average Pitch85 to 300 Hzn/aunder 85 or over 300 Hz

Test across 3000+ combinations of browsers, real devices & OS.

Book Demo

Help and Support

Related Articles