INDEPENDENT VOICE AI TESTING
Ship Voice AI without shipping avoidable failures.
VoiceNiche tests conversation behaviour, business policies, telephone conditions and external actions—then provides the recordings, evidence and severity-ranked findings.
We do not publish a downloadable sample report, because a redacted example from another customer's agent would tell you very little about yours. The interface below shows the structure of what you receive.
Production-readiness scorecard
suite: dental-inbound-v4
- Task completion92PASS
- Knowledge accuracy74WARNING
- Policy compliance88PASS
- Tool-call accuracy61FAIL
- Escalation accuracy69WARNING
- Latency90PASS
- Interruptibility94PASS
- Conversation quality96PASS
Overall readiness: not ready
Two critical failures override the average. An agent that sounds excellent but creates the wrong booking cannot pass.
Illustrative Voice QA report
THE PROBLEM
A good demo is not evidence of production readiness.
A voice agent can sound impressive for ninety seconds in a quiet room and still fail the first time reality touches it.
- The caller interrupts
- The caller changes their mind
- The line is noisy
- A tool returns unexpected data
- A calendar is unavailable
- A CRM write fails
- A transfer does not connect
- A prompt-injection attempt occurs
- A policy limit should block an action
- The agent is updated
SCENARIO MATRIX
Sixteen scenarios, defined before we dial.
Each carries an expected outcome, so the result is a pass or a fail rather than a matter of taste. Select any scenario to see what we look for.
Happy path
PASS- Expected action
- Booking created for requested slot
- Observed action
- Booking created
- Severity
- None
- Root-cause category
- —
Illustrative Voice QA report
WHAT MAKES THIS DIFFERENT
Simulated conversations are the easy part.
Anyone can generate synthetic calls. The work that matters is verifying what happened afterwards, attributing the fault correctly, and being able to run the whole thing again next week.
Vendor-neutral
We do not sell you the agent we are grading. Findings are not shaped by whose platform is at fault.
Tested through real deployment paths
Where possible we test over the telephone route your customers use, not a browser widget with a clean microphone.
Complete business journeys
The unit of measurement is the customer journey — enquiry through to completed action — not an isolated conversational turn.
External actions verified
We check whether the booking, the record and the message actually landed. A confident confirmation is not evidence.
Repeatable suites
Scenarios are defined once and re-run on demand, so you can compare versions instead of arguing about impressions.
Scenario-specific rubrics
An emergency triage call is not scored the same way as a booking enquiry. Each scenario carries its own pass criteria.
Evidence-backed results
Every finding ships with its recording and transcript. You can hear the failure, not just read our description of it.
Regression testing
Re-run after prompt, model or integration changes to catch what the update quietly broke.
Human review of serious findings
Critical and high-severity findings are checked by a person before they reach your inbox.
Faults attributed correctly
We separate agent, integration, telephony and provider failures so the fix lands with the right team.
EVALUATION MODEL
Seven scores, and one rule that overrides the average.
Scoring separately keeps a fluent agent from hiding a broken one. A single critical failure caps the overall result regardless of how well everything else performed.
01
Business outcome
Did the journey reach the result it was supposed to reach?
02
Knowledge accuracy
Was everything the agent stated true, approved and current?
03
Policy compliance
Did it stay inside your pricing, service-area and authority limits?
04
Tool and integration execution
Did the external action actually complete, and was it correct?
05
Escalation behaviour
Did it hand over when it should, and not when it should not?
06
Technical performance
Latency, interruptibility and stability on a real line.
07
Conversation quality
Was it clear, natural and appropriate to the caller's state?
Critical failures override the average
An agent that sounds excellent but creates the wrong booking must not receive a high production-readiness result. Conversation quality is the easiest axis to score well on and the least expensive one to get wrong.
THE REPORT
Every finding, with the evidence attached.
Select a row to see the expected action, what actually happened, where the fault sits and what we recommend.
Test suite
dental-inbound-v4
Agent version
prompt 2026.08.14 · model rev 3
Platform
Third-party agent (vendor-neutral test)
Telephone route
PSTN inbound, UK mobile and landline
Expected action
Fall back to a callback and notify the practice
Observed action
Confirmed a booking that was never written to the diary
Root-cause category
Integration
Recommendation
Gate the confirmation phrase on a successful write, not on the tool call being issued.
Retest status
Awaiting fix
Evidence
Call recording attachedFull transcript attached
Illustrative Voice QA interface
THE AUDIT
Production-Readiness Audit
Independent testing of an existing Voice AI agent. Scope is agreed before we start, and priced against it.
Quoted according to test scope
What an audit includes
- Defined customer journeys
- Real-world scenarios
- Recordings and transcripts
- Business-outcome scoring
- Integration verification
- Severity-ranked findings
- Recommendations
- One agreed retest
Scope varies by agent — confirmed in writing before work begins
FAQ
Voice QA questions
Let us find the failures before your customers do.
Tell us which journey matters most and what worries you about it. We will come back with a test scope, not a sales deck.
Request a Production-Readiness Audit