Skip to main content
VoiceNiche

INDEPENDENT VOICE AI TESTING

Ship Voice AI without shipping avoidable failures.

VoiceNiche tests conversation behaviour, business policies, telephone conditions and external actions—then provides the recordings, evidence and severity-ranked findings.

We do not publish a downloadable sample report, because a redacted example from another customer's agent would tell you very little about yours. The interface below shows the structure of what you receive.

Production-readiness scorecard

suite: dental-inbound-v4

  • Task completion92PASS
  • Knowledge accuracy74WARNING
  • Policy compliance88PASS
  • Tool-call accuracy61FAIL
  • Escalation accuracy69WARNING
  • Latency90PASS
  • Interruptibility94PASS
  • Conversation quality96PASS

Overall readiness: not ready

Two critical failures override the average. An agent that sounds excellent but creates the wrong booking cannot pass.

Illustrative Voice QA report

THE PROBLEM

A good demo is not evidence of production readiness.

A voice agent can sound impressive for ninety seconds in a quiet room and still fail the first time reality touches it.

  • The caller interrupts
  • The caller changes their mind
  • The line is noisy
  • A tool returns unexpected data
  • A calendar is unavailable
  • A CRM write fails
  • A transfer does not connect
  • A prompt-injection attempt occurs
  • A policy limit should block an action
  • The agent is updated

SCENARIO MATRIX

Sixteen scenarios, defined before we dial.

Each carries an expected outcome, so the result is a pass or a fail rather than a matter of taste. Select any scenario to see what we look for.

Happy path

PASS
Expected action
Booking created for requested slot
Observed action
Booking created
Severity
None
Root-cause category
—

Illustrative Voice QA report

WHAT MAKES THIS DIFFERENT

Simulated conversations are the easy part.

Anyone can generate synthetic calls. The work that matters is verifying what happened afterwards, attributing the fault correctly, and being able to run the whole thing again next week.

  • Vendor-neutral

    We do not sell you the agent we are grading. Findings are not shaped by whose platform is at fault.

  • Tested through real deployment paths

    Where possible we test over the telephone route your customers use, not a browser widget with a clean microphone.

  • Complete business journeys

    The unit of measurement is the customer journey — enquiry through to completed action — not an isolated conversational turn.

  • External actions verified

    We check whether the booking, the record and the message actually landed. A confident confirmation is not evidence.

  • Repeatable suites

    Scenarios are defined once and re-run on demand, so you can compare versions instead of arguing about impressions.

  • Scenario-specific rubrics

    An emergency triage call is not scored the same way as a booking enquiry. Each scenario carries its own pass criteria.

  • Evidence-backed results

    Every finding ships with its recording and transcript. You can hear the failure, not just read our description of it.

  • Regression testing

    Re-run after prompt, model or integration changes to catch what the update quietly broke.

  • Human review of serious findings

    Critical and high-severity findings are checked by a person before they reach your inbox.

  • Faults attributed correctly

    We separate agent, integration, telephony and provider failures so the fix lands with the right team.

EVALUATION MODEL

Seven scores, and one rule that overrides the average.

Scoring separately keeps a fluent agent from hiding a broken one. A single critical failure caps the overall result regardless of how well everything else performed.

  1. 01

    Business outcome

    Did the journey reach the result it was supposed to reach?

  2. 02

    Knowledge accuracy

    Was everything the agent stated true, approved and current?

  3. 03

    Policy compliance

    Did it stay inside your pricing, service-area and authority limits?

  4. 04

    Tool and integration execution

    Did the external action actually complete, and was it correct?

  5. 05

    Escalation behaviour

    Did it hand over when it should, and not when it should not?

  6. 06

    Technical performance

    Latency, interruptibility and stability on a real line.

  7. 07

    Conversation quality

    Was it clear, natural and appropriate to the caller's state?

Critical failures override the average

An agent that sounds excellent but creates the wrong booking must not receive a high production-readiness result. Conversation quality is the easiest axis to score well on and the least expensive one to get wrong.

THE REPORT

Every finding, with the evidence attached.

Select a row to see the expected action, what actually happened, where the fault sits and what we recommend.

Test suite

dental-inbound-v4

Agent version

prompt 2026.08.14 · model rev 3

Platform

Third-party agent (vendor-neutral test)

Telephone route

PSTN inbound, UK mobile and landline

  • Expected action

    Fall back to a callback and notify the practice

    Observed action

    Confirmed a booking that was never written to the diary

    Root-cause category

    Integration

    Recommendation

    Gate the confirmation phrase on a successful write, not on the tool call being issued.

    Retest status

    Awaiting fix

    Evidence

    Call recording attachedFull transcript attached

Illustrative Voice QA interface

THE AUDIT

Production-Readiness Audit

Independent testing of an existing Voice AI agent. Scope is agreed before we start, and priced against it.

Quoted according to test scope

What an audit includes

  • Defined customer journeys
  • Real-world scenarios
  • Recordings and transcripts
  • Business-outcome scoring
  • Integration verification
  • Severity-ranked findings
  • Recommendations
  • One agreed retest

Scope varies by agent — confirmed in writing before work begins

FAQ

Voice QA questions

Let us find the failures before your customers do.

Tell us which journey matters most and what worries you about it. We will come back with a test scope, not a sales deck.

Request a Production-Readiness Audit