How do you test a
Robot Phone Call?
VoiceIQ tests Retell voice agents before they go live โ benchmarked on EVA-Bench (Hugging Face), with LLM-as-judge scoring and a deploy gate. Retell Assure covers post-launch. Together that's the full quality loop.
Before launch + after launch
Pre-production ยท VoiceIQ EVA
EVA-Bench enterprise scenarios (airline, healthcare, ITSM), EVA-A/EVA-X scoring, Retell prompt import, and a deploy gate.
Post-production ยท Retell Assure
Retell's own QA layer monitors live calls. VoiceIQ is the missing pre-launch half of that story.
How it Works, Step-by-Step
The Mock Call
We load up two AI systems. One plays the role of your **AI Support Bot**. The other plays a **Mock Customer** (like an angry customer trying to cancel, or a confused buyer asking questions). They talk to each other just like a real phone call.
The AI Judge Grades
An objective **AI Referee** reads the entire typed conversation. It looks at key dimensions like *objection handling*, *empathy*, and whether the customer got what they wanted. It then awards a grade out of 100.
Verification (Calibration)
To make sure the AI Referee isn't just grading randomly, we compare its scores with **real human evaluations**. The closer the AI grades are to the human grades, the more we can trust it to test our product.
Try it in Action
In our live simulator tab, you can input your own prompt configurations, choose a custom scenario, click **Simulate**, and watch the AI systems execute the conversation loop live in front of you.