Automated Evaluation Engine
Score every production call with deterministic, rule, semantic, LLM-as-judge, audio, human and custom evaluators. Results include evidence — not absolute truth.
Avg accuracy
82
Safety score
98
Task success
76
Built-in evaluators
31
Evaluation sampling
How calls enter the scoring pipeline
Built-in evaluators (sample)
Call completed successfully
Customer goal achieved without failure
Correct greeting used
Required opening present
Identity verified
Auth steps completed when required
Customer intent identified
Intent correctly classified
Required information collected
Mandatory fields gathered
Correct workflow followed
State machine path valid
Correct tool called
Expected tools invoked
Tool output used correctly
Downstream use of tool data correct
Backend transaction completed
Side-effect committed
Customer request resolved
Request closed satisfactorily
Human handoff performed correctly
Escalation path correct
Transfer context preserved
Context passed to human
No hallucination
Claims grounded in tools/KB
No unsupported claim
No fabricated policy/pricing
No prohibited advice
No banned domains
No sensitive-data disclosure
No unexpected PII/PAN
Evaluation reliability
Judge comparison
OpenAI vs Claude panel
Calibration
Weekly golden set
Precision
0.91
Recall
0.87
F1
0.89
Inter-rater κ
0.78
False-positive review
Queue open
Evaluator drift
Stable
LLM evaluator output is not absolute fact — each result shows confidence and evidence for human review.
Recent evaluation runs
Score · pass/fail · confidence · reason · evidence · cost
| Call | Agent | Accuracy | Safety | Compliance | Task | Overall | Verdict | Confidence | Evidence |
|---|---|---|---|---|---|---|---|---|---|
| call_8f2a91 | Support Concierge | 48 | 96 | 90 | 0 | 41 | fail | 0.80 | task_incomplete, excessive_latency |
| call_3bc771 | Collections Agent | 95 | 100 | 98 | 100 | 93 | pass | 0.93 | spans + transcript |
| call_9912de | Banking IVR Assist | 62 | 99 | 88 | 40 | 55 | partial | 0.84 | escalation, asr_confidence |
| call_55aa10 | Appointment Scheduler | 97 | 100 | 95 | 100 | 96 | pass | 0.94 | spans + transcript |
| call_aa4411 | Outbound Survey Bot | 80 | 100 | 100 | 0 | 60 | review | 0.85 | abandoned |