3
AC
Quality/Evaluation

Automated Evaluation Engine

Score every production call with deterministic, rule, semantic, LLM-as-judge, audio, human and custom evaluators. Results include evidence — not absolute truth.

Avg accuracy

82

Safety score

98

Task success

76

Built-in evaluators

31

Evaluation sampling

How calls enter the scoring pipeline

DeterministicRule / Regex / KeywordSemantic similarityLLM-as-judgeMulti-model panelAudio qualityClassifierHuman reviewExternal APICustom codeSQL

Built-in evaluators (sample)

Call completed successfully

Customer goal achieved without failure

Outcome

Correct greeting used

Required opening present

Script

Identity verified

Auth steps completed when required

Security

Customer intent identified

Intent correctly classified

NLU

Required information collected

Mandatory fields gathered

Workflow

Correct workflow followed

State machine path valid

Workflow

Correct tool called

Expected tools invoked

Tools

Tool output used correctly

Downstream use of tool data correct

Tools

Backend transaction completed

Side-effect committed

Outcome

Customer request resolved

Request closed satisfactorily

Outcome

Human handoff performed correctly

Escalation path correct

Transfer

Transfer context preserved

Context passed to human

Transfer

No hallucination

Claims grounded in tools/KB

Safety

No unsupported claim

No fabricated policy/pricing

Safety

No prohibited advice

No banned domains

Compliance

No sensitive-data disclosure

No unexpected PII/PAN

Privacy
View all 31 + builder →

Evaluation reliability

Judge comparison

OpenAI vs Claude panel

Calibration

Weekly golden set

Precision

0.91

Recall

0.87

F1

0.89

Inter-rater κ

0.78

False-positive review

Queue open

Evaluator drift

Stable

LLM evaluator output is not absolute fact — each result shows confidence and evidence for human review.

Recent evaluation runs

Score · pass/fail · confidence · reason · evidence · cost

CallAgentAccuracySafetyComplianceTaskOverallVerdictConfidenceEvidence
call_8f2a91Support Concierge489690041fail0.80task_incomplete, excessive_latency
call_3bc771Collections Agent951009810093pass0.93spans + transcript
call_9912deBanking IVR Assist6299884055partial0.84escalation, asr_confidence
call_55aa10Appointment Scheduler971009510096pass0.94spans + transcript
call_aa4411Outbound Survey Bot80100100060review0.85abandoned