Eval Overview
116 answers across 18 sessions · judge: gpt-4o-mini · run: nightly-110
Headline metrics
Quality
Pass rate
0.0%
below 70% threshold
Faithfulness
46.8%
hallucinations detected
Voice
Opening
41.9%
meta-openings detected
Authenticity
43.6%
too formal / AI-sounding
Coverage
Correctness
51.4%
technical errors found
Completeness
42.8%
missing sub-questions
Sessions
weakest scoring sessions · click to drill in
Failure modes
why answers failed — most common first
Meta-opening ("The question is about...")99 · 85.3%
Incomplete answer (missed sub-question)92 · 79.3%
Hallucinated company / project not in resume90 · 77.6%
Voice: too formal, no contractions89 · 76.7%
Off-topic tangent (low context precision)83 · 71.6%
Per-Metric Detail
mean score across all 116 answers · worst: Opening (0.42)
Correctness
0.51
FAIL
Completeness
0.43
FAIL
Context Recall
0.50
FAIL
Context Precision
0.43
FAIL
Opening
0.42
FAIL
Voice Authenticity
0.44
FAIL
Faithfulness
0.47
FAIL