LIVE · 116 answersLast run: 68h agoJudge: gpt-4o-mini
Run eval

Per-Role Breakdown

Pass rates and metric means split by role template, sorted weakest first. Click a row to see which sessions contribute.

Role templateNPassCorrectnessCompletenessContextContextOpeningVoiceFaithfulness
unknown1160%0.510.430.500.430.420.440.47