Jul 10, 2026, 4:00 AM

A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

TickrWire Editorial Desk·Jul 10, 2026, 4:00 AM·1 min read AI-assisted, human-reviewed

Reported by arXiv cs.CL: A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents. Analysis and context written by TickrWire.

30-second summary

arXiv:2607.07985v1 Announce Type: new Abstract: We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.5 Flash, 3.5 Flash, and 3.1 Pro. Our primary evidence base uses Gemini 2.5 Flash as the ground-truth model, validated against three calibrated human raters on 209 stereo sessions, scored on 8 production dimensions: 152 full-duplex conversations across 13 accent-and-condition strata, together with 57 adversarial defect-injected clips. The evidence

TickrWire
Full story

arXiv:2607.07985v1 Announce Type: new

Abstract: We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.5 Flash, 3.5 Flash, and 3.1 Pro. Our primary evidence base uses Gemini 2.5 Flash as the ground-truth model, validated against three calibrated human raters on 209 stereo sessions, scored on 8 production dimensions: 152 full-duplex conversations across 13 accent-and-condition strata, together with 57 adversarial defect-injected clips. The evidence

Sources · 1
More stories