SOURCE-LINKED INTELLIGENCE
Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis
Evaluating live streaming speech synthesis (TTS) requires assessing fine-grained, highly expressive prosody such as emotion, intonation, and energy which traditional MOS predictors fail to capture. While proprietary Large Language Models (LLMs) like Gemini can evaluate these aspects, they are too costly for massive inference and reinforcement learning feedback. To address this, we first introduce Live-ProsodyJudge (LPJ), a cost-effective pairwise evaluator distilled from Gemini into Qwen3-Omni. However, we identify a critical flaw in standard multi-dimensional evaluation: verdict coupling. The
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-17T12:20:48.000Z
- arXiv · Artificial Intelligence · 2026-09-17T12:20:48.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.