AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis

arXiv · AI, language, vision and robotics · article · Aug 24, 2026 · UTC

Psychological research on emotion dynamics has established that human affect is a continuous, evolving process: emotions rise, decay, and transition within seconds. Current emotional text-to-speech (TTS) systems, however, condition on a single discrete label or static embedding per utterance, fundamentally misaligning with the temporal nature of affect. While recent LLM-based TTS systems may implicitly vary prosody through text understanding, such variation is neither explicitly controllable nor precise enough for targeted intra-utterance transitions. We address three challenges: (1) a multi-p

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-21T10:22:00.206Z. This is not the publication date.