SOURCE-LINKED INTELLIGENCE
GrainSpeech: Less Context, More Detail for Compact Speech Synthesis
Compact acoustic models face a challenging quality-capacity trade-off. We investigate two factors in this regime: encoder context and Mel-spectrogram supervision. A receptive-field-scaling study shows that expanding self-attention beyond 15 phonemes provides no consistent gains in pitch, energy, or duration prediction. Guided by this finding, we introduce a fixed-receptive-field convolutional encoder that reduces the respective prediction errors by 36.0%, 17.3%, and 3.4%. We further show that directly transferring image-domain gradient-variance supervision restores fine-scale variation but deg
Read original source ↗ Open in workspace
- recordType
- paper
- region
- Global
Evidence & attribution
- arXiv · AI, language, vision and robotics · 2026-09-16T15:59:20.000Z
- arXiv · Artificial Intelligence · 2026-09-16T15:59:20.000Z
First collected: 2026-09-19T20:26:32.566Z. This is not the publication date.