AIIC AI Intelligence Centre

SOURCE-LINKED INTELLIGENCE

Voice or Stereotype? Disentangling Acoustic and Content-Based Gender in Speech-to-Speech Models

arXiv · AI, language, vision and robotics · article · Sep 8, 2026 · UTC

Speech-to-speech (S2S) models now run inside dubbing, translation, and voice agents. Unlike text models, they hear the speaker's voice, which carries the speaker's gender. A faithful system should treat a speaker as who they sound like, not as whoever usually says what they said. Testing this is harder than it looks, since most S2S models answer in a single, fixed output voice, hard-coded so it cannot drift toward a stereotype. Checking the output voice comes back clean even when the model is biased. We therefore ask two questions. When a model re-speaks the input, does the stereotype in the w

Read original source ↗ Open in workspace

recordType
paper
region
Global

Evidence & attribution

First collected: 2026-09-20T20:02:11.508Z. This is not the publication date.