multimodal
fact
bullish
XTTSv2's voice cloning capabilities can preserve prosodic structure independently of speaker identity, enabling it to be repurposed for speaker anonymization without retraining
Our key insight is that XTTSv2's voice cloning capabilities preserve prosodic structure independently of speaker identity, enabling voice conversion by conditioning on a pseudo-speaker.
Computation and Language30 Aug 2026