Abstract: Modeling speech variation is key to natural, expressive generation. Speaker embeddings are commonly used to condition personalized speech systems, but they are typically trained for speaker recognition, where intra-speaker variability is suppressed and inter-speaker separation is maximized. This objective leads to overly compact representations that may discard variations crucial for generation. We revisit this design choice and propose a sub-center modeling framework for speaker embeddings. Instead of a single prototype per speaker, we learn multiple sub-centers during discriminative training, allowing utterances to align with different prototypes. This strategy preserves structured intra-speaker variability while maintaining discriminability. In zero-shot voice conversion, our method improves intelligibility, increases pitch variability, achieves higher naturalness ratings, and retains strong speaker verification performance.
------------------> Sub-center Speaker Embeddings for Voice Conversion <--------
-----------------------------> Speech Samples <---------------------------
The samples are from speakers that are unseen during the VC training. For each conversion, a random reference utterance from target speaker(~3s) is used
Methods:
- VC with baseline ECAPA-TDNN [1,2]
- VC with Sub-center ECAPA-TDNN, C=10, T=0.1 (least intra-class variance)
- Proposed Method: VC with Sub-center ECAPA-TDNN, C=20 (most intra-class variance)
Zero-shot Voice Conversion
| Ground-Truth | VC with Baseline ECAPA-TDNN[1,2] | VC with Sub-center ECAPA-TDNN, C=10, T=0.1 (least intra-class variance) | Proposed Method: VC with Sub-center ECAPA-TDNN, C=20 (most intra-class variance) | Female-to-Male |
|
|---|---|---|---|---|---|
| Source: p229 | Target: p345 | ||||
| Source: p308 | Target: p260 | Female-to-Female |
|||
| Source: p329 | Target: p305 | ||||
| Source: p265 | Target: s5 | Male-to-Female |
|||
| Source: p260 | Target: p310 | ||||
| Source: 246 | Target: 305 | Male-to-Male |
|||
| Source: p298 | Target: p345 | ||||
| Source: p246 | Target: p260 | ||||
[2] Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, & Emmanuel Dupoux (2021). Speech Resynthesis from Discrete Disentangled Self-Supervised Representations. In Proc. Interspeech 2021 (pp. 3615–3619).