Abstract: Modeling speech variation is key to natural, expressive generation. Speaker embeddings are commonly used to condition personalized speech systems, but they are typically trained for speaker recognition, where intra-speaker variability is suppressed and inter-speaker separation is maximized. This objective leads to overly compact representations that may discard variations crucial for generation. We revisit this design choice and propose a sub-center modeling framework for speaker embeddings. Instead of a single prototype per speaker, we learn multiple sub-centers during discriminative training, allowing utterances to align with different prototypes. This strategy preserves structured intra-speaker variability while maintaining discriminability. In zero-shot voice conversion, our method improves intelligibility, increases pitch variability, achieves higher naturalness ratings, and retains strong speaker verification performance.

------------------> Sub-center Speaker Embeddings for Voice Conversion <--------


a) Proposed sub-center modeling applied to the traditional ECAPA-TDNN network. b) VC framework using the proposed embeddings

-----------------------------> Speech Samples <---------------------------

Experimental Setup:

The samples are from speakers that are unseen during the VC training. For each conversion, a random reference utterance from target speaker(~3s) is used
Methods:
  • VC with baseline ECAPA-TDNN [1,2]
  • VC with Sub-center ECAPA-TDNN, C=10, T=0.1 (least intra-class variance)
  • Proposed Method: VC with Sub-center ECAPA-TDNN, C=20 (most intra-class variance)



Zero-shot Voice Conversion


Ground-Truth VC with Baseline ECAPA-TDNN[1,2] VC with Sub-center ECAPA-TDNN, C=10, T=0.1 (least intra-class variance) Proposed Method: VC with Sub-center ECAPA-TDNN, C=20 (most intra-class variance)

Female-to-Male

Source: p229 Target: p345
Source: p308 Target: p260

Female-to-Female

Source: p329 Target: p305
Source: p265 Target: s5

Male-to-Female

Source: p260 Target: p310
Source: 246 Target: 305

Male-to-Male

Source: p298 Target: p345
Source: p246 Target: p260
[1] Brecht Desplanques, Jenthe Thienpondt, & Kris Demuynck (2020). ECAPA-TDNN: Emphasized Channel Attention, Propagation and Aggregation in TDNN Based Speaker Verification. In Proc. Interspeech 2020 (pp. 3830–3834).
[2] Adam Polyak, Yossi Adi, Jade Copet, Eugene Kharitonov, Kushal Lakhotia, Wei-Ning Hsu, Abdelrahman Mohamed, & Emmanuel Dupoux (2021). Speech Resynthesis from Discrete Disentangled Self-Supervised Representations. In Proc. Interspeech 2021 (pp. 3615–3619).