Training speakers — ESD dataset

The two ESD speakers below were used during training exclusively with their neutral utterances (~350 samples each). All emotional samples from both speakers were held out entirely and used only for evaluation in the cross-speaker emotional style transfer setup — the model must generalise to emotions it has never seen the target speaker express.

Speaker 0011

Reference neutral sample — training distribution

Speaker 0015

Reference neutral sample — training distribution

GT (ground truth) E3 TTS VECL Proposed Proposed (self) Proposed (VC)

😠 Angry

Speaker 0011

Emotion reference
angry_0011_000433
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)
angry_0011_000478
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)

Speaker 0015

angry_0015_000384
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)
angry_0015_000619
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)

😊 Happy

Speaker 0011

Emotion reference
happy_0011_000716
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)
happy_0011_000939
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)

Speaker 0015

happy_0015_000769
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)
happy_0015_000879
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)

😢 Sad

Speaker 0011

Emotion reference
sad_0011_001163
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)
sad_0011_001376
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)

Speaker 0015

sad_0015_001245
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)
sad_0015_001258
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)

😲 Surprise

Speaker 0011

Emotion reference
surprise_0011_001576
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)
surprise_0011_001661
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)

Speaker 0015

surprise_0015_001571
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)
surprise_0015_001682
GT
E3 TTS
VECL
Proposed
Proposed (self)
Proposed (VC)

🌐 Cross-corpus generalization

The cross-corpus speakers (LJ Speech, p226, and p231 from VCTK) were trained with a similar data budget to the neutral-only ESD speakers — approximately 350 utterances each — in order to maintain a fixed and comparable experimental setup across all conditions. The Proposed (VC) variant utilizes the voice conversion branch of the model to perform style transfer.
Note on naturalness and intelligibility: Outputs from cross-corpus speakers may exhibit reduced naturalness and intelligibility due to the small amount of training data and the acoustic mismatch between the source corpus and ESD during style transfer. Scaling the data would likely mitigate these effects. However, we deliberately keep the data budget fixed to isolate the contributions of the proposed approach and to more clearly expose both its benefits and its current limitations.
Speaker identity references
LJ Speech
p226 (VCTK)
p231 (VCTK)
Speaker 😠 Angry 😊 Happy 😢 Sad 😲 Surprise 😐 Neutral
Emotion ref
(ESD 0011 GT)
LJ Speech
p226 (VCTK)
p231 (VCTK)