Javascript must be enabled for the correct page display

Speaker Differentiation in Synthetic Child-Like Voices: Perceptual Validation and UW-Related Acoustic Evidence from a Zero-Shot Voice Conversion Pipeline

Gao, Yuxiang (2026) Speaker Differentiation in Synthetic Child-Like Voices: Perceptual Validation and UW-Related Acoustic Evidence from a Zero-Shot Voice Conversion Pipeline. Master thesis, Voice Technology (VT).

[img]
Preview
PDF
MA5332370YGao.pdf

Download (277kB) | Preview

Abstract

Children’s speech is acoustically and developmentally variable, making child-like speaker differentiation challenging for both perception and speech technology. This thesis examines whether controlled synthetic child-like speech can be used to study speaker differentiation, and which acoustic cues support separation among synthetic child-like voices. A TTS + Seed-VC pipeline generated 35 synthetic child-like samples from five source sentences and seven target voices. A same/different listening test with 32 complete-case participants was combined with acoustic and computational analyses, including whole-file measures, UW vowel/formant features, speaker embeddings, MFCC classification, and adult-control validation. Listeners distinguished different synthetic target voices with high accuracy. Overall same/different accuracy was 85.9%, and different-speaker pairs reached 96.1% accuracy both within and across intended gender categories. However, same-target cross-utterance accuracy was lower at 65.6%, showing that target separation and same-target consistency are distinct challenges. The main acoustic finding was that UW-like vowel contexts, especially F2, F3, and ΔF2–F3, provided a promising interpretable cue area for speaker separation. Whole-file F0 was informative but insufficient alone, while embedding and MFCC analyses further supported target-related separability. Together, the findings provide perceptual validation for a controlled synthetic child-like voice space, identify UW-related upper-formant structure as a promising acoustic cue area, and demonstrate a mixed-evidence pipeline for studying child-like speaker separation.

Item Type: Thesis (Master)
Name supervisor: Coler, M.L. and Mouw, J.M. and Verkhodanova, V.
Date Deposited: 10 Jun 2026 13:13
Last Modified: 10 Jun 2026 13:13
URI: https://campus-fryslan.studenttheses.ub.rug.nl/id/eprint/840

Actions (login required)

View Item View Item