Wang, Junjuan (2026) Integrating Speaker Embeddings into TIGER for Target Speaker Extraction: A Comparative Study of Injection Strategies. Master thesis, Voice Technology (VT).
|
PDF
MScs6295509JWang.pdf Download (6MB) | Preview |
Abstract
This study investigates the integration of d-vector speaker embeddings into the TIGER (Time- frequency Interleaved Gain Extraction and Reconstruction) architecture for target speaker extraction, with a systematic comparison of three injection strategies, namely bottleneck-level conditioning, intra-separator modulation, and a dual-level hybrid approach, as well as three projection dimension- alities. To enhance the utilisation of speaker embeddings, a combination of cosine-similarity-based gating and Feature-wise Linear Modulation (FiLM) is employed. Experiments on the Libri2Mix dataset show that intra-separator FiLM conditioning achieves the strongest signal-level performance and the lowest speaker confusion rate among the three strategies, while bottleneck-level injection and dual-level conditioning perform at a comparable but lower level. A complementary experiment on projection dimensionality finds that performance is not monotonically related to dimensionality, and that the optimal choice is determined by the conditioning mechanism and internal feature di- mensions of the architecture rather than by scale alone. The results show that speaker confusion is the most common problem across all models, which means that handling speaker identity more effectively is still a major challenge for future work.
| Item Type: | Thesis (Master) |
|---|---|
| Name supervisor: | Nayak, S. |
| Date Deposited: | 15 Jun 2026 08:20 |
| Last Modified: | 15 Jun 2026 08:20 |
| URI: | https://campus-fryslan.studenttheses.ub.rug.nl/id/eprint/842 |
Actions (login required)
![]() |
View Item |
