su, sulede (2026) Seeing to Understand: Audio-Visual Chain-of-Thought Reasoning for Speech Emotion Recognition. Master thesis, Voice Technology (VT).
|
PDF
MAS6215254Sulede.pdf Download (459kB) | Preview |
Abstract
Multimodal large language models capable of processing audio and video have opened a new paradigm for speech emotion recognition (SER): rather than mapping acoustic features to an emotion label via a classification head, the model generates a reasoning chain that articulates why the perceived cues signal a particular emotion before committing to a prediction. However, when these models rea- son directly over raw video, they face a structural hallucination risk: without verifiable anchors, the model can generate confident but factually incorrect descriptions of facial expressions that are ab- sent from the recording. This thesis investigates whether feature-grounded audio-visual reasoning — in which a teacher language model (Gemini 2.5 Flash) reasons over pre-extracted facial action unit (AU) intensities and symbolic prosodic descriptors, rather than raw media — can suppress this hallucination and improve emotion recognition on the MELD benchmark. The AV-EmotionCoT dataset is constructed by extracting facial AU measurements via a Reti- naFace / TalkNet / OpenFace 3.0 pipeline and prosodic features via a librosa / pYIN / WhiStress pipeline, then using these symbolic features as teacher inputs to generate modality-specific reason- ing traces. Qwen2.5-Omni-7B is fine-tuned on these traces with LoRA and subsequently refined with Group Relative Policy Optimisation (GRPO) using a binary emotion-accuracy reward. Experiments across three parts reveal three key findings. First, a preliminary SFT study shows that simple one-sentence targets outperform elaborate chain-of-thought annotations (59.6% vs. 48.2% accuracy), demonstrating that reasoning-chain quality matters more than length. Second, feature- grounded SFT outperforms end-to-end generative multimodal SFT by 11 percentage points (57.8% vs. 46.6%), with the neutral-class F1 improving from a hallucination-driven collapse to 0.719, di- rectly confirming the hallucination hypothesis. Third, GRPO post-training provides consistent but modest improvements of 1–1.5 percentage points over the SFT base, with both training lines ex- hibiting early-peak over-training behaviour. The visual modality benefit anticipated from combining audio and video features does not materialise over audio-only baselines on MELD, suggesting that audio and visual emotion signals are largely redundant in the sitcom domain. These results establish feature-grounded reasoning as an effective approach for reducing hallu- cination in multimodal emotion reasoning, and highlight that supervision quality at the SFT stage is the dominant driver of performance — more so than either reasoning-chain elaborateness or rein- forcement learning post-training.
| Item Type: | Thesis (Master) |
|---|---|
| Name supervisor: | Schauble, J.K. |
| Date Deposited: | 30 Jun 2026 12:55 |
| Last Modified: | 30 Jun 2026 12:55 |
| URI: | https://campus-fryslan.studenttheses.ub.rug.nl/id/eprint/864 |
Actions (login required)
![]() |
View Item |
