Zheng, Kaiwen (2026) Investigating ASR Error Correction with a Large Language Model under the 1-best Hypothesis Setting. Master thesis, Voice Technology (VT).
|
PDF
MScs6438156KZheng.pdf Download (285kB) | Preview |
Abstract
Large language models (LLMs) have recently shown strong potential for automatic speech recognition (ASR) post-correction by leveraging contextual and semantic information to repair recognition errors. However, most practical ASR systems provide only a single decoding result, making it unclear how effective LLM-based correction remains under the 1-best hypothesis setting. In this thesis, I investigate LLM-based ASR error correction using Whisper-small and GPT-4o mini under realistic and domain-diverse speech conditions. The first experiment in this study was conducted on LibriSpeech and Earnings-22 to compare relatively clean read speech with real-world financial speech. Additional prompting and accent-region experiments were conducted to examine the impact of prompting constraints and regional accent variation on post-correction effectiveness. The results showed that GPT-4o mini consistently improved transcription quality under the 1-best setting. Compared with LibriSpeech, Earnings-22 has larger correction benefits. Furthermore, constrained prompting was much better than unconstrained prompts, and the results also demonstrated that few-shot prompting under the 1-best setting only gets limited additional benefits. The accent-region experiments further showed that the post-correction effectiveness varied greatly in different regional accent groups. Overall, this study shows that LLM-based ASR post-correction is a feasible method under the 1-best hypothesis setting, but its effectiveness strongly depends on the speech conditions, prompting strategy designs, and the contextual information preserved in the ASR output.
| Item Type: | Thesis (Master) |
|---|---|
| Name supervisor: | Nayak, S. |
| Date Deposited: | 15 Jun 2026 08:21 |
| Last Modified: | 15 Jun 2026 08:21 |
| URI: | https://campus-fryslan.studenttheses.ub.rug.nl/id/eprint/844 |
Actions (login required)
![]() |
View Item |
