Javascript must be enabled for the correct page display

When Replaying Fools the Detector: Applying Grad-SAM to Explain False Positives in Audio Deepfake Detection Under Replay Attack Conditions

Toma, Victor (2026) When Replaying Fools the Detector: Applying Grad-SAM to Explain False Positives in Audio Deepfake Detection Under Replay Attack Conditions. Master thesis, Voice Technology (VT).

[img]
Preview
PDF
MScS4553020VToma.pdf

Download (3MB) | Preview

Abstract

Synthetic speech generation has advanced to the point where audio deepfakes, artificially produced voice recordings generated by text-to-speech (TTS) systems or voice conversion algorithms, can be almost impossible to distinguish from genuine human speech by ear. This has given rise to the development of automatic deepfake detection systems, which learn to identify the subtle artefacts that synthesis processes leave in the acoustic signal. The ASVspoof challenge series has been the primary driver of progress in this field, and the best-performing detection systems now combine self-supervised speech representations with graph attention networks to achieve Equal Error Rates (EERs) below 1% on standard evaluation benchmarks (Tak et al., 2022; Kang et al., 2024). The robustness of these systems to real-world acoustic conditions is considerably less well understood. Müller et al. (2025) demonstrated that a loudspeaker, a microphone, and a few minutes of effort are enouugh to significantly degrade one of the best audio deepfake detector models. The method consists of playing TTS-generated audio through a speaker and re-recording it. This replay process is not complex, but what it reveals about how these detectors actually work is significant. The top-performing model in Müller et al.'s study, W2V2-AASIST (Tak et al., 2022), saw its EER climb from 4.7% to 18.2% when applied to replayed samples, a near four time degradation in detection performance. The question of why this degradation occurs is not straightforwardly answered. Detection models built on self-supervised speech representations do not provide explanations alongside their decisions. When a replayed deepfake fools the model, there is no direct way to determine what happened internally. One possibility is that replay redistributes acoustic energy in ways that mask the synthesis artefacts the model relies upon, causing it to attend to room acoustic characteristics that more closely resemble genuine speech. To the best of our knowledge, no published work has examined this question at the level of internal model representations, leaving the mechanistic basis of the performance collapse open.

Item Type: Thesis (Master)
Name supervisor: Do, T.P.
Date Deposited: 15 Jun 2026 08:21
Last Modified: 15 Jun 2026 08:21
URI: https://campus-fryslan.studenttheses.ub.rug.nl/id/eprint/843

Actions (login required)

View Item View Item