Skip to main navigation Skip to search Skip to main content

Heterogeneity over homogeneity: investigating multilingual speech pre-trained models for detecting audio deepfake

Orchid Chetia Phukan*, Gautam Siddharth Kashyap, Arun Balaji Buduru, Rajesh Sharma

*Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference proceeding contributionpeer-review

Abstract

In this work, we investigate multilingual speech Pre-Trained models (PTMs) for Audio deepfake detection (ADD). We hypothesize that multilingual PTMs trained on large-scale diverse multilingual data gain knowledge about diverse pitches, accents, and tones, during their pre-training phase and making them more robust to variations. As a result, they will be more effective for detecting audio deepfakes. To validate our hypothesis, we extract representations from state-of-the-art (SOTA) PTMs including monolingual, multilingual as well as PTMs trained for speaker and emotion recognition, and evaluated them on ASVSpoof 2019 (ASV), In-the-Wild (ITW), and DECRO benchmark databases. We show that representations from multilingual PTMs, with simple downstream networks, attain the best performance for ADD compared to other PTM representations, which validates our hypothesis. We also explore the possibility of fusion of selected PTM representations for further improvements in ADD, and we propose a framework, MiO (Merge into One) for this purpose. With MiO, we achieve SOTA performance on ASV and ITW and comparable performance on DECRO with current SOTA works.

Original languageEnglish
Title of host publicationFindings of the Association for Computational Linguistics
Subtitle of host publicationNAACL 2024
Place of PublicationKerrville, TX
PublisherAssociation for Computational Linguistics
Pages2496-2506
Number of pages11
ISBN (Electronic)9798891761193
DOIs
Publication statusPublished - 2024
Externally publishedYes
Event2024 Annual Conference of the North American Association for Computational Linguistics - Hybrid, Mexico City, Mexico
Duration: 16 Jun 202421 Jun 2024

Conference

Conference2024 Annual Conference of the North American Association for Computational Linguistics
Abbreviated titleNAACL 2024
Country/TerritoryMexico
CityMexico City
Period16/06/2421/06/24

Fingerprint

Dive into the research topics of 'Heterogeneity over homogeneity: investigating multilingual speech pre-trained models for detecting audio deepfake'. Together they form a unique fingerprint.

Cite this