Skip to main navigation Skip to search Skip to main content

Revealing the truth with ConLLM for detecting multi-modal deepfakes

Gautam Siddharth Kashyap, Harsh Joshi, Niharika Jain, Ebad Shabbir, Jiechao Gao*, Nipun Joshi, Usman Naseem*

*Corresponding author for this work

Research output: Chapter in Book/Report/Conference proceedingConference proceeding contributionpeer-review

Abstract

The rapid rise of deepfake technology poses a severe threat to social and political stability by enabling hyper-realistic synthetic media capable of manipulating public perception. However, existing detection methods struggle with two core limitations: (1) modality fragmentation, which leads to poor generalization across diverse and adversarial deepfake modalities; and (2) shallow inter-modal reasoning, resulting in limited detection of fine-grained semantic inconsistencies. To address these, we propose ConLLM (Contrastive Learning with Large Language Models), a hybrid framework for robust multimodal deepfake detection. ConLLM employs a two-stage architecture: stage 1 uses Pre-Trained Models (PTMs) to extract modality-specific embeddings; stage 2 aligns these embeddings via contrastive learning to mitigate modality fragmentation, and refines them using LLM-based reasoning to address shallow inter-modal reasoning by capturing semantic inconsistencies. ConLLM demonstrates strong performance across audio, video, and audio-visual modalities. It reduces audio deepfake EER by up to 50%, improves video accuracy by up to 8%, and achieves approximately 9% accuracy gains in audio-visual tasks. Ablation studies confirm that PTM-based embeddings contribute 9%–10% consistent improvements across modalities. Our code and data is available at: https://github.com/gskgautam/ConLLM/tree/main

Original languageEnglish
Title of host publicationFindings of the Association for Computational Linguistics
Subtitle of host publicationEACL 2026
EditorsVera Demberg, Kentaro Inui, Lluís Marquez
Place of PublicationKerrville, TX
PublisherAssociation for Computational Linguistics (ACL)
Pages1968-1978
Number of pages11
ISBN (Electronic)9798891763869
DOIs
Publication statusPublished - 2026
Event19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026 - Rabat, Morocco
Duration: 24 Mar 202629 Mar 2026

Conference

Conference19th Conference of the European Chapter of the Association for Computational Linguistics, Findings of EACL 2026
Country/TerritoryMorocco
CityRabat
Period24/03/2629/03/26

Fingerprint

Dive into the research topics of 'Revealing the truth with ConLLM for detecting multi-modal deepfakes'. Together they form a unique fingerprint.

Cite this