Abstract
Memes are important because they serve as conduits for expressing emotions, opinions, and social commentary online, providing valuable insight into public sentiment, trends, and social interactions. By combining textual and visual elements, multi-modal fusion techniques enhance meme analysis, enabling the classification of offensive and sentimental memes effectively. Early and late fusion methods effectively integrate multi-modal data but face limitations. Early fusion integrates features from different modalities before classification. Late fusion combines classification outcomes from each modality after individual classification and reclassifies the combined results. This paper compares early and late fusion models in meme analysis. It showcases their efficacy in extracting meme concepts and classifying meme reasoning. Pre-trained vision encoders, including ViT and VGG-16, and language encoders such as BERT, AlBERT, and DistilBERT, were employed to extract image and text features. These features were subsequently utilized for performing both early and late fusion techniques. This paper further compares the explainability of fusion models through SHAP analysis. In comprehensive experiments, various classifiers such as XGBoost and Random Forest, along with combinations of different vision and text features across multiple sentiment scenarios, showcased the superior effectiveness of late fusion over early fusion.
| Original language | English |
|---|---|
| Title of host publication | WWW '24 Companion |
| Subtitle of host publication | Companion proceedings of the ACM on Web Conference 2024 |
| Place of Publication | New York |
| Publisher | Association for Computing Machinery |
| Pages | 1681-1689 |
| Number of pages | 9 |
| ISBN (Electronic) | 9798400701726 |
| DOIs | |
| Publication status | Published - 2024 |
| Event | 33rd ACM Web Conference, WWW 2024 - Singapore, Singapore Duration: 13 May 2024 → 17 May 2024 |
Conference
| Conference | 33rd ACM Web Conference, WWW 2024 |
|---|---|
| Country/Territory | Singapore |
| City | Singapore |
| Period | 13/05/24 → 17/05/24 |
Bibliographical note
Copyright the Author(s) 2024. Version archived for private and non-commercial use with the permission of the author/s and according to publisher conditions. For further rights please contact the publisher.Keywords
- Multi-modal Meme analysis
- fusion
- explainability
Fingerprint
Dive into the research topics of 'Decoding memes: a comprehensive analysis of late and early fusion models for explainable meme analysis'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver