Abstract
The proliferation of information-sharing platforms and the ease of access to diverse resources have led to an overwhelming volume of multimodal data that is increasingly difficult to process effectively. The integration of multiple data types, including text, images, video, and audio, highlights the growing importance of Multimodal Text Summarization (MMTS). Collecting and synthesizing existing research on this topic can provide a comprehensive foundation for advancing the field. Following a Systematic Literature Review (SLR) methodology, we addressed three pivotal research questions concerning methodologies, evaluation measures, and datasets in MMTS. Through a systematic analysis of 132 papers, we examined the strategies employed to address MMTS challenges, assessed the evaluation methods used to quantify performance, and compiled a detailed list of available datasets along with their limitations. This review offers critical insights and identifies future research directions, aiming to inform and guide continued innovation in this dynamic and evolving domain.
| Original language | English |
|---|---|
| Article number | 68 |
| Pages (from-to) | 1-38 |
| Number of pages | 38 |
| Journal | ACM Computing Surveys |
| Volume | 58 |
| Issue number | 3 |
| DOIs | |
| Publication status | Published - Feb 2026 |
Keywords
- fusion strategies
- multimodal evaluation
- Multimodal text summarization
- systematic literature review
Fingerprint
Dive into the research topics of 'A Systematic literature review on multimodal text summarization'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver