Skip to main navigation Skip to search Skip to main content

FMSA: few-shot multimodal sentiment analysis for social media via integrated prompt learning and vision-language models

Xianxun Zhu, Heyang Feng, Erik Cambria, Xiaohan Yu, Jose Santamaria, Xuhui Fan, Rui Wang*

*Corresponding author for this work

Research output: Contribution to journalArticlepeer-review

Abstract

Multimodal sentiment analysis (MSA) on social media is increasingly critical for understanding complex emotional expressions, yet it faces significant challenges in data-scarce environments where annotated multimodal content is limited. Here, we present FMSA, a novel framework that integrates prompt-based learning with advanced vision-language models to enable robust few-shot MSA. Our approach leverages an instruction-aware Query Transformer (Q-Former) to dynamically extract and align visual features with task-specific textual prompts, enhancing cross-modal fusion. We introduce a distributed consistency sampling strategy to construct representative few-shot datasets, ensuring statistical diversity under constrained conditions. Evaluated across six benchmark social media datasets-including MVSA-S, MVSA-M, and Twitter-Depression-FMSA outperforms state-of-the-art methods, achieving an accuracy of 63.47% and an F1 score of 57.34% on MVSA-S with full data, and 61.25% accuracy with 56.06% F1 in few-shot settings using just 1% of the data. By fine-tuning lightweight components while preserving pretrained model robustness, FMSA mitigates overfitting and delivers generalizable performance. We also release a curated few-shot dataset as a community resource. This framework advances MSA by offering an efficient, scalable solution for interpreting multimodal emotions in low-resource scenarios.

Original languageEnglish
JournalIEEE Transactions on Affective Computing
Early online date6 Apr 2026
DOIs
Publication statusE-pub ahead of print - 6 Apr 2026

Keywords

  • Few-shot
  • Multimodal Sentiment
  • Prompt Learning
  • Social Media

Fingerprint

Dive into the research topics of 'FMSA: few-shot multimodal sentiment analysis for social media via integrated prompt learning and vision-language models'. Together they form a unique fingerprint.

Cite this