Abstract
Multimodal sentiment analysis (MSA) on social media is increasingly critical for understanding complex emotional expressions, yet it faces significant challenges in data-scarce environments where annotated multimodal content is limited. Here, we present FMSA, a novel framework that integrates prompt-based learning with advanced vision-language models to enable robust few-shot MSA. Our approach leverages an instruction-aware Query Transformer (Q-Former) to dynamically extract and align visual features with task-specific textual prompts, enhancing cross-modal fusion. We introduce a distributed consistency sampling strategy to construct representative few-shot datasets, ensuring statistical diversity under constrained conditions. Evaluated across six benchmark social media datasets-including MVSA-S, MVSA-M, and Twitter-Depression-FMSA outperforms state-of-the-art methods, achieving an accuracy of 63.47% and an F1 score of 57.34% on MVSA-S with full data, and 61.25% accuracy with 56.06% F1 in few-shot settings using just 1% of the data. By fine-tuning lightweight components while preserving pretrained model robustness, FMSA mitigates overfitting and delivers generalizable performance. We also release a curated few-shot dataset as a community resource. This framework advances MSA by offering an efficient, scalable solution for interpreting multimodal emotions in low-resource scenarios.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Affective Computing |
| Early online date | 6 Apr 2026 |
| DOIs | |
| Publication status | E-pub ahead of print - 6 Apr 2026 |
Keywords
- Few-shot
- Multimodal Sentiment
- Prompt Learning
- Social Media
Fingerprint
Dive into the research topics of 'FMSA: few-shot multimodal sentiment analysis for social media via integrated prompt learning and vision-language models'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver