Computer Science ›› 2026, Vol. 53 ›› Issue (9): 299-308.doi: 10.11896/jsjkx.250600227

• Artificial Intelligence • Previous Articles     Next Articles

Multimodal Sentiment Analysis Based on Prompt Learning and Guided Gated Fusion Mechanism

FENG Guang1, SUN Xiangli2, LIN Yibao2, LIU Xinting1, CAO Yuqiao1, HUANG Junhui2, LIAO Beirong2   

  1. 1 School of Automation,Guangdong University of Technology,Guangzhou 510006,China
    2 School of Computer Science,Guangdong University of Technology,Guangzhou 510006,China
  • Received:2025-06-30 Revised:2025-11-27 Online:2026-09-15 Published:2026-09-10
  • About author:FENG Guang,born in 1973,Ph.D,professor.His main research interests include classroom streaming,big data and artificial intelligence.
  • Supported by:
    Key Program of the National Natural Science Foundation of China(62237001) and Youth Project of Guangdong Provincial Philosophy and Social Sciences(GD23YJY08).

Abstract: Multimodal sentiment analysis has emerged as one of the most prominent tasks in the multimodal learning domain,aiming to predict emotions by leveraging the complementary strengths of different modalities.Due to the significant heterogeneity among different modalities in terms of temporal structure,semantic space,and representational scales,achieving deep and consis-tent alignment remains inherently challenging,which impedes the full exploitation of affective features from audio and visual modalities.Moreover,most existing models adopt simple concatenation or weighted summation mechanisms during the fusion stage,lacking dynamic modeling and selective filtering of critical affective cues, thereby resulting in limited fusion expressiveness.To address these challenges,this paper proposes TPGF(Text-Guided Prompt and Gated Fusion Model),a novel multimodal sentiment analysis framework that integrates text-guided prompt learning and gated fusion mechanisms.Specifically,to mitigate modality heterogeneity,TPGF introduces a prompt learning module during cross-modal semantic interaction,leveraging the dominant semantic information from text to guide the alignment and fusion of audio and visual modalities.To overcome the limitations of simplistic fusion,a hierarchical cross-modal gated fusion mechanism is designed that performs dynamic selection and integration of salient features.Furthermore,semantic augmentation strategies for both text and audio modalities are introduced to enhance the robustness and expressive power of the model.Extensive experiments on two benchmark datasets,CMU-MOSI and CMU-MOSEI,demonstrate that TPGF significantly outperforms state-of-the-art models in terms of classification accuracy,F1 score,and other evaluation metrics,validating the effectiveness and superiority of the proposed framework.

Key words: Multimodal learning, Prompt learning, Gated fusion, Semantic enhancement, Sentiment analysis

CLC Number: 

  • TP391
[1] FENG G,LIU T X,YANG Y R,et al.A multimodal named entity recognition method based on image-text collaborative hierarchical fusion [J].Application Research of Computers,2025,42(6):1-9.
[2] ZHANG X M,WEI W,ZOU S H.Modalfeature optimization network with prompt for multimodal sentiment analysis[C] //Proceedings of the 31st International Conference on Computational Linguistics.Abu Dhabi,UAE:ACL,2025:4611-4621.
[3] GEETHANJALI R,VALARMATHI A.A novel hybrid deep learning IChOA-CNN-LSTM model for modality-enriched and multilingual emotion recognition in social media[J].Scientific Reports,2024,14(1):22270.
[4] QIN Z,LUO Q,ZANG Z,et al.Multimodal GRU with directed pairwise cross-modal attention for sentiment analysis[J].Scientific Reports,2025,15:10112.
[5] WU Y,LIU H,LU P,et al.Design and implementation of virtual fitting system based on gesture recognition and clothing transfer algorithm[J].Scientific Reports,2022,12:18356.
[6] MIAH M S U,KABIR M M,SARWAR T B,et al.A multimodal approach to cross-lingual sentiment analysis with ensemble of transformer and LLM[J].Scientific Reports,2024,14(1):9603.
[7] PANG L,ZHU S,NGO C W.Deep Multimodallearning for affective analysis and retrieval[J].IEEE Transactions on Multimedia,2015,17(11):2008-2020.
[8] FENG G,ZHOU Y H,ZHONG T,et al.Multimodal sentiment analysis combining adaptive feature weighting and weight optimization strategy [J].Computer Engineering and Applications,2025,61(12):1-12.
[9] ZHANG Y X,LIN X X,WANG S,et al.Multi-modal fusion emotional recognition algorithm based on facial images and HRV[J].Journal of Nanjing University of Information Science & Technology,2026,18(2):183-191.
[10] SAILUNAZ K,ALHAJJ R.Emotion and sentiment analysisfrom twitter text[J].Journal of Computer Science,2019,36:101003.
[11] JUNCHI M,CHAUDHRY H N,KULSOOM F,et al.MULTICAUSENET temporal attention for multimodal emotion cause pair extraction[J].Scientific Reports,2025,15:19372.
[12] SONG X J.The study of multimodal emotion recognition based on text,speech and video[D].Jinan:Shandong University,2019.
[13] DEVLIN J,CHANG M W,LEE K,et al.BERT:Pre-training of deep bidirectional transformers for language understanding[J].arXiv:1810.04805,2018.
[14] YU B G,XING Y,ZHANG S W.Aspect-level sentiment analysis model based on multimodal collaborative contrastive learning[J].Data Analysis and Knowledge Discovery,2024,8(11):22-32.
[15] HUANG J,TAO J,LIU B,et al.Multimodal transformer fusion for continuous emotion recognition[C] //ICASSP 2020.IEEE,2020:3507-3511.
[16] PRAVEEN R G,GRANGER E,CARDINAL P.Cross atten-tional audio-visual fusion for dimensional emotion recognition[C] //2021 IEEE International Conference on Automatic Face and Gesture Recognition.2021:1-8.
[17] WILLIAMS J,KLEINEGESSE S,COMANESCU R,et al.Recognizing emotions in video using multimodal DNN feature fusion[C] //Proceedings of the 2018 Grand Challenge Workshop on Human Multimodal Language.ACL,2018:11-19.
[18] WU H,WU Q,LIN B,et al.Network analysis and sentimentclassification ofminnan nursery rhymes[J].npj Heritage Science,2025,13:172.
[19] PORIA S,CHATURVEDI I,CAMBRIA E,et al.Convolutional MKL based multimodal emotion recognition and sentiment analysis[C] //2016 IEEE International Conference on Data Mining.2016:439-448.
[20] SONG Y F,REN G,YANG Y,et al.Multimodal sentiment analysis based on hybrid feature fusion of multi-level attention mechanism and multi-task learning[J].Application Research of Computers,2022,39(3):716-720.
[21] ZADEH A,LIANG P P,MAZUMDER N,et al.Memory fusion network for multi-view sequential learning[C] //Proceedings of the AAAI Conference on Artificial Intelligence.2018.
[22] LIU Y,OTT M,GOYAL N,et al.Roberta:A robustly opti-mizedBERT pretraining approach[J].arXiv:1907.11692,2019.
[23] ZADEH A,ZELLERS R,PINCUS E,et al.MOSI:Multimodal corpus of sentiment intensity and subjectivity analysis in online opinion videos[J].arXiv:1606.06259,2016.
[24] ZADEH A A B,LIANG P P,PORIA S,et al.Multimodal language analysis in the wild:CMU-MOSEI dataset and interpretable dynamic fusion graph[C] //Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics(Volume 1:Long Papers).ACL,2018:2236-2246.
[25] ZADEH A,CHEN M,PORIA S.Tensor fusion network formultimodal sentiment analysis[C] //Proceedings of EMNLP.2017:1103-1114.
[26] YU W,XU H,YUAN Z,et al.Learning modality-specific representations with self-supervised multi-task learning for multimodal sentiment analysis[C] //AAAI.2021:10790-10797.
[27] LIU Z,SHEN Y,LAKSHMINARASIMHAN V B,et al.Efficient low-rank multimodal fusion with modality-specific factors[C] //Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics(Volume 1:Long Papers).ACL,2018:2247-2256.
[28] RAHMAN W,HASAN M K,LEE S,et al.Integrating multimodal information in large pretrained transformers[C] //Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.ACL,2020:2359-2369.
[29] HAZARIKA D,ZIMMERMANN R,PORIA S.MISA:Modality-invariant and-specific representations for multimodal sentiment analysis[C] //Proceedings of ACM Multimedia.2020:1122-1131.
[30] HUANG J,ZHOU J,TANG Z,et al.TMBL:Transformer-based multimodal binding learning model for multimodal sentiment analysis[J].Knowledge-Based Systems,2024,285:111346.
[31] LIU W,XU H,HUA Y,et al.AdaFN-AG:Enhancing multimodal interaction with adaptive feature normalization for multimodal sentiment analysis[J].Intelligent Systems with Applications,2024,23:200410.
[32] LI M,ZHU Z,LI K,et al.Joint training strategy of unimodal and multimodal for multimodal sentiment analysis[J].Image and Vision Computing,2024,149:105172.
[33] WANG P,ZHOU Q,WU Y,et al.DLF:Disentangled-language-focused multimodal sentiment analysis[C] //Proceedings of the AAAI Conference on Artificial Intelligence.2025:21180-21188.
[34] SONG Y,CHO S.Leveraging CLIPencoder for multimodal emotion recognition[C] //2025 IEEE/CVF Winter Conference on Applications of Computer Vision(WACV).IEEE,2025:6115-6124.
[1] WANG Xingyue, YE Hongting, XU Honghua, ZHOU Suyang, KONG Youyong. Multimodal Renewable Energy Data Feature Graph Modeling Method Based on Hard Prompts [J]. Computer Science, 2026, 53(9): 145-156.
[2] ZHU Yuchao, ZHANG Shunxiang, WEN Boyu, SUN Liang, XU Yang. Frequency-augmented and Multi-level Feature Fusion for Image-Text Sentiment Analyzer [J]. Computer Science, 2026, 53(7): 71-79.
[3] DING Zhijun, AI Fangju, LIU Aihuan. Document-level Event Argument Extraction Model Based on Hierarchical Dependency Aggregation and Event Enhancement [J]. Computer Science, 2026, 53(7): 156-167.
[4] SHEN Jianwei, CHEN Jiawen, CHEN Hanlin, MA Xinjian, CHEN Xing. Construction and Application of Dataset Knowledge Graph Based on Metadata Semantic Enhancement [J]. Computer Science, 2026, 53(6A): 250500052-10.
[5] ZHANG Yongyu, GUO Chenjuan, FEI Xueqin, LI Feng. Study on Financial Text Sentiment Analysis Method Based on Large Language Models with Market Feedback Supervision [J]. Computer Science, 2026, 53(6A): 250500073-14.
[6] ZHAO Jingyun, LIU Keying, GUO Wenke. FIN-GDAN:Sentiment Adversarial Transfer Network for Shanghai Gold Futures News [J]. Computer Science, 2026, 53(6A): 250700179-9.
[7] KE Changbo, LI Tianhao, ZHANG Bolei, XIAO Fu, XU Kang. Teaching Evaluation Sentiment Analysis Method Based on Capsule Network [J]. Computer Science, 2026, 53(6): 10-18.
[8] SHEN Ao, ZHOU Qingkai, XIA Tian, GAO Ruiling. Span-based Aspect Sentiment Triplet Extraction Based on Multi-view Graph Neural Networks [J]. Computer Science, 2026, 53(5): 319-327.
[9] ZHENG Cheng, BAN Qingqing. Knowledge-assisted and Reinforced Syntax-driven for Aspect-based Sentiment Analysis [J]. Computer Science, 2026, 53(4): 406-414.
[10] TAN Pingping, XU Ji, LI Yijun, WANG Hai. Dynamic Interaction Dual-channel Graph Attention Network for Chinese and English SarcasmDetection [J]. Computer Science, 2026, 53(2): 300-311.
[11] CHEN Lin, MA Longxuan, ZHANG Yongbing, HUANG Yuxin, GAO Shengxiang, YU Zhengtao. Industrial Text Classification for Chinese and Vietnamese Based on Prompt Learning and AdaptiveLoss Weighting [J]. Computer Science, 2026, 53(2): 312-321.
[12] BU Yunyang, QI Binting, BU Fanliang. Multimodal Sentiment Analysis for Interactive Fusion of Dual Perspectives Under Cross-modalInconsistent Perception [J]. Computer Science, 2026, 53(1): 187-194.
[13] CHEN Qian, CHENG Kaixuan, GUO Xin, ZHANG Xiaoxia, WANG Suge, LI Yanhong. Bidirectional Prompt-Tuning for Event Argument Extraction with Topic and Entity Embeddings [J]. Computer Science, 2026, 53(1): 278-284.
[14] CAI Qihang, XU Bin, DONG Xiaodi. Knowledge Graph Completion Model Using Semantically Enhanced Prompts and Structural Information [J]. Computer Science, 2025, 52(9): 282-293.
[15] CHENG Zhangtao, HUANG Haoran, XUE He, LIU Leyuan, ZHONG Ting, ZHOU Fan. Event Causality Identification Model Based on Prompt Learning and Hypergraph [J]. Computer Science, 2025, 52(9): 303-312.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!