计算机科学 ›› 2026, Vol. 53 ›› Issue (7): 71-79.doi: 10.11896/jsjkx.250900117

• 计算机图形学 & 多媒体 • 上一篇    下一篇

基于频域增强与多层次特征融合的图文情感分析模型

朱宇超1, 张顺香1,2,3, 闻博宇1, 孙亮1, 徐洋1   

  1. 1 安徽理工大学计算机科学与工程学院 安徽 淮南 232001
    2 合肥综合性国家科学中心人工智能研究院 合肥 230000
    3 淮南师范学院计算机学院 安徽 淮南 232038
  • 收稿日期:2025-09-18 修回日期:2025-12-29 出版日期:2026-07-15 发布日期:2026-07-10
  • 通讯作者: 张顺香(sxzhang@aust.edu.cn)
  • 作者简介:(906220114@qq.com)
  • 基金资助:
    国家自然科学基金面上项目(62476005,62076006);认知智能全国重点实验室开放课题(COGOS-2023HE02);安徽高校协同创新项目(GXXT-2021-008);安徽理工大学研究生创新基金(2025cx2104)

Frequency-augmented and Multi-level Feature Fusion for Image-Text Sentiment Analyzer

ZHU Yuchao1, ZHANG Shunxiang1,2,3, WEN Boyu1, SUN Liang1, XU Yang1   

  1. 1 School of Computer Science and Engineering,Anhui University of Science and Technology,Huainan,Anhui 232001,China
    2 Artificial Intelligence Research Institute of Hefei Comprehensive National Science Center,Hefei 230000,China
    3 School of Computer Science,Huainan Normal University,Huainan,Anhui 232038,China
  • Received:2025-09-18 Revised:2025-12-29 Published:2026-07-15 Online:2026-07-10
  • About author:ZHU Yuchao,born in 2001,postgra-duate,is a member of CCF(No.Z6016G).His main research interest is multimodal sentiment analysis.
    ZHANG Shunxiang,born in 1970,Ph.D,professor,Ph.D supervisor.His main research interests include Web mining,semantic search and complex network.
  • Supported by:
    National Natural Science Foundation of China(62476005,62076006),Opening Foundation of State Key Laboratory of Cognitive Intelligence, iFLYTEK(COGOS-2023HE02),University Synergy Innovation Program of Anhui Province(GXXT-2021-008) and Graduate Innovation Fund of Anhui University of Science and Technology(2025cx2104).

摘要: 图文情感分析利用文本和图像的互补信息,精准挖掘用户情感倾向。现有研究大多聚焦于深层语义特征交互,对图文模态浅层细节特征利用不足,造成情感极性识别准确率下降。针对这一问题,提出一种基于频域增强与多层次特征融合的图文情感分析模型。首先,通过特征提取模块提取图文模态的多层次特征,涵盖从浅层细节到深层语义的不同层次特征;其次,在频域增强模块中,通过傅里叶变换和坐标注意力机制协同优化,增强浅层特征的高频细节与深层特征的低频语义,从而提升不同层次特征表征能力;接着,在双模态协同交互模块中,通过双向注意力与动态校验,有效实现了图文多层次特征的双向跨模态交互;最后,在渐进融合模块中逐层融合图文特征,实现了从浅层细节特征到深层语义特征的有效利用。在数据集MVSA和HFM上进行的实验验证了所提模型的有效性。

关键词: 图文情感分析, 多层次特征, 频域增强, 跨模态交互, 双向交互

Abstract: Image-text sentiment analysis accurately mines users' emotional tendencies by utilizing the complementary information of text and images.Existing research mostly focuses on deep semantic feature interactions,and underutilizes shallow detailed features of the image-text modality,resulting in a decrease in the accuracy of sentiment polarity recognition.To address this pro-blem,this paper proposes a image-text sentiment analysis model based on frequency domain enhancement and multi-level feature fusion.Firstly,multi-level featuresfrom the image-text modality are extracted through the feature extraction module,covering different levels of features from shallow details to deep semantics.Secondly,in the frequency domain enhancement module,the high-frequency details of the shallow features and the low-frequency semantics of the deep features are enhanced through the synergistic optimization of the Fourier Transform and the coordinate attention mechanism,so as to improve the ability of different levels of feature characte-rization.Then,in the bimodal synergistic interaction module,through bidirectional attention and dynamic calibration,it effectively realizes the bidirectional cross-modal interaction of multi-level image and text features.Finally,the image and text features are fused layer by layer in the progressive fusion module,which realizes the effective utilization of the shallow detailed features to the deep semantic features.Experiments conducted on the dataset MVSA and HFM verify the effectiveness of the proposed model.

Key words: Image-text sentiment analysis, Multi-level features, Frequency domain enhancement, Cross-modal interaction, Bidirectional interaction

中图分类号: 

  • TP391
[1]GANDHI A,ADHVARYU K,PORIA S,et al.Multimodal sentiment analysis:A systematic review of history,datasets,multimodal fusion methods,applications,challenges and future directions[J].Information Fusion,2023,91:424-444.
[2]SINGH U,ABHISHEK K,AZAD H K.A survey of cutting-edge multimodal sentiment analysis[J].ACM Computing Surveys,2024,56(9):1-38.
[3]GUO X,BUYIDAN G,GULANBAIR T.A review of research on sentiment analysis algorithms based on multimodal fusion[J].Computer Engineering and Applications,2024,60(2):1-18.
[4]PANDEY A,VISHWAKARMA D K.Progress,achievements,and challenges in multimodal sentiment analysis using deep learning:A survey[J].Applied Soft Computing,2024,152:111206.
[5]LU Q,SUN X,LONG Y,et al.Sentiment analysis:Comprehensive reviews,recent advances,and open challenges[J].IEEE Transactions on Neural Networks and Learning Systems,2024,35(11):15092-15112.
[6]WANG D,GUO X,TIAN Y,et al.TETFN:A text enhanced transformer fusion network for multimodal sentiment analysis[J].Pattern Recognition,2023,136:109259.
[7]ZHANG S X,LIU J J,JIAO Y X,et al.A graphical sentiment analysis model based on sentiment representation calibration[J/OL].Journal of Beijing University of Aeronautics and Astronautics,1-14[2025-05-29].https://doi.org/10.13700/j.bh.1001-5965.2024.0545.
[8]TABOADA M,BROOKE J,TOFILOSKI M,et al.Lexicon-based methods for sentiment analysis[J].Computational linguistics,2011,37(2):267-307.
[9]DEY S,WASIF S,TONMOY D S,et al.A comparative study of support vector machine and Naive Bayes classifier for sentiment analysis on Amazon product reviews[C]//2020 International Conference on Contemporary Computing and Applications(IC3A).IEEE,2020:217-220.
[10]LI W,QI F,TANG M,et al.Bidirectional LSTM with self-attention mechanism and multi-channel features for sentiment classification[J].Neurocomputing,2020,387:63-77.
[11]DEVLIN J,CHANG M W,LEE K,et al.Bert:Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies.2019:4171-4186.
[12]YANG Y,SUN X,LU Q,et al.A sentiment and syntactic-aware graph convolutional network for aspect-level sentiment classification[C]//2023 IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP 2023).IEEE,2023:1-5.
[13]ZADEH A,CHEN M,PORIA S,et al.Tensor Fusion Network for Multimodal Sentiment Analysis[C]//Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing.2017:1103-1114.
[14]MAJUMDER N,HAZARIKA D,GELBUKH A,et al.Multi-modal sentiment analysis using hierarchical fusion with context modeling[J].Knowledge-based Systems,2018,161:124-133.
[15]ZHU T,LI L,YANG J,et al.Multimodal sentiment analysis with image-text interaction network[J].IEEE Transactions on Multimedia,2022,25:3375-3385.
[16]YANG J,YU Y,NIU D,et al.Confede:Contrastive feature decomposition for multimodal sentiment analysis[C]//Proceedings of the 61st Annual Meeting of the Association for Computa-tional Linguistics.2023:7617-7630.
[17]JI Y,WANG J,GONG Y,et al.Map:Multimodal uncertainty-aware vision-language pre-training model[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:23262-23271.
[18]SUN W,YAN X,SU Y,et al.MSDSANet:Multimodal Emotion Recognition Based on Multi-Stream Network and Dual-Scale Attention Network Feature Representation[J].Sensors(Basel,Switzerland),2025,25(7):2029.
[19]LI Z,XU B,ZHU C,et al.CLMLF:A Contrastive Learning and Multi-Layer Fusion Method for Multimodal Sentiment Detection[C]//Findings of the Association for Computational Linguistics:NAACL 2022.2022:2282-2294.
[20]LI Y,DING H,LIN Y,et al.Multi-level textual-visual align-ment and fusion network for multimodal aspect-based sentiment analysis[J].Artificial Intelligence Review,2024,57(4):78.
[21]NIU T,ZHU S,PANG L,et al.Sentiment analysis on multi-view social data[C]//Multi Media Modeling:22nd International Conference.Springer,2016:15-27.
[22]CAI Y,CAI H,WAN X.Multi-modal sarcasm detection in twitter with hierarchical fusion model[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.2019:2506-2515.
[23]ZHOU P,SHI W,TIAN J,et al.Attention-based bidirectional long short-term memory networks for relation classification[C]//Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics.2016:207-212.
[24]HE K,ZHANG X,REN S,et al.Deep residual learning for image recognition[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.2016:770-778.
[25]HAN K,WANG Y,CHEN H,et al.A survey on vision transformer[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2022,45(1):87-110.
[26]XU N.Analyzing multimodal public sentiment based on hierarchical semantic attentional network[C]//2017 IEEE International Conference on Intelligence and Security Informatics(ISI).IEEE,2017:152-154.
[27]XU N,MAO W,CHEN G.A co-memory network for multimodal sentiment analysis[C]//The 41st International ACM SIGIR Conference on Research & Development in Information Retrie-val.2018:929-932.
[28]PENG C,ZHANG C,XUE X,et al.Cross-modal complementary network with hierarchical fusion for multimodal sentiment classification[J].Tsinghua Science and Technology,2021,27(4):664-679.
[29]YANG X,FENG S,ZHANG Y,et al.Multimodal sentiment detection based on multi-channel graph neural networks[C]//Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing.2021:328-339.
[30]YU B,LI C,SHI Z.Multi-grained feature gating fusion network for multimodal sentiment analysis[J].Knowledge and Information Systems,2025,67:6879-6905.
[31]SCHIFANELLA R,DE JUAN P,TETREAULT J,et al.Detecting sarcasm in multimodal social platforms[C]//Proceedings of the 24th ACM International Conference on Multimedia.2016:1136-1145.
[32]XU N,ZENG Z,MAO W.Reasoning with multimodal sarcastictweets via modeling cross-modality contrast and semantic association[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.2020:3777-3786.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!