计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250400127-6.doi: 10.11896/jsjkx.250400127

• 人工智能 • 上一篇    下一篇

融合语义建模与协同注意力机制的多模态讽刺识别方法

魏巍1, 李弼程1, 朱振水2, 左军2   

  1. 1 华侨大学计算机科学与技术学院 福建 厦门 361021
    2 厦门市美亚柏科信息安全研究所有限公司 福建 厦门 361008
  • 出版日期:2026-06-16 发布日期:2026-06-12
  • 通讯作者: 李弼程(lbclm@163.com)
  • 作者简介:(23014083056@stu.hqu.edu.cn)
  • 基金资助:
    公共安全领域多模态大模型构建及产品研发与产业化应用(3502Z20241029)

Semantic Modeling and Co-attention Mechanism for Multimodal Sarcasm Detection Method

WEI Wei1, LI Bicheng1, ZHU Zhenshui2, ZUO Jun2   

  1. 1 College of Computer Science and Technology,Huaqiao University,Xiamen,Fujian 361021,China
    2 Xiamen Meiya Boke Information Security Research Institute Company Limited,Xiamen,Fujian 361008,China
  • Published:2026-06-16 Online:2026-06-12
  • About author:WEI Wei,born in 1999,postgraduate.His main research interests include natural language processing,multimodal network public opinion knowledge graph.
    LI Bicheng,born in 1970,professor,doctoral supervisor.His main research interests include intelligent information processing,network ideological security,online public opinion monitoring and guidance,as well as big data analysis and mining.
  • Supported by:
    Construction of Multimodal Large Models in the Public Safety Domain,Product Development,and Industrial Application(3502Z20241029).

摘要: 讽刺被广泛应用于社交媒体和其他形式的以计算机为媒介的通信中,结合文本和图像信息的多模态讽刺识别,面临着多样化和复杂性的挑战,常常依赖于语言与图像等多模态信息之间的隐含对比与语义冲突。为了更有效地捕捉这种跨模态语义差异,文中提出了一种融合语义建模与协同注意力机制(Co-Attention Transformer)的多模态讽刺识别方法。该方法结合CLIP预训练模型的文本和图像的特征表示能力,采用协同注意力机制融合文本和图像特征,以更好地捕捉多模态间的深度交互与特征融合。此外,结合依存树信息进行图结构建模,并引入语义相似度增强,来有效捕捉文本和图像之间的语义一致性,从而提升讽刺识别的精度。使用公开的讽刺检测数据集进行实验验证,结果表明了所提方法相较于传统方法取得更优的性能。

关键词: 多模态讽刺识别, CLIP模型, 语义相似度, 协同注意力机制, 依存树, 图结构

Abstract: Sarcasm is widely used in social media and other forms of computer-mediated communication.Multimodal sarcasm detection,which leverages both textual and visual information,faces challenges due to the diversity and complexity of content,often relying on implicit contrast and semantic conflict across modalities.To better capture such cross-modal semantic discrepancies,this paper proposes a method that integrates semantic modeling with a co-attention mechanism(Co-Attention Transformer).Leveraging the representational power of the CLIP pre-trained model,the approach employs co-attention to enhance deep interaction and feature fusion across modalities.Moreover,it incorporates syntactic dependency trees for graph-based modeling and introduces semantic similarity enhancement to improve semantic alignment between text and image.Experiments on a public sarcasm detection dataset demonstrate the superiority of the proposed method over traditional baselines.

Key words: Multimodal sarcasm detection, CLIP model, Semantic similarity, Co-attention mechanism, Dependency tree, Graph structure

中图分类号: 

  • TP391
[1] TIWARI D,KANOJIA D,RAY A,et al.Predict and use:Harnessing predicted gaze to improve multimodal sarcasm detection[C]//Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing.2023:15933-15948.
[2] HUANG B,YU G.Research on the mining of opinion communi-ty for social media based on sentiment analysis and regional distribution[C]//2016 Chinese Control and Decision Conference(CCDC).IEEE,2016:6900-6905.
[3] LI L,JIN D,WANG X,et al.Multi-modal sarcasm detectionbased on cross-modal composition of inscribed entity relations[C]//2023 IEEE 35th International Conference on Tools with Artificial Intelligence(ICTAI).IEEE,2023:918-925.
[4] TAY Y,TUAN L A,HUI S C,et al.Reasoning with sarcasm by reading in-between[J].arXiv:1805.02856,2018.
[5] LOU C,LIANG B,GUI L,et al.Affective dependency graph for sarcasm detection[C]//Proceedings of the 44th international ACM SIGIR Conference on Research and Development in Information Retrieval.2021:1844-1849.
[6] WANG R,WANG Q,LIANG B,et al.Masking and generation:An unsupervised method for sarcasm detection[C]//Proceedings of the 45th International ACM SIGIR Conference on Research and Development in Information Retrieval.2022:2172-2177.
[7] FRENDA S,CIGNARELLA A T,BASILE V,et al.The unbearable hurtfulness of sarcasm[J].Expert Systems with Applications,2022,193:116398.
[8] YUE T,MAO R,WANG H,et al.KnowleNet:Knowledge fusion network for multimodal sarcasm detection[J].Information Fusion,2023,100:101921.
[9] LIU H,WEI R,TU G,et al.Sarcasm driven by sentiment:Asentiment-aware hierarchical fusion network for multimodal sarcasm detection[J].Information Fusion,2024,108:102353.
[10] TIAN Y,XU N,ZHANG R,et al.Dynamic routing transformer network for multimodal sarcasm detection[C]//Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics(Volume 1:Long Papers).2023:2468-2480.
[11] LIANG B,LOU C,LI X,et al.Multi-modal sarcasm detection via cross-modal graph convolutional network[C]//Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics(Volume 1:Long Papers).2022:1767-1777.
[12] LIU H,WANG W,LI H.Towards Multi-Modal Sarcasm Detection via Hierarchical Congruity Modeling with Knowledge Enhancement[C]//2022 Conference on Empirical Methods in Natural Language Processing(EMNLP 2022).Association for Computational Linguistics,2022:4995-5006.
[13] SCHIFANELLA R,DE JUAN P,TETREAULT J,et al.Detecting sarcasm in multimodal social platforms[C]//Proceedings of the 24th ACM international conference on Multimedia.2016:1136-1145.
[14] CAI Y,CAI H,WAN X.Multi-modal sarcasm detection in twitter with hierarchical fusion model[C]//Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics.2019:2506-2515.
[15] XU N,ZENG Z,MAO W.Reasoning with multimodal sarcastic tweets via modeling cross-modality contrast and semantic association[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.2020:3777-3786.
[16] LIANG B,LOU C,LI X,et al.Multi-modal sarcasm detectionwith interactive in-modal and cross-modal graphs[C]//Proceedings of the 29th ACM International Conference on Multimedia.2021:4707-4715.
[17] PAN H,LIN Z,FU P,et al.Modeling intra and inter-modality incongruity for multi-modal sarcasm detection[C]//Findings of the Association for Computational Linguistics:EMNLP 2020.2020:1383-1392.
[18] WU Q,FANG W,ZHONG W,et al.Dual-level adaptive incongruity-enhanced model for multimodal sarcasm detection[J].Neurocomputing,2025,612:128689.
[19] QIN L,HUANG S,CHEN Q,et al.MMSD2.0:Towards a reliable multi-modal sarcasm detection system[J].arXiv:2307.07135,2023.
[20] XU N,ZENG Z,MAO W.Reasoning with multimodal sarcastic tweets via modeling cross-modality contrast and semantic association[C]//Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics.2020:3777-3786.
[21] DOSOVITSKIY A,BEYER L,KOLESNIKOV A,et al.Animage is worth 16x16 words:Transformers for image recognition at scale[C]//9th International Conference on Learning Representations(ICLR 2021).VirtualEvent,OpenReview.net,2021.
[22] CHEN Y.Convolutional neural network for sentence classification[D].University of Waterloo,2015.
[23] DEVLIN J,CHANG M W,LEE K,et al.Bert:Pre-training of deep bidirectional transformers for language understanding[C]//Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies,volume 1(Long and Short Papers).2019:4171-4186.
[24] PAN H,LIN Z,FU P,et al.Modeling intra and inter-modality incongruity for multi-modal sarcasm detection[C]//Findings of the Association for Computational Linguistics(EMNLP 2020).2020:1383-1392.
[25] LIANG B,LOU C,LI X,et al.Multi-modal sarcasm detectionvia cross-modal graph convolutional network[C]//Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics(Volume 1:Long Papers).2022:1767-1777.
[26] YUE T,MAO R,WANG H,et al.KnowleNet:Knowledge fusion network for multimodal sarcasm detection[J].Information Fusion,2023,100:101921.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!