Computer Science ›› 2026, Vol. 53 ›› Issue (9): 196-208.doi: 10.11896/jsjkx.250800071

• Computer Graphics & Multimedia • Previous Articles     Next Articles

MSBF-Net:Multi-modal Synergistic Fusion Method for Image Manipulation Detection and Localization

YANG Gang, LU Tianliang, ZHANG Peiyuan   

  1. College of Information Network Security,People’s Public Security University of China,Beijing 100038,China
  • Received:2025-08-18 Revised:2025-12-08 Online:2026-09-15 Published:2026-09-10
  • About author:YANG Gang,born in 2000,postgra-duate,is a member of CCF(No.A04488G).His main research interests include image forensics,and so on.
    LU Tianliang,born in 1985,Ph.D,professor,Ph.D supervisor.His main research interests include cyber security,artificial intelligence,and so on.
  • Supported by:
    Ministry of Public Security Science and Technology Plan Project(2024LL25) and Research and Public Security AI Model Research and Application Laboratory Project(2024300050036).

Abstract: To address the challenges of increasingly diverse and covert manipulation techniques in the field of image tampering detection,existing methods have attempted to fuse auxiliary modal information to enhance detection capabilities.However,they often only consider fusing a single auxiliary modality,which limits the comprehensive capture of diverse tampering clues and thereby suppresses overall fusion performance.To solve this key challenge,firstly,a multi-stage synergistic fusion network named MSBF-Net is proposed.This network adopts a Transformer-based encoder-decoder architecture capable of effectively processing feature information at different hierarchical levels.Secondly,a multi-branch modality-cooperative encoder is designed to jointly model RGB visual features with three complementary noise fingerprints-DCT,SRM,and NoisePrint++-to capture rich forgery clues from multiple dimensions and enhance the representation capability for diverse manipulations.Thirdly,to achieve efficient inter-action between the feature branches,a tri-source attention adaptive fusion module is designed.It deeply integrates the multi-path feature streams at different stages of the encoder,strengthening the complementary information between modalities through an adaptive attention mechanism.Finally,to solve the critical problem of modal imbalance in multi-modal learning and to improve localization precision,a mechanism based on class prototypes,comprising prototypical class entropy(PCE) loss and prototypical entropy regularization(PER),is designed to guide the balanced development of each modality.The effectiveness of this method is verified through comprehensive experiments on five public datasets,with the results demonstrating the comprehensive perfor-mance advantages of the MSBF-Net model.Its advanced nature is particularly evident when tackling the CocoGlide dataset,which is generated by advanced diffusion models.On this dataset,the model achieves state-of-the-art results in the two critical tasks of pixel-level localization and image-level detection,with performance improvements of approximately 3.82% and 5.95%,respectively,compared to the previous state-of-the-art model.Simultaneously,on multiple benchmark datasets featuring traditional manipulation types such as splicing and copy-move,the model’s core metrics also consistently remain in the top tier,with its average metrics being optimal,which verifies its strong generalization ability.Furthermore,ablation studies and robustness tests further confirm that the model’s internal components form an effective synergy and that its performance is stable under various image distortion conditions.

Key words: Image manipulation detection, Multi-modal fusion, Modal imbalance, Synergistic fusion, Image processing

CLC Number: 

  • TP391
[1] VACCARI C,CHADWICK A.Deepfakes and Disinformation:Exploring the Impact of Synthetic Political Video on Deception,Uncertainty,and Trust in News[J].Social Media + Society,2020,6(1):2056305120903408.
[2] GAFNI O,WOLF L.Wish you were here:Context-aware human generation[C] //Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2020:7840-7849.
[3] AVRAHAMI O,FRIED O,LISCHINSKI D.Blended LatentDiffusion[J].ACM Transactions on Graphics,2023,42(4):1-11.
[4] NICHOL A,DHARIWAL P,RAMESH A,et al.GLIDE:Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models[J].arXiv:2112.10741,2021.
[5] CAO Y,LI S,LIU Y,et al.A Comprehensive Survey of AI-Ge-nerated Content(AIGC):A History of Generative AI from GAN to ChatGPT[J].arXiv:2303.04226,2022.
[6] GUILLARO F,COZZOLINO D,SUD A,et al.TruFor:Leveraging all-round clues for trustworthy image forgery detection and localization[J].arXiv:2212.10957,2023.
[7] KWON M J,NAM S H,YU I J,et al.Learning JPEG compression artifacts for image manipulation detection and localization[J].International Journal of Computer Vision,2022,130(8):1875-1895.
[8] TRIARIDIS K,MEZARIS V.Exploring multi-modal fusion for image manipulation detection and localization[J].arXiv:2312.01790,2023.
[9] LI S,MA W,GUO J,et al.UnionFormer:Unified-LearningTransformer with Multi-View Representation for Image Manipulation Detection and Localization[C] //2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).IEEE,2024:12523-12533.
[10] WU Y,ABDALMAGEED W,NATARAJAN P.ManTra-Net:Manipulation Tracing Network for Detection and Localization of Image Forgeries With Anomalous Features[C] //2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).IEEE,2019:9535-9544.
[11] HU X,ZHANG Z,JIANG Z,et al.SPAN:Spatial Pyramid Attention Network for Image Manipulation Localization[M] //Computer Vision-ECCV 2020.Cham:Springer,2020:312-328.
[12] ZHOU P,HAN X,MORARIU V I,et al.Learning rich features for image manipulation detection[C] //2018 IEEE/CVF Confe-rence on Computer Vision and Pattern Recognition.IEEE,2018:1053-1061.
[13] MA X,ZHU X,SU L,et al.IMDL-BenCo:A Comprehensive Benchmark and Codebase for Image Manipulation Detection & Localization[J].arXiv:2406.10580,2024.
[14] FAN Y,XU W,WANG H,et al.PMR:Prototypical Modal Rebalance for Multimodal Learning[C] //2023 IEEE/CVF Confe-rence on Computer Vision and Pattern Recognition(CVPR).IEEE,2023:20029-20038.
[15] ZHANG J,LIU H,YANG K,et al.CMX:Cross-modal fusionfor RGB-X semantic segmentation with transformers[J].IEEE Transactions on Intelligent Transportation Systems,2023,24(12):14679-14694.
[16] KNIAZ V V,KNYAZ V,REMONDINO F.The point where reality meets fantasy:Mixed adversarial generators for image splice detection[J].Advances in Neural Information Processing Systems,2019,32:215-226.
[17] BAYAR B,STAMM M C.A Deep Learning Approach to Universal Image Manipulation Detection Using a New Convolutional Layer[C] //Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Security.ACM,2016:5-10.
[18] YANG C,LI H,LIN F,et al.Constrained R-Cnn:A GeneralImage Manipulation Detection Model[C] //2020 IEEE International Conference on Multimedia and Expo(ICME).IEEE,2020:1-6.
[19] CHEN X,DONG C,JI J,et al.Image Manipulation Detection by Multi-View Multi-Scale Supervision[C] //2021 IEEE/CVF International Conference on Computer Vision(ICCV).IEEE,2021:14165-14173.
[20] LI D,ZHU J,WANG M,et al.Edge-aware Regional Message Passing Controller for Image Forgery Localization[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).IEEE,2023:8222-8232.
[21] LIU Y,LV B,JIN X,et al.Tbformer:Two-branch transformer for image forgery localization[J].IEEE Signal Processing Letters,2023,30:623-627.
[22] COZZOLINO D,VERDOLIVA L.Noiseprint:a CNN-basedcamera model fingerprint[J].IEEE Transactions on Information Forensics and Security,2019,15:144-159.
[23] GU A R,NAM J H,LEE S C.FBI-Net:Frequency-based image forgery localization via multitask learning with self-attention[J].IEEE Access,2022,10:62751-62762.
[24] WANG W,TRAN D,FEISZLI M.What makes training multi-modal classification networks hard?[C] //Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2020:12695-12705.
[25] PENG X,WEI Y,DENG A,et al.Balanced multimodal learning via on-the-fly gradient modulation[C] //Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:8238-8247.
[26] XIAO F,LEE Y J,GRAUMAN K,et al.Audiovisual SlowFast Networks for Video Recognition[J].arXiv:2001.08740,2020.
[27] BIANCHI T,DE ROSA A,PIVA A.Improved DCT coefficient analysis for forgery localization in JPEG images[C] //2011 IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP).IEEE,2011:2444-2447.
[28] YANG L,ZHANG R Y,LI L,et al.Simam:A simple,parameter-free attention module for convolutional neural networks[C] //International Conference on Machine Learning.PMLR,2021:11863-11874.
[29] MILLETARI F,NAVAB N,AHMADI S A.V-Net:Fully Con-volutional Neural Networks for Volumetric Medical Image Segmentation[C] //2016 Fourth International Conference on 3D Vision(3DV).2016:565-571.
[30] DONG J,WANG W,TAN T.CASIA Image Tampering Detection Evaluation Database[C] //2013 IEEE China Summit and International Conference on Signal and Information Processing.IEEE,2013:422-426.
[31] NOVOZAMSKY A,MAHDIAN B,SAIC S.IMD2020:A large-scale annotated dataset tailored for detecting manipulated images[C] //Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision Workshops.2020:71-80.
[32] LIN T Y,MAIRE M,BELONGIE S,et al.Microsoft COCO:Common Objects in Context[M] //Computer Vision-ECCV 2014.Cham:Springer,2014:740-755.
[33] DANG-NGUYEN D T,PASQUINI C,CONOTTER V,et al.RAISE:a raw images dataset for digital image forensics[C] //Proceedings of the 6th ACM Multimedia Systems Conference.ACM,2015:219-224.
[34] WEN B,ZHU Y,SUBRAMANIAN R,et al.COVERAGE-A novel database for copy-move forgery detection[C] //2016 IEEE International Conference on Image Processing(ICIP).IEEE,2016:161-165.
[35] HSU Y F,CHANG S F.Detecting Image Splicing using Geometry Invariants and Camera Characteristics Consistency[C] //2006 IEEE International Conference on Multimedia and Expo.IEEE,2006:549-552.
[36] GUAN H,KOZAK M,ROBERTSON E,et al.MFC Datasets:Large-Scale Benchmark Datasets for Media Forensic Challenge Evaluation[C] //2019 IEEE Winter Applications of Computer Vision Workshops(WACVW).IEEE,2019:63-72.
[37] WU Y,ABD-ALMAGEED W,NATARAJAN P.Image copy-move forgery detection via an end-to-end deep neural network[C] //2018 IEEE Winter Conference on Applications of Computer Vision(WACV).IEEE,2018:1907-1915.
[38] KWON M J,YU I J,NAM S H,et al.CAT-net:Compression artifact tracing network for detection and localization of image splicing[C] //Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.2021:375-384.
[39] WANG J Z,LI J,WIEDERHOLD G.SIMPLIcity:Semantics-sensitive integrated matching for picture libraries[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2001,23(9):947-963.
[40] XIE E,WANG W,YU Z,et al.SegFormer:Simple and efficient design for semantic segmentation with transformers[C] //Advances in Neural Information Processing Systems.Curran Associates Inc.,2021:12077-12090.
[41] ZHANG J,LIU R,SHI H,et al.Delivering arbitrary-modal semantic segmentation[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).2023:1136-1147.
[42] LIU X,LIU Y,CHEN J,et al.PSCC-net:Progressive spatio-channel correlation network for image manipulation detection and localization[J].IEEE Transactions on Circuits and Systems for Video Technology,2022,32(11):7505-7517.
[43] DONG C,CHEN X,HU R,et al.MVSS-net:Multi-view multi-scale supervised networks for image manipulation detection[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2023,45(3):3539-3553.
[44] WU H,ZHOU J,TIAN J,et al.Robust image forgery detection over online social network shared images[C] //Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:13440-13449.
[45] SU L,MA X,ZHU X,et al.Can we get rid of handcrafted fea-ture extractors?SparseViT:Nonsemantics-centered,parameter-efficient image manipulation localization through sparse-coding transformer[C] //Proceedings of the AAAI Conference on Artificial Intelligence.2025:7024-7032.
[1] CHU Chunyu, JIANG Feilong. Water Meter Reading Recognition Based on Deep Learning and Prior Correction [J]. Computer Science, 2026, 53(6A): 250300143-7.
[2] SU Ye, XU Xin, ZHAO Longlong, LI Xiaoli, CHEN Pan, CHEN Jinsong. LitchiNet:Lightweight Litchi Variety Recognition Network with Fused Multi-scale Gated Attention and Class Imbalance Awareness [J]. Computer Science, 2026, 53(6A): 250600127-8.
[3] DONG Ye, LIAN Xinyue, WANG Yuyang, OU Xinyu. RGB-IR Multi-modal Fusion-based Tomato Small Object Detection [J]. Computer Science, 2026, 53(6A): 250700173-8.
[4] XU Cheng, LIU Yuxuan, WANG Xin, ZHANG Cheng, YAO Dengfeng, YUAN Jiazheng. Review of Speech Disorder Assessment Methods Driven by Large Language Models [J]. Computer Science, 2026, 53(3): 307-320.
[5] DU Jiantong, GUAN Zeli, XUE Zhe. Multi-task Learning-based Ophthalmic Video Feature Fusion and Multi-dimensional Profiling [J]. Computer Science, 2026, 53(3): 383-391.
[6] SHI Xincheng, WANG Baohui, YU Litao, DU Hui. Study on Segmentation Algorithm of Lower Limb Bone Anatomical Structure Based on 3D CTImages [J]. Computer Science, 2025, 52(6A): 240500119-9.
[7] SHANG Yunxian, CAI Guoyong, LIU Qinghua, JIANG Yiming. Active Learning-based Multi-modal Fusion Rumor Detection [J]. Computer Science, 2025, 52(12): 391-399.
[8] XU Ying, LI Xiaoming, YU Fenghao. Drug Name Recognition Method Based on CRAFT and OCR Technology [J]. Computer Science, 2025, 52(11A): 241200160-7.
[9] LUO Xin, LIANG Bo. RFI Suppression of the Yunnan 40-meter Radio Telescope Based on Deep Learning [J]. Computer Science, 2025, 52(11A): 250300044-7.
[10] SONG Lei, WANG Baohui, DU Hui. Optimization Study of Segmentation Algorithms for Lower Limb Bone on 3D CT Slices [J]. Computer Science, 2025, 52(11A): 240900072-7.
[11] LIN Zukai, HOU Guojia, WANG Guodong, PAN Zhenkuan. Image Deraining Based on Union Attention Mechanism and Multi-stage Feature Extraction [J]. Computer Science, 2025, 52(11): 206-212.
[12] LIU Yuming, DAI Yu, CHEN Gongping. Review of Federated Learning in Medical Image Processing [J]. Computer Science, 2025, 52(1): 183-193.
[13] HUANG Xiaofei, GUO Weibin. Multi-modal Fusion Method Based on Dual Encoders [J]. Computer Science, 2024, 51(9): 207-213.
[14] ZHANG Tianchi, LIU Yuxuan. Research Progress of Underwater Image Processing Based on Deep Learning [J]. Computer Science, 2024, 51(6A): 230400107-12.
[15] TAN Peng, OU Bo. Medical Image Reversible Contrast Enhancement Based on Adaptive Histogram Equalization [J]. Computer Science, 2024, 51(6A): 230700124-7.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!