Computer Science ›› 2026, Vol. 53 ›› Issue (8): 182-190.doi: 10.11896/jsjkx.251200069

• Computer Graphics & Multimedia • Previous Articles     Next Articles

Localization of Diffusion-based Image Editing via Inversion-Reconstruction

LING Yun, LI Haodong   

  1. Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen, Guangdong 518060, China
    Shenzhen Key Laboratory of Media Security, Shenzhen, Guangdong 518060, China
    SZU-AFS Joint Innovation Center for AI Technology, Shenzhen, Guangdong 518060, China
  • Received:2025-12-10 Revised:2026-04-13 Online:2026-08-15 Published:2026-08-17
  • About author:LING Yun,born in 1999,postgraduate.His main research interest is digital image forensics.
    LI Haodong,born in 1990,Ph.D,asso-ciate professor.His main research in-terest is multimedia forensics and security.
  • Supported by:
    National Natural Science Foundation of China(62572325,U23B2022),Guangdong Basic and Applied Basic Research Foundation(2025A1515010234) and Shenzhen R & D Program(SYSPG20241211174032004).

Abstract: Image tampering localization aims to detect and locate regions in an image that have been edited or forged.With the rapid proliferation of diffusion models for image generation and editing,even ordinary users can locally modify images with high photorealism via simple textual prompts,posing new challenges to existing forensic techniques.To tackle diffusion-based image editing,this paper proposes a tampering localization method built upon DDIM inversion and multi-step reconstruction consistency.The key observation is that when an edited image undergoes DDIM inversion followed by reconstruction along the denoising trajectory,edited(diffusion-generated) regions can be stably recovered,whereas pristine,unedited regions present noticeable inconsistencies due to their mismatch with the diffusion prior.Leveraging this insight,this paper designs a dual-stream network in which one branch processes the original image and the other one handles a series of reconstructed images obtained at multiple diffusion timesteps.A self-attention mechanism is introduced to explicitly model the dynamic reconstruction behavior across timesteps,while cross-attention is used to fuse domain-specific features from the pristine and diffusion-edited regions.To mitigate data scarcity,it also constructs a diffusion-edited image dataset with ground-truth masks by using several mainstream diffusion models.Experimental results on multiple benchmarks show that the proposed me-thod substantially outperforms existing approaches on both seen and unseen types of images,demonstrating strong robustness and generalization.

Key words: Image tampering localization, Diffusion-based image editing, Image inversion and reconstruction, Forgery image dataset

CLC Number: 

  • TP391
[1] WANG Z,BAO J,ZHOU W,et al.Dire for diffusion-generated image detection[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2023:22445-22455.
[2] BAYAR B,STAMM M C.A deep learning approach to universal image manipulation detection using a new convolutional layer[C]//Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Security.2016:5-10.
[3] RAO Y,NI J.A deep learning approach to detection of splicing and copy-move forgeries in images[C]//2016 IEEE Internatio-nal Workshop on Information Forensics and Security(WIFS).IEEE,2016:1-6.
[4] FRIDRICH J,KODOVSKY J.Rich models for steganalysis of digital images[J].IEEE Transactions on Information Forensics and Security,2012,7(3):868-882.
[5] COZZOLINO D,VERDOLIVA L.Noiseprint:A CNN-basedcamera model fingerprint[J].IEEE Transactions on Information Forensics and Security,2019,15:144-159.
[6] GUILLARO F,COZZOLINO D,SUD A,et al.Trufor:Leveraging all-round clues for trustworthy image forgery detection and localization[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:20606-20615.
[7] XIE E,WANG W,YU Z,et al.SegFormer:Simple and efficientdesign for semantic segmentation with transformers[J].Advances in Neural Information Processing Systems,2021,34:12077-12090.
[8] DONG C,CHEN X,HU R,et al.MVSS-Net:Multi-view multi-scale supervised networks for image manipulation detection[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2022,45(3):3539-3553.
[9] KWON M J,YU I J,NAM S H,et al.CAT-Net:Compression artifact tracing network for detection and localization of image splicing[C]//Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.2021:375-384.
[10] LIU X,LIU Y,CHEN J,et al.PSCC-Net:Progressive spatio-channel correlation network for image manipulation detection and localization[J].IEEE Transactions on Circuits and Systems for Video Technology,2022,32(11):7505-7517.
[11] GUO X,LIU X,REN Z,et al.Hierarchical fine-grained image forgery detection and localization[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:3155-3165.
[12] YU Z,NI J,LIN Y,et al.Diffforensics:Leveraging diffusion prior to image forgery detection and localization[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2024:12765-12774.
[13] XU Z,ZHANG X,LI R,et al.FakeShield:Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models[C]//The Thirteenth International Conference on Learning Representations.2025.
[14] SOHL-DICKSTEIN J,WEISS E,MAHESWARANATHANN,et al.Deep unsupervised learning using nonequilibrium thermodynamics[C]//International Conference on Machine Lear-ning.PMLR,2015:2256-2265.
[15] SONG J,MENG C,ERMON S.Denoising diffusion implicitmodels[J].arXiv:2010.02502,2020.
[16] LIU Z,MAO H,WU C Y,et al.A convnet for the 2020s[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:11976-11986.
[17] DONG J,WANG W,TAN T.Casia image tampering detection evaluation database[C]//2013 IEEE China Summit and International Conference on Signal and Information Processing.IEEE,2013:422-426.
[18] WEN B,ZHU Y,SUBRAMANIAN R,et al.COVERAGE—A novel database for copy-move forgery detection[C]//2016 IEEE International Conference on Image Processing(ICIP).IEEE,2016:161-165.
[19] GUAN H,KOZAK M,ROBERTSON E,et al.MFC datasets:Large-scale benchmark datasets for media forensic challenge evaluation[C]//2019 IEEE Winter Applications of Computer Vision Workshops(WACVW).IEEE,2019:63-72.
[20] DE CARVALHO T J,RIESS C,ANGELOPOULOU E,et al.Exposing digital image forgeries by illumination color classification[J].IEEE Transactions on Information Forensics and Secu-rity,2013,8(7):1182-1194.
[21] HSU Y F,CHANG S F.Detecting image splicing using geometry invariants and camera characteristics consistency[C]//2006 IEEE International Conference on Multimedia and Expo.IEEE,2006:549-552.
[22] ROMBACH R,BLATTMANN A,LORENZ D,et al.High-re-solution image synthesis with latent diffusion models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:10684-10695.
[23] JIA S,HUANG M,ZHOU Z,et al.Autosplice:A text-prompt manipulated image dataset for media forensics[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:893-903.
[24] NICHOL A,DHARIWAL P,RAMESH A,et al.Glide:To-wards photorealistic image generation and editing with text-guided diffusion models[J].arXiv:2112.10741,2021.
[25] LIN T Y,MAIRE M,BELONGIE S,et al.Microsoft coco:Common objects in context[C]//European Conference on Computer Vision.Cham:Springer,2014:740-755.
[26] MAREEN H,KARAGEORGIOU D,VAN WALLENDAEL G,et al.TGIF:Text-guided inpainting forgery dataset[C]//2024 IEEE International Workshop on Information Forensics and Security(WIFS).IEEE,2024:1-6.
[27] REN S,HE K,GIRSHICK R,et al.Faster R-CNN:Towardsreal-time object detection with region proposal networks[C]//Advances in Neural Information Processing Systems.2015.
[28] KIRILLOV A,MINTUN E,RAVI N,et al.Segment anything[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2023:4015-4026.
[29] LI J,LI D,XIONG C,et al.Blip:Bootstrapping language-image pre-training for unified vision-language understanding and ge-neration[C]//International conference on Machine Learning.PMLR,2022:12888-12900.
[30] JU X,LIU X,WANG X,et al.Brushnet:A plug-and-play image inpainting model with decomposed dual-branch diffusion[C]//European Conference on Computer Vision.Cham:Springer,2024:150-168.
[31] ZHUANG J,ZENG Y,LIU W,et al.A task is worth one word:Learning with task prompts for high-quality versatile image inpainting[C]//European Conference on Computer Vision.Cham:Springer,2024:195-211.
[32] AVRAHAMI O,FRIED O,LISCHINSKI D.Blended latent diffusion[J].ACM Transactions on Graphics,2023,42(4):1-11.
[1] SHEN Fengyin, FAN Hongjie. Research on Design of Multi-level SQL Optimization Framework Based on Big Data Business Workflows [J]. Computer Science, 2026, 53(8): 1-8.
[2] LAI Yingxu, ZHANG Yuwei, ZHUANG Junxi. Construction and Generation of SOR Label System for Learner Personalized Portrait [J]. Computer Science, 2026, 53(8): 9-19.
[3] LIU Yanze, HAN Bo, YUAN Jidong, SU Dongliang, REN Jia, CAI Zhiming, WANG Zhihai. Anomaly Detection in Time Series Based on Time-Frequency Contrastive Learning [J]. Computer Science, 2026, 53(8): 20-28.
[4] GUO Peilin, ZOU Zhiyi, WANG Bo, LUO Jiawei. STMVF:Novel Multi-view Graph Convolutional Network for Spatial Transcriptomics Cell Deconvolution with Dual Cross-attention Mechanism [J]. Computer Science, 2026, 53(8): 40-49.
[5] PAN Yuquan, YUAN Deyu, WANG Anran, JIA Yuan. Enhanced GNNs Across Social Networks User Identity Linkage Algorithm Based on HiddenFeatures [J]. Computer Science, 2026, 53(8): 50-60.
[6] LIAO Xuechao, ZOU Hang, LYU Peidong, ZENG Zhiqiang. Lithium-ion Battery Capacity Data Augmentation and Prediction Based on Diffusion,Denoise and Coding-Decoding Attention [J]. Computer Science, 2026, 53(8): 103-116.
[7] DENG Jiayan, TIAN Shirui, LIU Hou, ZHU Ningbo, DUAN Mingxing. Zero-shot Pedestrian Trajectory Prediction Method Based on Compositional Motion [J]. Computer Science, 2026, 53(8): 117-126.
[8] GUO Chenhui, CHEN Zebin, TAN Guang. Sparse-view Gaussian Splatting Consistent Reconstruction with Mask-guided Generative Prior [J]. Computer Science, 2026, 53(8): 127-138.
[9] ZHANG Zhe, LIU Junjie, ZHANG Baili. Improved DETR-based Method for CAD Graphic Element Recognition [J]. Computer Science, 2026, 53(8): 139-147.
[10] CAI Yi, WANG Xiaobin, CHEN Ruili, XU Jinfeng. Handwriting Gender Recognition Method Based on Multi-scale Directional Attention Transformer [J]. Computer Science, 2026, 53(8): 156-164.
[11] WANG Jiahui, WANG Hongyu, HAO Yingguang. Feature Aggregation with Joint Tracking:Video Object Detection in Occlusion Scenarios [J]. Computer Science, 2026, 53(8): 165-173.
[12] HU Changyu, FAN Xinyu, ZHANG Zhi, DING Zixu, ZHANG Zhengyue, PENG Juhong. Hierarchical Lightweight Micro-expression Recognition Based on Optical Flow Partitioned FeatureFusion [J]. Computer Science, 2026, 53(8): 174-181.
[13] WANG Yifan, YANG Peixuan, LU Yuansuo, LIU Mengjun. Teacher Trajectory Recognition and Reconstruction in Smart Classrooms via Multi-source Fusion [J]. Computer Science, 2026, 53(8): 191-200.
[14] WANG Jingyang, XUE Weimin, HUANG Min, WU Shaoguang. Improved YOLOv11n Model for Small Target Detection in UAV Aerial Images [J]. Computer Science, 2026, 53(8): 201-208.
[15] ZHONG Rui, YAN Hongwei, LIU Jiawei. Lightweight Low-resolution Face Recognition via Hierarchical Dynamic Feature Generation Distillation [J]. Computer Science, 2026, 53(8): 209-218.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!