计算机科学 ›› 2026, Vol. 53 ›› Issue (8): 182-190.doi: 10.11896/jsjkx.251200069

• 计算机图形学 & 多媒体 • 上一篇    下一篇

基于反演重建的扩散模型图像编辑区域定位方法

凌云, 李昊东   

  1. 广东省智能信息处理重点实验室 广东 深圳 518060
    深圳市媒体信息内容安全重点实验室 广东 深圳 518060
    深大-软牛人工智能技术联合创新中心 广东 深圳 518060
  • 收稿日期:2025-12-10 修回日期:2026-04-13 出版日期:2026-08-15 发布日期:2026-08-17
  • 通讯作者: 李昊东(lihaodong@szu.edu.cn)
  • 作者简介:(1179300981@qq.com)
  • 基金资助:
    国家自然科学基金(62572325,U23B2022);广东省自然科学基金(2025A1515010234);深圳市科技计划项目(SYSPG20241211174032004)

Localization of Diffusion-based Image Editing via Inversion-Reconstruction

LING Yun, LI Haodong   

  1. Guangdong Provincial Key Laboratory of Intelligent Information Processing, Shenzhen, Guangdong 518060, China
    Shenzhen Key Laboratory of Media Security, Shenzhen, Guangdong 518060, China
    SZU-AFS Joint Innovation Center for AI Technology, Shenzhen, Guangdong 518060, China
  • Received:2025-12-10 Revised:2026-04-13 Published:2026-08-15 Online:2026-08-17
  • About author:LING Yun,born in 1999,postgraduate.His main research interest is digital image forensics.
    LI Haodong,born in 1990,Ph.D,asso-ciate professor.His main research in-terest is multimedia forensics and security.
  • Supported by:
    National Natural Science Foundation of China(62572325,U23B2022),Guangdong Basic and Applied Basic Research Foundation(2025A1515010234) and Shenzhen R & D Program(SYSPG20241211174032004).

摘要: 图像篡改定位旨在检测并标定图像中被编辑或伪造的区域。随着扩散模型在图像生成与编辑领域的迅速普及,普通用户仅需通过简单指令即可对图像局部进行高逼真度的修改,这对图像篡改定位技术提出了新挑战。针对基于扩散模型的图像编辑,提出了一种基于DDIM反演与多步重建一致性分析的篡改定位方法。该方法的核心发现是:对经扩散模型编辑的图像进行DDIM反演并逐步重建时,编辑区域因符合扩散先验而能够被稳定复原,而源自真实图像的未编辑区域则表现出显著的不一致性。基于该现象,设计了一个双流网络:一个分支接收原始输入图像,另一分支接收通过多个时间步重建获得的图像。网络通过自注意力显式建模跨时间步的重建动态,并利用交叉注意力融合未编辑区域与编辑区域的特征,从而实现精确的篡改定位。为缓解训练数据不足的问题,构建了一个涵盖多种主流扩散模型及具有伪造区域掩码标注的篡改图像数据集。实验结果表明,所提方法在多个测试数据集上对已知和未知模型编辑的图像都取得了显著优于现有方法的性能,展现了良好的鲁棒性与泛化能力。

关键词: 图像篡改定位, 扩散模型图像编辑, 图像反演与重建, 伪造图像数据集

Abstract: Image tampering localization aims to detect and locate regions in an image that have been edited or forged.With the rapid proliferation of diffusion models for image generation and editing,even ordinary users can locally modify images with high photorealism via simple textual prompts,posing new challenges to existing forensic techniques.To tackle diffusion-based image editing,this paper proposes a tampering localization method built upon DDIM inversion and multi-step reconstruction consistency.The key observation is that when an edited image undergoes DDIM inversion followed by reconstruction along the denoising trajectory,edited(diffusion-generated) regions can be stably recovered,whereas pristine,unedited regions present noticeable inconsistencies due to their mismatch with the diffusion prior.Leveraging this insight,this paper designs a dual-stream network in which one branch processes the original image and the other one handles a series of reconstructed images obtained at multiple diffusion timesteps.A self-attention mechanism is introduced to explicitly model the dynamic reconstruction behavior across timesteps,while cross-attention is used to fuse domain-specific features from the pristine and diffusion-edited regions.To mitigate data scarcity,it also constructs a diffusion-edited image dataset with ground-truth masks by using several mainstream diffusion models.Experimental results on multiple benchmarks show that the proposed me-thod substantially outperforms existing approaches on both seen and unseen types of images,demonstrating strong robustness and generalization.

Key words: Image tampering localization, Diffusion-based image editing, Image inversion and reconstruction, Forgery image dataset

中图分类号: 

  • TP391
[1] WANG Z,BAO J,ZHOU W,et al.Dire for diffusion-generated image detection[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2023:22445-22455.
[2] BAYAR B,STAMM M C.A deep learning approach to universal image manipulation detection using a new convolutional layer[C]//Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Security.2016:5-10.
[3] RAO Y,NI J.A deep learning approach to detection of splicing and copy-move forgeries in images[C]//2016 IEEE Internatio-nal Workshop on Information Forensics and Security(WIFS).IEEE,2016:1-6.
[4] FRIDRICH J,KODOVSKY J.Rich models for steganalysis of digital images[J].IEEE Transactions on Information Forensics and Security,2012,7(3):868-882.
[5] COZZOLINO D,VERDOLIVA L.Noiseprint:A CNN-basedcamera model fingerprint[J].IEEE Transactions on Information Forensics and Security,2019,15:144-159.
[6] GUILLARO F,COZZOLINO D,SUD A,et al.Trufor:Leveraging all-round clues for trustworthy image forgery detection and localization[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:20606-20615.
[7] XIE E,WANG W,YU Z,et al.SegFormer:Simple and efficientdesign for semantic segmentation with transformers[J].Advances in Neural Information Processing Systems,2021,34:12077-12090.
[8] DONG C,CHEN X,HU R,et al.MVSS-Net:Multi-view multi-scale supervised networks for image manipulation detection[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2022,45(3):3539-3553.
[9] KWON M J,YU I J,NAM S H,et al.CAT-Net:Compression artifact tracing network for detection and localization of image splicing[C]//Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.2021:375-384.
[10] LIU X,LIU Y,CHEN J,et al.PSCC-Net:Progressive spatio-channel correlation network for image manipulation detection and localization[J].IEEE Transactions on Circuits and Systems for Video Technology,2022,32(11):7505-7517.
[11] GUO X,LIU X,REN Z,et al.Hierarchical fine-grained image forgery detection and localization[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:3155-3165.
[12] YU Z,NI J,LIN Y,et al.Diffforensics:Leveraging diffusion prior to image forgery detection and localization[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2024:12765-12774.
[13] XU Z,ZHANG X,LI R,et al.FakeShield:Explainable Image Forgery Detection and Localization via Multi-modal Large Language Models[C]//The Thirteenth International Conference on Learning Representations.2025.
[14] SOHL-DICKSTEIN J,WEISS E,MAHESWARANATHANN,et al.Deep unsupervised learning using nonequilibrium thermodynamics[C]//International Conference on Machine Lear-ning.PMLR,2015:2256-2265.
[15] SONG J,MENG C,ERMON S.Denoising diffusion implicitmodels[J].arXiv:2010.02502,2020.
[16] LIU Z,MAO H,WU C Y,et al.A convnet for the 2020s[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:11976-11986.
[17] DONG J,WANG W,TAN T.Casia image tampering detection evaluation database[C]//2013 IEEE China Summit and International Conference on Signal and Information Processing.IEEE,2013:422-426.
[18] WEN B,ZHU Y,SUBRAMANIAN R,et al.COVERAGE—A novel database for copy-move forgery detection[C]//2016 IEEE International Conference on Image Processing(ICIP).IEEE,2016:161-165.
[19] GUAN H,KOZAK M,ROBERTSON E,et al.MFC datasets:Large-scale benchmark datasets for media forensic challenge evaluation[C]//2019 IEEE Winter Applications of Computer Vision Workshops(WACVW).IEEE,2019:63-72.
[20] DE CARVALHO T J,RIESS C,ANGELOPOULOU E,et al.Exposing digital image forgeries by illumination color classification[J].IEEE Transactions on Information Forensics and Secu-rity,2013,8(7):1182-1194.
[21] HSU Y F,CHANG S F.Detecting image splicing using geometry invariants and camera characteristics consistency[C]//2006 IEEE International Conference on Multimedia and Expo.IEEE,2006:549-552.
[22] ROMBACH R,BLATTMANN A,LORENZ D,et al.High-re-solution image synthesis with latent diffusion models[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:10684-10695.
[23] JIA S,HUANG M,ZHOU Z,et al.Autosplice:A text-prompt manipulated image dataset for media forensics[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:893-903.
[24] NICHOL A,DHARIWAL P,RAMESH A,et al.Glide:To-wards photorealistic image generation and editing with text-guided diffusion models[J].arXiv:2112.10741,2021.
[25] LIN T Y,MAIRE M,BELONGIE S,et al.Microsoft coco:Common objects in context[C]//European Conference on Computer Vision.Cham:Springer,2014:740-755.
[26] MAREEN H,KARAGEORGIOU D,VAN WALLENDAEL G,et al.TGIF:Text-guided inpainting forgery dataset[C]//2024 IEEE International Workshop on Information Forensics and Security(WIFS).IEEE,2024:1-6.
[27] REN S,HE K,GIRSHICK R,et al.Faster R-CNN:Towardsreal-time object detection with region proposal networks[C]//Advances in Neural Information Processing Systems.2015.
[28] KIRILLOV A,MINTUN E,RAVI N,et al.Segment anything[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2023:4015-4026.
[29] LI J,LI D,XIONG C,et al.Blip:Bootstrapping language-image pre-training for unified vision-language understanding and ge-neration[C]//International conference on Machine Learning.PMLR,2022:12888-12900.
[30] JU X,LIU X,WANG X,et al.Brushnet:A plug-and-play image inpainting model with decomposed dual-branch diffusion[C]//European Conference on Computer Vision.Cham:Springer,2024:150-168.
[31] ZHUANG J,ZENG Y,LIU W,et al.A task is worth one word:Learning with task prompts for high-quality versatile image inpainting[C]//European Conference on Computer Vision.Cham:Springer,2024:195-211.
[32] AVRAHAMI O,FRIED O,LISCHINSKI D.Blended latent diffusion[J].ACM Transactions on Graphics,2023,42(4):1-11.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!