计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250700023-10.doi: 10.11896/jsjkx.250700023

• 信息安全 • 上一篇    下一篇

面向多模态深度检测器的黑盒物理对抗攻击

郑海斌1,2,3,4, 林秀豪2, 韩烨2, 陈晋音1,2, 李贝贝3   

  1. 1 浙江工业大学计算机学院 杭州 310023
    2 浙江工业大学信息工程学院 杭州 310023
    3 四川大学数据安全防护与智能治理教育部重点实验室 成都 610065
    4 北京生命科技研究院有限公司重点实验室 北京 102206
  • 出版日期:2026-06-16 发布日期:2026-06-12
  • 通讯作者: 陈晋音(chenjinyin@zjut.edu.cn)
  • 作者简介:(haibinzheng320@gmail.com)
  • 基金资助:
    国家自然科学基金(62406286);浙江省自然科学基金(LDQ23F020001);四川大学数据安全防护与智能治理教育部重点实验室放课题(SCUSAKFKT202402Z);北京生命科技研究院有限公司开放基金(2024200CD0210)

Black-box Physical Adversarial Attack Against Multimodal Object Detector

ZHENG Haibin1,2,3,4, LIN Xiuhao2, HAN Ye2, CHEN Jinyin1,2, LI Beibei3   

  1. 1 College of Computer Science,Zhejiang University of Technology,Hangzhou 310023,China
    2 College of Information Engineering,Zhejiang University of Technology,Hangzhou 310023,China
    3 Key Laboratory of Data Protection and Intelligent Management,Ministry of Education,Sichuan University,Chengdu 610065,China
    4 Key Laboratory of Beijing Life Science Academy,Beijing 102206,China
  • Published:2026-06-16 Online:2026-06-12
  • About author:ZHENG Haibin,born in 1995,Ph.D,lecturer.His main research interests include deep learning and artificial intelligence security.
    CHEN Jinyin,born in 1982,Ph.D,professor,is a member of CCF(No.14348M).Her main research interests include data mining,intelligent computing and complex network analysis.
  • Supported by:
    National Natural Science Foundation of China(62406286),Zhejiang Provincial Natural Science Foundation(LDQ23F020001),Key Laboratory of Data Protection and Intelligent Management,Ministry of Education,Sichuan University(SCUSAKFKT202402Z) and Beijing Life Science Academy(2024200CD0210).

摘要: 针对实际复杂工况,基于深度学习的多模态(可见光、红外等)目标检测器通过融合不同波段数据特征,提高检测效果。然而,已有研究发现,多模态检测器容易遭受对抗攻击,导致输出检测框严重偏离目标或检测框消失,降低其在物理世界使用的可靠性。已有工作探索了面向多模态深度检测器的黑盒物理对抗攻击,但依然存在模态攻击低效、目标检测器与融合策略受限、物理域攻击赋形效果差等问题。针对上述问题,提出了一种多模态对抗色块补丁(Multimodal Adversarial Color Patch,MAC-Patch)生成方法,实现对多模态深度检测器的高效、通用、鲁棒攻击。具体而言,在等效模型上利用随机梯度下降优化器生成针对不同模态的强对抗补丁,并在无需访问目标模型内部结构的黑盒设置下,依然能够有效干扰目标模型;提出了基于差分进化的补丁位置优化方法,自适应多种目标融合策略、目标检测模型以及防御设置下的最优攻击位置选择。最后,分别在2个模型、3种图像融合策略和4个数据集上测试MAC-Patch的攻击有效性、通用性和迁移性;在物理域采用期望平移变换,在不同环境亮度、不同补丁旋转角下的实际攻击效果,验证其鲁棒性。实验结果表明,MAC-Patch在攻击成功率、AP降低值等指标均最优,如相较于MAP,MIC,UAP 3种先进攻击,MAC-Patch的AP降低值提高了62.6%1)

关键词: 目标检测, 多模态, 对抗攻击, 物理攻击, 防御

Abstract: For real-world complex working conditions,deep learning-based multimodal(visible,infrared,etc.) target detectors improve the detection effect by fusing data features from different bands.However,it has been found that multimodal detectors are susceptible to adversarial attacks,resulting in the output detection frames being severely off-target or the detection frames disappearing,which reduces their reliability for use in the physical world.Work has been done to explore black-box physical adversarial attacks for multimodal depth detectors,but there are still problems such as inefficient modal attacks,limited target detectors and fusion strategies,and poor physical domain attack assignment.Aiming at the above problems,this paper proposes a multimodal adversarial color patch(MAC-Patch) generation method to achieve efficient,general,and robust attacks on multimodal deep detectors.Specifically,a stochastic gradient descent optimizer is utilized to generate strong adversarial patches against different modalities on the equivalent model,and can still effectively interfere with the target model in a black-box setting without accessing the internal structure of the target model.A patch location optimization method based on differential evolution is proposed to adaptively select the optimal attack location under multiple target fusion strategies,target detection models,and defense settings.Finally,the attack effectiveness,generalization and migration of MAC-Patch are tested on 2 models,3 image fusion strategies and 4 datasets respectively;the actual attack effect under different environment brightness and different patch rotation angles is adopted in the physical domain with expectation translation transformation to verify its robustness.Experimental results show that MAC-Patch is optimal in terms of attack success rate,AP reduction value and other indexes,such as compared with the three advanced attacks of MAP,MIC,and UAP,the AP reduction value of MAC-Patch is improved by 62.6%.

Key words: Object detection, Multimodality, Adversarial attack, Physical attack, Defense

中图分类号: 

  • TP391.4
[1] WANG X X,CHEN J,HE K,et al.A survey on adversarial attack and defense for object detection[J].Journal on Communications,2023,44(11):260-277.
[2] JONES J.Tesla,Ideal,Azure,Xiaopeng Intelligent TechnologyLayout Differences[EB/OL].https://news.qq.com/rain/a/20211012A0B9HX00.
[3] LI T H.Research on multimodal pedestrian recognition methods in harsh environments[D].Xi'an:Xi'an Technological University,2023.
[4] WEI X,YU J,HUANG Y.Infrared adversarial patches withlearnable shapes and locations in the physical world [J].International Journal of Computer Vision,2024,132:1928-1944.
[5] CHENG Y H,SHI W W,TIAN L.Adversarial color projection:A projector-based physical-world attack to DNNs[J].Image and Vision Computing,2023,140:104861.
[6] ZHU X P,HU Z H,HUANG S Y,et al.Infrared invisible clothing hiding from infrared detectors at multiple angles in the real world[C]//Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).Piscataway,NJ:IEEE,2022:13317-13326.
[7] REDMON J,FARHADI A.YOLOv3:an incremental improvement[J].arXiv:1804.02767,2018.
[8] REN S,HE K,GIRSHICK R,et al.Faster R-CNN:Towards real-time object detection with region proposal networks[C]//Proceedings of Advances in Neural Information Processing Systems(NeurIPS).Curran Associates Inc.,2015:91-99.
[9] LI H,WU X.DenseFuse:A fusion approach to infrared and visible images[J].IEEE Transactions on Image Processing,2019,28(5):2614-2623.
[10] LIU J,FAN X,HUANG Z,et al.Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection[C]//Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).2022:5792-5801.
[11] WANG Z,CHEN Y,SHAO W,et al.SwinFuse:A residual swin transformer fusion network for infrared and visible images[J].IEEE Transactions on Instrumentation and Measurement,2022,71:1-12.
[12] LI G,XU Y,DING J,et al.Toward generic and controllable attacks against object detection [J].IEEE Transactions on Geoscience and Remote Sensing,2024,62:1-12.
[13] AMIRA G,RUITIAN D,MUHAMMAD A H,et al.A dynamic adversarial patch for evading person detectors[J].arXiv:2305.11618,2023.
[14] WANG Y,LI X,YANG L,et al.Adaptive oriented adversarial attacks on visible and infrared image fusion models [C]//2024 IEEE International Conference on Multimedia and Expo(ICME).IEEE,2024:1-6.
[15] HUANG Y,DONG Y,RUAN S,et al.Towards transferabletargeted 3d adversarial attack in the physical world[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2024:24512-24522.
[16] GOODFELLOW I J,SHLENS J,SZEGERDY C.Explaining and harnessing adversarial examples[C]//International Conference on Learning Representations(ICLR),2015.
[17] SHAFAHI A,HUANG W R,NAJIBI M,et al.Are adversarial examples inevitable?[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision(ICCV).Springer,2019:3407-3416.
[18] HWANG S,PARK J,KIM N,et al.Multispectral pedestrian detection:Benchmark dataset and baselines[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).IEEE Computer Society,2018:5386-5394.
[19] JIA X,ZHU C,LI M,et al.LLVIP:A visible-infrared paired dataset for low-light vision[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision(ICCV).Springer,2021:2380-2389.
[20] LIU J,FAN X,HUANG Z,et al.Target-aware dual adversarial learning and a multi-scenario multi-modality benchmark to fuse infrared and visible for object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).IEEE Computer Society,2021:2379-2388.
[21] XU H,MA J,LE Z,et al.FusionDN:A unified densely connected network for image fusion[C]//Proceedings of the AAAI Conference on Artificial Intelligence.2020:12484-12491.
[22] WEI X X,HUANG Y,SUN Y T,et al.Unified adversarial patch for cross-modal attacks in the physical world[J].arXiv:2307.07859,2023.
[23] KIM T,LEE H J,RO Y M.MAP:Multispectral adversarialpatch to attack person detection[C]//Proceedings of the 2022 IEEE International Conference on Acoustics,Speech and Signal Processing(ICASSP).2022:4853-4857.
[24] KIM T,YU Y,RO Y M.Multispectral invisible coating:Laminated visible-thermal physical attack against multispectral object detectors using transparent low-e films[C]//Proceedings of the Thirty-Seventh AAAI Conference on Artificial Intelligence and Thirty-Fifth Conference on Innovative Applications of Artificial Intelligence and Thirteenth Symposium on Educational Advances in Artificial Intelligence.AAAI,2023:1151-1159.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!