计算机科学 ›› 2026, Vol. 53 ›› Issue (6): 263-269.doi: 10.11896/jsjkx.250700103
刘继康1, 黄磊1, 张科1, 聂婕1, 魏志强1,2
LIU Jikang1, HUANG Lei1, ZHANG Ke1, NIE Jie1, WEI Zhiqiang1,2
摘要: 目标检测是计算机视觉领域一项重要的基础性任务,已在自动驾驶、智慧交通和医疗诊断等领域得到广泛应用。在检测领域,DETR(DEtection TRansformer)摒弃了手动设计组件的需求,构建了一个完全端到端的目标检测框架。然而,该框架解码器各层对前序输出的高度依赖导致了解码阶段出现级联负优化的问题。针对上述问题,提出了一种基于动态特征融合的目标检测方法——DFF DETR(Dynamic Feature Fusion DETR)。该方法通过动态融合解码器跨层特征,实现目标查询在解码层间的有效传递,缓解解码器后续阶段对前序输出的高度依赖。在反向传播阶段,引入层间传递的监督信号,提升中间层目标查询的特征拟合效果,修正解码器前序输出,缓解解码阶段的级联负优化问题。在多种基于DETR框架的目标检测算法上进行的大量实验表明,嵌入DFF方法可实现近1.0% mAP的性能增益,其中小目标检测精度的提升尤其明显。
中图分类号:
| [1]LI B,YAN J,WU W,et al.High performance visual tracking with siamese region proposal network[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2018:8971-8980. [2]MENG L,YANG X.A Survey of Object Tracking Algorithms [J].Acta Automatica Sinica,2019,45(7):1244-1260. [3]HE K,GKIOXARI G,DOLLÁR P,et al.Mask r-cnn[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.Piscataway,NJ:IEEE,2017:2961-2969. [4]WANG Z Y,YUAN C,LI J C.Instance Segmentation with Se-parable Convolutions and Multi-level Features [J].Journal of Software,2019,30(4):954-961. [5]XIAO T,LIU Y,ZHOU B,et al.Unified perceptual parsing for scene understanding[C]//Proceedings of the European Confe-rence on Computer Vision.Berlin:Springer,2018:418-434. [6]WU D H,YE X Q,GU W K.An Uncertain Knowledge Based Real Time Road Scene Understanding Algorithm [J].Journal of Image and Graphics,2002(1):71-76. [7]TANG W B,LI F.Semi-supervised object detection algorithm based on feature alignment and feature fusion [J].Journal of Chongqing Technology and Business University(Natural Science Edition),2025,42(1):35-41. [8]LIU S,HUANG D,WANG Y.Adaptive nms:Refining pedestrian detection in a crowd[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Pisca-taway,NJ:IEEE,2019:6459-6468. [9]SUN Z,BEBIS G,MILLER R.On-road vehicle detection:A review [J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2006,28(5):694-711. [10]HUANG K Q,CHEN X T,KANG Y F,et al.Intelligent Visual Surveillance:A Review [J].Chinese Journal of Computers,2015,38(6):1093-1118. [11]VASWANI A.Attention is all you need [C]//Proceedings of the 31st International Conference on Neural Information Processing Systems.2017:6000-6010. [12]TIAN Y L,WANG Y T,WANG J G,et al.Key Problems and Progress of Vision Transformers:The State of the Art and Prospects [J].Acta Automatica Sinica,2022,48(4):957-979. [13]CARION N,MASSA F,SYNNAEVE G,et al.End-to-end object detection with transformers [C]//Proceedings of the European Conference on Computer Vision.Berlin:Springer,2020:213-229. [14]CHEN F,ZHANG H,HU K,et al.Enhanced training of query-based object detection via selective query recollection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2023:23756-23765. [15]LIU S,LI F,ZHANG H,et al.DAB-DETR:Dynamic anchorboxes are better queries for DETR[C]//Proceedings of the International Conference on Learning Representations.2022. [16]ZHU X,SU W,LU L,et al.Deformable detr:Deformable transformers for end-to-end object detection[C]//Proceedings of the International Conference on Learning Representations.2021. [17]LAW H,DENG J.Cornernet:Detecting objects as paired keypoints[C]//Proceedings of the European Conference on Computer Vision.Berlin:Springer,2018:734-750. [18]MENG D,CHEN X,FAN Z,et al.Conditional detr for fasttraining convergence[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.Piscataway,NJ:IEEE,2021:3651-3660. [19]WANG Y,ZHANG X,YANG T,et al.Anchor detr:Query design for transformer-based detector[C]//Proceedings of the AAAI Conference on Artificial Intelligence.Menlo Park,CA:AAAI,2022:2567-2575. [20]LI F,ZHANG H,LIU S,et al.Dn-detr:Accelerate detr training by introducing query denoising[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2022:13619-13627. [21]ZHANG H,LI F,LIU S,et al.DINO:DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection[C]//Proceedings of the International Conference on Learning Representations.2023. [22]CHEN Q,CHEN X,WANG J,et al.Group detr:Fast detr trai-ning with group-wise one-to-many assignment[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.Piscataway,NJ:IEEE,2023:6633-6642. [23]JIA D,YUAN Y,HE H,et al.Detrs with hybrid matching[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2023:19702-19712. [24]HOU X,LIU M,ZHANG S,et al.Salience detr:Enhancing detection transformer with hierarchical salience filtering refinement[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2024:17574-17583. [25]HUANG Y X,LIU H I,SHUAI H H,et al.Dq-detr:Detr with dynamic query for tiny object detection[C]//Proceedings of the European Conference on Computer Vision.Berlin:Springer,2024:290-305. [26]HOU X,LIU M,ZHANG S,et al.Relation detr:Exploring explicit position relation prior for object detection[C]//Procee-dings of the European Conference on Computer Vision.Berlin:Springer,2024:89-105. [27]TEED Z,DENG J.Raft:Recurrent all-pairs field transforms for optical flow[C]//Proceedings of the European Conference on Computer Vision.Berlin:Springer,2020:402-419. [28]LIN T Y,MAIRE M,BELONGIE S,et al.Microsoft coco:Common objects in contex[C]//Proceedings of the European Conference on Computer Vision.Berlin:Springer,2014:740-755. [29]WANG C Y,YEH I H,MARK L H Y.Yolov9:Learning what you want to learn using programmable gradient information[C]//Proceedings of the European Conference on Computer Vision.Berlin:Springer,2024:1-21. [30]ZHAO C,SUN Y,WANG W,et al.MS-DETR:Efficient DETR training with mixed supervision[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2024:17027-17036. [31]ZHAO Y,LYU W,XU S,et al.Detrs beat yolos on real-time object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2024:16965-16974. [32]HOU X,LIU M,ZHANG S,et al.Salience detr:Enhancing detection transformer with hierarchical salience filtering refinement[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2024:17574-17583. [33]HUANG S,LU Z,CUN X,et al.Deim:Detr with improved matching for fast convergence[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2025:15162-15171. |
|
||