计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250700088-7.doi: 10.11896/jsjkx.250700088

• 图像处理&多媒体技术 • 上一篇    下一篇

基于分阶段训练策略与多尺度特征融合的目标检测方法

瞿洁武1, 路新喜2, 孙健1, 柳燕1, 高玲1, 徐彬彬1   

  1. 1 北京智网数科技术有限公司 北京 102200
    2 北京航空航天大学软件学院 北京 100191
  • 出版日期:2026-06-16 发布日期:2026-06-12
  • 通讯作者: 路新喜(lxx@buaa.edu.cn)
  • 作者简介:(ford.qu@163.com)

Object Detection Method Based on Phased Training Strategy and Multi-scale Feature Fusion

QU Jiewu1, LU Xinxi2, SUN Jian1, LIU Yan1, GAO Ling1, XU Binbin1   

  1. 1 Pipechina Digital Co.,Ltd.,Beijing 102200,China
    2 School of Software,Beihang University,Beijing 100191,China
  • Published:2026-06-16 Online:2026-06-12
  • About author:QU Jiewu,born in 1985.His main research interests include network communication and artificial intelligence.
    LU Xinxi,born in 1978,Ph.D,associate professor,master's supervisor.His main research interests include intelligent software and artificial intelligence.

摘要: 针对DETR(Detection Transformer)系列目标检测方法在推理阶段存在计算瓶颈、难以兼顾实时性与精度的问题,提出了一种基于分阶段训练策略与多尺度特征融合的改进方法。具体而言,通过简化DETR的多层编码器结构降低计算复杂度,并采用分阶段训练策略提升特征表达能力和模型收敛速度。第一阶段采用一对多标签匹配获取高质量二维多尺度特征,第二阶段冻结第一阶段的网络权重,并引入并行注意力卷积融合模块进一步细化特征。实验结果表明,所提方法在COCO数据集上较基线模型实现了5倍的推理速度提升,并带来了1.5个百分点的AP增益,有效缓解了DETR在推理阶段效率低下的问题;在BitVehicle数据集上也取得了1.4个百分点的AP提升。

关键词: 目标检测, 并行注意力卷积, 分阶段训练策略, 多尺度特征融合

Abstract: To overcome these limitations,such as the computational bottlenecks and the challenge of balancing real-time perfor-mance with accuracy in the DETR(Detection Transformer) family of object detection methods during inference,this paper proposes an enhanced approach that combines a phased training strategy with multi-scale feature fusion.Specifically,the multi-layer encoder structure of DETR is simplified to reduce computational complexity,while the phased training strategy improves feature representation and accelerates model convergence.In the first phase,one-to-many label matching is adopted to obtain high-quality two-dimensional multi-scale features.In the second phase,the weights from the first phase are frozen,and a parallel attention-convolutional fusion module is introduced to further refine the features.Experimental results demonstrate that the proposed method achieves a 5× increase in inference speed and a 1.5-point AP gain over the baseline model on the COCO dataset,effectively alleviating DETR's inference inefficiency.In addition,it yields a 1.4-point AP improvement on the BitVehicle dataset.

Key words: Object detection, Parallel attention convolution, Phased training strategy, Multi-scale feature fusion

中图分类号: 

  • TP311
[1] CARION N,MASSA F,SYNNAEVE G,et al.End-to-end object detection with transformers[C]//European Conference on Computer Vision.Cham:Springer International Publishing,2020:213-229.
[2] ZHU X,SU W,LU L,et al.Deformable detr:Deformable transformers for end-to-end object detection[J].arXiv:2010.04159,2020.
[3] LIU S,LI F,ZHANG H,et al.Dab-detr:Dynamic anchor boxes are better queries for detr[J].arXiv:2201.12329,2022.
[4] ZHANG H,LI F,LIU S,et al.Dino:Detr with improved denoi-sing anchor boxes for end-to-end object detection[J].arXiv:2203.03605,2022.
[5] LI F,ZENG A,LIU S,et al.Lite detr:An interleaved multi-scale encoder for efficient detr[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:18558-18567.
[6] REN S,HE K,GIRSHICK R,et al.Faster R-CNN:Towards real-time object detection with region proposal networks[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2016,39(6):1137-1149.
[7] HE K,GKIOXARI G,DOLLÁR P,et al.Mask r-cnn[C]//Proceedings of the IEEE International Conference on Computer Vision.2017:2961-2969.
[8] CAI Z,VASCONCELOS N.Cascade r-cnn:Delving into highquality object detection[C]//Proceedings of the IEEE Confe-rence on Computer Vision and Pattern Recognition.2018:6154-6162.
[9] WANG C Y,BOCHKOVSKIY A,LIAO H Y M.YOLOv7:Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors[C]//Proceedings of the IEEE/CVF Confe-rence on Computer Vision and Pattern Recognition.2023:7464-7475.
[10] LIN T Y,GOYAL P,GIRSHICK R,et al.Focal loss for dense object detection[C]//Proceedings of the IEEE International Conference on Computer Vision.2017:2980-2988.
[11] WU J,ZHAO C.Small Object Detection Method Based on Improved DETR Algorithm[J/OL].Computer Applications,2025.
[12] MENG D,CHEN X,FAN Z,et al.Conditional detr for fasttraining convergence[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2021:3651-3660.
[13] YAO Z,AI J,LI B,et al.Efficient detr:improving end-to-endobject detector with dense prior[J].arXiv:2104.01318,2021.
[14] WANG Y,ZHANG X,YANG T,et al.Anchor detr:Query design for transformer-based detector[C]//Proceedings of the AAAI Conference on Artificial Intelligence.2022:2567-2575.
[15] ZHANG D P,WEI Y Y,HE S J,et al.Feature Fusion and Inter-layer Transmission:An Improved Object Detection Method Based on Anchor DETR[J].Journal of Graphics,2024,45(5):968-978.
[16] LI F,ZHANG H,LIU S,et al.Dn-detr:Accelerate detr training by introducing query denoising[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:13619-13627.
[17] ZHENG D,DONG W,HU H,et al.Less is more:Focus attention for efficient detr[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2023:6674-6683.
[18] ZHAO Y,LV W,XU S,et al.Detrs beat yolos on real-time object detection[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2024:16965-16974.
[19] GUAN Y,LIAO S,YANG W.AParC-DETR:accelerate DETR training by introducing adaptive position-aware circular convolution[J].The Visual Computer,2025,41(2):1319-1333.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!