计算机科学 ›› 2026, Vol. 53 ›› Issue (8): 139-147.doi: 10.11896/jsjkx.251100082

• 计算机图形学 & 多媒体 • 上一篇    下一篇

基于DETR改进的CAD图元识别方法

张哲1,2, 刘俊杰1,2, 张柏礼1,2,3   

  1. 1 东南大学计算机科学与工程学院 南京 211189
    2 新一代人工智能技术与交叉应用教育部重点实验室(东南大学) 南京 211189
    3 中国最高人民法院司法大数据研究基地 南京 211189
  • 收稿日期:2025-11-17 修回日期:2026-04-22 出版日期:2026-08-15 发布日期:2026-08-17
  • 通讯作者: 张柏礼(zhangbl@seu.edu.cn)
  • 作者简介:(zhangz_seu@163.com)
  • 基金资助:
    国家重点研发计划(2023YFC3806004);国家自然科学基金(62373104);海南省科技专项(ZDYF2023GXJS150);中央高校基本科研业务费专项资金(2242024k30035)

Improved DETR-based Method for CAD Graphic Element Recognition

ZHANG Zhe1,2, LIU Junjie1,2, ZHANG Baili1,2,3   

  1. 1 School of Computer Science and Engineering, Southeast University, Nanjing 211189, China
    2 Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications(Southeast University), Ministry of Education, Nanjing 211189, China
    3 Research Center for Judicial Big Data, Supreme Count of China, Nanjing 211189, China
  • Received:2025-11-17 Revised:2026-04-22 Published:2026-08-15 Online:2026-08-17
  • About author:ZHANG Zhe,born in 2001,postgra-duate.His main research interests include computer vision,object detection and artificial intelligence.
    ZHANG Baili,born in 1970,Ph.D,professor,Ph.D supervisor.His main research interests include big data analysis,artificial intelligence and data warehouse.
  • Supported by:
    National Key R & D Program of China(2023YFC3806004),National Natural Science Foundation of China(62373104),Hainan Provincial Science and Technology Special Project(ZDYF2023GXJS150) and Fundamental Research Funds for the Central Universities(2242024k30035).

摘要: 图元识别是建筑图纸智能化分析与合规性评审中的关键步骤。尽管DETR(Detection Transformer)框架在通用目标检测领域展现出强大潜力,但其在建筑CAD图元识别这种特定任务中仍面临一些问题。具体而言,DETR在进行CAD图元识别时主要受限于两方面:其一,其特征提取机制对图元尺度变化的适应性不足,难以同时有效表征微小结构细节与大跨度结构信息,导致对小尺度图元漏检或对大尺度图元识别模糊;其二,由于图元有效信息在图像上分布稀疏,其注意力机制难以高效聚焦于关键区域,易受大量无效背景干扰。针对上述问题,在DETR框架的基础上进行三方面的改进:采用基于Swin Transfor-mer的多尺度特征提取模块,增强模型对不同尺寸图元的鲁棒表征能力;引入空间稀疏采样注意力机制,优化特征图中有效信息的采样效率,提升关键区域的感知性能;结合对比去噪训练策略,强化模型对稀疏有效特征的定位能力。在实际建筑数据集上的实验表明,所提方法取得了0.675的mAP[0.5-0.95],显著优于现有主流目标检测模型,验证了其在建筑图元识别任务上的有效性与工程实用价值。

关键词: 图元识别, 计算机视觉, 目标检测, 计算机辅助设计

Abstract: Graphic element recognition constitutes a critical step in the intelligent analysis and compliance review of architectural drawings.Although the DETR(Detection Transformer) framework demonstrates significant potential in general object detection,notable challenges persist when applied to complex CAD graphic element recognition tasks.Specifically,DETR's limitations primarily stem from two aspects:1) Its feature extraction mechanism exhibits insufficient adaptability to scale variations among graphic elements,failing to simultaneously capture fine details of small components and large-scale structural information effectively,which results in missed detections of small elements and ambiguous boundary recognition for large ones;2) The sparse spatial distribution of informative pixels within graphic elements hinders efficient attention focusing on key regions,making the mechanism susceptible to substantial background interference.To address these issues,three key improvements to the DETR framework are proposed.A Swin Transformer based multi-scale feature extraction module is adopted to enhance robust representational capability for elements of varying sizes.A spatially sparse sampling attention mechanism is introduced to optimize the sampling efficiency of informative features and improve perception performance in critical areas.A contrastive denoising training strategy is integrated to strengthen model localization capability for sparse informative features.Experimental results on a real-world architectural dataset demonstrate that the proposed method achieves a mAP[0.5-0.95] of 0.675,significantly outperforming existing mainstream object detection models.This validates the proposed method's effectiveness and practical engineering value for architectural graphic element recognition task.

Key words: Graphic element recognition, Computer vision, Object detection, Computer-Aided Design(CAD)

中图分类号: 

  • TP391
[1] GIRSHICK R.Fast r-cnn[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2015:1440-1448.
[2] REN S,HE K,GIRSHICK R,et al.Faster R-CNN:Towards real-time object detection with region proposal networks[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2016,39(6):1137-1149.
[3] HE K,GKIOXARI G,DOLLÁR P,et al.Mask r-cnn[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2017:2961-2969.
[4] REDMON J,DIVVALA S,GIRSHICK R,et al.You only look once:Unified,real-time object detection[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2016:779-788.
[5] WANG C Y,BOCHKOVSKIY A,LIAO H Y M.YOLOv7:Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2023:7464-7475.
[6] YASEEN M.What is yolov9:An in-depth exploration of the internal features of the next-generation object detector[J].arXiv:2409.07813,2024.
[7] WANG C Y,YEH I H,MARK LIAO H Y.Yolov9:Learning what you want to learn using programmable gradient information[C]//European Conference on Computer Vision(ECCV).Cham:Springer Nature Switzerland,2024:1-21.
[8] GE Z,LIU S,WANG F,et al.Yolox:Exceeding yolo series in 2021[J].arXiv:2107.08430,2021.
[9] CHEN Q,WANG Y,YANG T,et al.You only look one-level feature[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2021:13039-13048.
[10] VASWANI A,SHAZEER N,PARMAR N,et al.Attention is all you need[C]//Proceedings of the 31st International Confe-rence on Neural Information Processing Systems.2017:6000-6010.
[11] CARION N,MASSA F,SYNNAEVEG,et al.End-to-end object detection with transformers[C]//European Conference on Computer Vision(ECCV).Springer International Publishing,2020:213-229.
[12] YUAN X.Research on the Recognition Method of Wall Symbols in Architectural Floor Plans[D].Harbin:Harbin Institute of Technology,2005.
[13] HILAIRE X,TOMBRE K.Robust and accurate vectorization of line drawings[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2006,28(6):890-904.
[14] UPADHYAY A,DUBEY A,KURIAKOSE S M.FPNet:Deep attention network for automated floor plan analysis[C]//International Conference on Document Analysis and Recognition.Cham:Springer Nature Switzerland,2023:163-176.
[15] LIN T Y,GOYAL P,GIRSHICK R,et al.Focal loss for dense object detection[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2017:2980-2988.
[16] 张小虎.建筑图纸构件识别模型构建方法、识别方法及相关设备:CN111476138B[P].2023-08-18.
[17] YANG F,MU J,ZHANG Y,et al.Cadspotting:Robust panoptic symbol spotting on large-scale cad drawings[J].arXiv:2412.07377,2024.
[18] FELZENSZWALB P F,GIRSHICK R B,MCALLESTER D,et al.Object detection with discriminatively trained part-based models[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2009,32(9):1627-1645.
[19] GIRSHICK R,DONAHUE J,DARRELLT,et al.Rich feature hierarchies for accurate object detection and semantic segmentation[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2014:580-587.
[20] WANG A,CHEN H,LIU L,et al.Yolov10:Real-time end-to-end object detection[J].Advances in Neural Information Processing Systems,2024,37:107984-108011.
[21] KHANAM R,HUSSAIN M.Yolov11:An overview of the key architectural enhancements[J].arXiv:2410.17725,2024.
[22] JIA D,YUAN Y,HE H,et al.Detrs with hybrid matching[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2023:19702-19712.
[23] ZONG Z,SONG G,LIU Y.Detrs with collaborative hybrid assignments training[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2023:6748-6758.
[24] LIU Z,LIN Y,CAOY,et al.Swin transformer:Hierarchical vision transformer using shifted windows[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2021:10012-10022.
[25] DOSOVITSKIY A,BEYER L,KOLESNIKOV A,et al.Anlimage is worth 16x16 words:Transformers for image recognition at scale[J].arXiv:2010.11929,2020.
[26] ZHANG H,LI F,LIU S,et al.DINO:DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection[C]//International Conference on Learning Representations(ICLR).2023.
[27] KALERVO A,YLIOINAS J,HÄIKIÖ M,et al.CubiCasa5K:A Dataset and an Improved Multi-task Model for Floorplan Image Analysis[C]//Scandinavian Conference on Image Analysis.Springer,2019:28-40.
[28] HE K,ZHANG X,REN S,et al.Deep residual learning forimage recognition[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2016:770-778.
[29] ZHU X,SU W,LU L,et al.Deformable DETR:DeformableTransformers for End-to-End Object Detection[C]//International Conference on Learning Representations(ICLR).2021.
[30] MENG D,CHEN X,FAN Z,et al.Conditional detr for fast trai-ning convergence[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2021:3651-3660.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!