Computer Science ›› 2026, Vol. 53 ›› Issue (8): 139-147.doi: 10.11896/jsjkx.251100082

• Computer Graphics & Multimedia • Previous Articles     Next Articles

Improved DETR-based Method for CAD Graphic Element Recognition

ZHANG Zhe1,2, LIU Junjie1,2, ZHANG Baili1,2,3   

  1. 1 School of Computer Science and Engineering, Southeast University, Nanjing 211189, China
    2 Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications(Southeast University), Ministry of Education, Nanjing 211189, China
    3 Research Center for Judicial Big Data, Supreme Count of China, Nanjing 211189, China
  • Received:2025-11-17 Revised:2026-04-22 Online:2026-08-15 Published:2026-08-17
  • About author:ZHANG Zhe,born in 2001,postgra-duate.His main research interests include computer vision,object detection and artificial intelligence.
    ZHANG Baili,born in 1970,Ph.D,professor,Ph.D supervisor.His main research interests include big data analysis,artificial intelligence and data warehouse.
  • Supported by:
    National Key R & D Program of China(2023YFC3806004),National Natural Science Foundation of China(62373104),Hainan Provincial Science and Technology Special Project(ZDYF2023GXJS150) and Fundamental Research Funds for the Central Universities(2242024k30035).

Abstract: Graphic element recognition constitutes a critical step in the intelligent analysis and compliance review of architectural drawings.Although the DETR(Detection Transformer) framework demonstrates significant potential in general object detection,notable challenges persist when applied to complex CAD graphic element recognition tasks.Specifically,DETR's limitations primarily stem from two aspects:1) Its feature extraction mechanism exhibits insufficient adaptability to scale variations among graphic elements,failing to simultaneously capture fine details of small components and large-scale structural information effectively,which results in missed detections of small elements and ambiguous boundary recognition for large ones;2) The sparse spatial distribution of informative pixels within graphic elements hinders efficient attention focusing on key regions,making the mechanism susceptible to substantial background interference.To address these issues,three key improvements to the DETR framework are proposed.A Swin Transformer based multi-scale feature extraction module is adopted to enhance robust representational capability for elements of varying sizes.A spatially sparse sampling attention mechanism is introduced to optimize the sampling efficiency of informative features and improve perception performance in critical areas.A contrastive denoising training strategy is integrated to strengthen model localization capability for sparse informative features.Experimental results on a real-world architectural dataset demonstrate that the proposed method achieves a mAP[0.5-0.95] of 0.675,significantly outperforming existing mainstream object detection models.This validates the proposed method's effectiveness and practical engineering value for architectural graphic element recognition task.

Key words: Graphic element recognition, Computer vision, Object detection, Computer-Aided Design(CAD)

CLC Number: 

  • TP391
[1] GIRSHICK R.Fast r-cnn[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2015:1440-1448.
[2] REN S,HE K,GIRSHICK R,et al.Faster R-CNN:Towards real-time object detection with region proposal networks[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2016,39(6):1137-1149.
[3] HE K,GKIOXARI G,DOLLÁR P,et al.Mask r-cnn[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2017:2961-2969.
[4] REDMON J,DIVVALA S,GIRSHICK R,et al.You only look once:Unified,real-time object detection[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2016:779-788.
[5] WANG C Y,BOCHKOVSKIY A,LIAO H Y M.YOLOv7:Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2023:7464-7475.
[6] YASEEN M.What is yolov9:An in-depth exploration of the internal features of the next-generation object detector[J].arXiv:2409.07813,2024.
[7] WANG C Y,YEH I H,MARK LIAO H Y.Yolov9:Learning what you want to learn using programmable gradient information[C]//European Conference on Computer Vision(ECCV).Cham:Springer Nature Switzerland,2024:1-21.
[8] GE Z,LIU S,WANG F,et al.Yolox:Exceeding yolo series in 2021[J].arXiv:2107.08430,2021.
[9] CHEN Q,WANG Y,YANG T,et al.You only look one-level feature[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2021:13039-13048.
[10] VASWANI A,SHAZEER N,PARMAR N,et al.Attention is all you need[C]//Proceedings of the 31st International Confe-rence on Neural Information Processing Systems.2017:6000-6010.
[11] CARION N,MASSA F,SYNNAEVEG,et al.End-to-end object detection with transformers[C]//European Conference on Computer Vision(ECCV).Springer International Publishing,2020:213-229.
[12] YUAN X.Research on the Recognition Method of Wall Symbols in Architectural Floor Plans[D].Harbin:Harbin Institute of Technology,2005.
[13] HILAIRE X,TOMBRE K.Robust and accurate vectorization of line drawings[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2006,28(6):890-904.
[14] UPADHYAY A,DUBEY A,KURIAKOSE S M.FPNet:Deep attention network for automated floor plan analysis[C]//International Conference on Document Analysis and Recognition.Cham:Springer Nature Switzerland,2023:163-176.
[15] LIN T Y,GOYAL P,GIRSHICK R,et al.Focal loss for dense object detection[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2017:2980-2988.
[16] 张小虎.建筑图纸构件识别模型构建方法、识别方法及相关设备:CN111476138B[P].2023-08-18.
[17] YANG F,MU J,ZHANG Y,et al.Cadspotting:Robust panoptic symbol spotting on large-scale cad drawings[J].arXiv:2412.07377,2024.
[18] FELZENSZWALB P F,GIRSHICK R B,MCALLESTER D,et al.Object detection with discriminatively trained part-based models[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2009,32(9):1627-1645.
[19] GIRSHICK R,DONAHUE J,DARRELLT,et al.Rich feature hierarchies for accurate object detection and semantic segmentation[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2014:580-587.
[20] WANG A,CHEN H,LIU L,et al.Yolov10:Real-time end-to-end object detection[J].Advances in Neural Information Processing Systems,2024,37:107984-108011.
[21] KHANAM R,HUSSAIN M.Yolov11:An overview of the key architectural enhancements[J].arXiv:2410.17725,2024.
[22] JIA D,YUAN Y,HE H,et al.Detrs with hybrid matching[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2023:19702-19712.
[23] ZONG Z,SONG G,LIU Y.Detrs with collaborative hybrid assignments training[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2023:6748-6758.
[24] LIU Z,LIN Y,CAOY,et al.Swin transformer:Hierarchical vision transformer using shifted windows[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2021:10012-10022.
[25] DOSOVITSKIY A,BEYER L,KOLESNIKOV A,et al.Anlimage is worth 16x16 words:Transformers for image recognition at scale[J].arXiv:2010.11929,2020.
[26] ZHANG H,LI F,LIU S,et al.DINO:DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection[C]//International Conference on Learning Representations(ICLR).2023.
[27] KALERVO A,YLIOINAS J,HÄIKIÖ M,et al.CubiCasa5K:A Dataset and an Improved Multi-task Model for Floorplan Image Analysis[C]//Scandinavian Conference on Image Analysis.Springer,2019:28-40.
[28] HE K,ZHANG X,REN S,et al.Deep residual learning forimage recognition[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2016:770-778.
[29] ZHU X,SU W,LU L,et al.Deformable DETR:DeformableTransformers for End-to-End Object Detection[C]//International Conference on Learning Representations(ICLR).2021.
[30] MENG D,CHEN X,FAN Z,et al.Conditional detr for fast trai-ning convergence[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2021:3651-3660.
[1] WANG Jiahui, WANG Hongyu, HAO Yingguang. Feature Aggregation with Joint Tracking:Video Object Detection in Occlusion Scenarios [J]. Computer Science, 2026, 53(8): 165-173.
[2] ZHU Yifei, LIU Tianpeng, SUN Tengzhong, LI Yanchen, CHEN Zhihong, FANG Pengfei. Survey of Hyperbolic Geometry in Computer Vision [J]. Computer Science, 2026, 53(7): 9-23.
[3] ZHANG Shouyi, SHEN Qiang, GUO Yiran, WANG Hanyu. Rain and Fog Weather Object Detection Algorithm Based on Improved YOLOv8 Model [J]. Computer Science, 2026, 53(6A): 250300090-7.
[4] LIU Dai, AN Pengyu, WANG Kai. Improved YOLOv5s-based Algorithm for Emergency Situation Detection in Airport Terminals [J]. Computer Science, 2026, 53(6A): 250300174-7.
[5] MAO Lihong, TANG Jianjun, CHEN Tong, ZHANG Rui. Aerial Image Object Detection Model Based on Dual-domain Attention and Feature Fusion [J]. Computer Science, 2026, 53(6A): 250600036-7.
[6] HUANG Haixin, HOU Guangshuai, HE Tianyu. SeguGAN:Research on Super-resolution Reconstruction of License Plate Images UtilizingGenerative Adversarial Networks [J]. Computer Science, 2026, 53(6A): 250600070-5.
[7] HUANG Haixin, HE Tianyu, HOU Guangshuai. Multi-layer Graph Convolutional Action Recognition Method Based on Topological Information [J]. Computer Science, 2026, 53(6A): 250600147-5.
[8] SHAN Chengcheng, MEI Chun, LI Weiting, GUO Yuanyuan, QIAN Weixing, XIONG Zhi. Semantic Perception Active Learning Method for the Datum Map of Scene Matching Navigation System [J]. Computer Science, 2026, 53(6A): 250600228-8.
[9] CHEN Nuo, ZHAO Peng, HUAN Haisheng. Review of Small Object Detection Based on Deep Learning [J]. Computer Science, 2026, 53(6A): 250700022-9.
[10] ZHENG Haibin, LIN Xiuhao, HAN Ye, CHEN Jinyin, LI Beibei. Black-box Physical Adversarial Attack Against Multimodal Object Detector [J]. Computer Science, 2026, 53(6A): 250700023-10.
[11] QU Jiewu, LU Xinxi, SUN Jian, LIU Yan, GAO Ling, XU Binbin. Object Detection Method Based on Phased Training Strategy and Multi-scale Feature Fusion [J]. Computer Science, 2026, 53(6A): 250700088-7.
[12] DONG Ye, LIAN Xinyue, WANG Yuyang, OU Xinyu. RGB-IR Multi-modal Fusion-based Tomato Small Object Detection [J]. Computer Science, 2026, 53(6A): 250700173-8.
[13] ZHOU Wenwu, LEI Lei, XUAN Xin. Armory Equipment Detection Based on Improved YOLOv5 [J]. Computer Science, 2026, 53(6A): 250800049-6.
[14] JI Wenyu, LI Yang, WANG Jiabao, FU Ruizhi, LIU Xiaoyu, MIAO Zhuang. Review of 3D Object Detection Based on LiDAR-camera Fusion [J]. Computer Science, 2026, 53(6): 214-231.
[15] LI Peng, ZHANG Zihao, HAN Yahong. Primitive Dynamic Weighting for Multi-modal Salient Object Detection [J]. Computer Science, 2026, 53(6): 242-251.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!