计算机科学 ›› 2026, Vol. 53 ›› Issue (8): 139-147.doi: 10.11896/jsjkx.251100082
张哲1,2, 刘俊杰1,2, 张柏礼1,2,3
ZHANG Zhe1,2, LIU Junjie1,2, ZHANG Baili1,2,3
摘要: 图元识别是建筑图纸智能化分析与合规性评审中的关键步骤。尽管DETR(Detection Transformer)框架在通用目标检测领域展现出强大潜力,但其在建筑CAD图元识别这种特定任务中仍面临一些问题。具体而言,DETR在进行CAD图元识别时主要受限于两方面:其一,其特征提取机制对图元尺度变化的适应性不足,难以同时有效表征微小结构细节与大跨度结构信息,导致对小尺度图元漏检或对大尺度图元识别模糊;其二,由于图元有效信息在图像上分布稀疏,其注意力机制难以高效聚焦于关键区域,易受大量无效背景干扰。针对上述问题,在DETR框架的基础上进行三方面的改进:采用基于Swin Transfor-mer的多尺度特征提取模块,增强模型对不同尺寸图元的鲁棒表征能力;引入空间稀疏采样注意力机制,优化特征图中有效信息的采样效率,提升关键区域的感知性能;结合对比去噪训练策略,强化模型对稀疏有效特征的定位能力。在实际建筑数据集上的实验表明,所提方法取得了0.675的mAP[0.5-0.95],显著优于现有主流目标检测模型,验证了其在建筑图元识别任务上的有效性与工程实用价值。
中图分类号:
| [1] GIRSHICK R.Fast r-cnn[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2015:1440-1448. [2] REN S,HE K,GIRSHICK R,et al.Faster R-CNN:Towards real-time object detection with region proposal networks[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2016,39(6):1137-1149. [3] HE K,GKIOXARI G,DOLLÁR P,et al.Mask r-cnn[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2017:2961-2969. [4] REDMON J,DIVVALA S,GIRSHICK R,et al.You only look once:Unified,real-time object detection[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2016:779-788. [5] WANG C Y,BOCHKOVSKIY A,LIAO H Y M.YOLOv7:Trainable bag-of-freebies sets new state-of-the-art for real-time object detectors[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2023:7464-7475. [6] YASEEN M.What is yolov9:An in-depth exploration of the internal features of the next-generation object detector[J].arXiv:2409.07813,2024. [7] WANG C Y,YEH I H,MARK LIAO H Y.Yolov9:Learning what you want to learn using programmable gradient information[C]//European Conference on Computer Vision(ECCV).Cham:Springer Nature Switzerland,2024:1-21. [8] GE Z,LIU S,WANG F,et al.Yolox:Exceeding yolo series in 2021[J].arXiv:2107.08430,2021. [9] CHEN Q,WANG Y,YANG T,et al.You only look one-level feature[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2021:13039-13048. [10] VASWANI A,SHAZEER N,PARMAR N,et al.Attention is all you need[C]//Proceedings of the 31st International Confe-rence on Neural Information Processing Systems.2017:6000-6010. [11] CARION N,MASSA F,SYNNAEVEG,et al.End-to-end object detection with transformers[C]//European Conference on Computer Vision(ECCV).Springer International Publishing,2020:213-229. [12] YUAN X.Research on the Recognition Method of Wall Symbols in Architectural Floor Plans[D].Harbin:Harbin Institute of Technology,2005. [13] HILAIRE X,TOMBRE K.Robust and accurate vectorization of line drawings[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2006,28(6):890-904. [14] UPADHYAY A,DUBEY A,KURIAKOSE S M.FPNet:Deep attention network for automated floor plan analysis[C]//International Conference on Document Analysis and Recognition.Cham:Springer Nature Switzerland,2023:163-176. [15] LIN T Y,GOYAL P,GIRSHICK R,et al.Focal loss for dense object detection[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2017:2980-2988. [16] 张小虎.建筑图纸构件识别模型构建方法、识别方法及相关设备:CN111476138B[P].2023-08-18. [17] YANG F,MU J,ZHANG Y,et al.Cadspotting:Robust panoptic symbol spotting on large-scale cad drawings[J].arXiv:2412.07377,2024. [18] FELZENSZWALB P F,GIRSHICK R B,MCALLESTER D,et al.Object detection with discriminatively trained part-based models[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2009,32(9):1627-1645. [19] GIRSHICK R,DONAHUE J,DARRELLT,et al.Rich feature hierarchies for accurate object detection and semantic segmentation[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2014:580-587. [20] WANG A,CHEN H,LIU L,et al.Yolov10:Real-time end-to-end object detection[J].Advances in Neural Information Processing Systems,2024,37:107984-108011. [21] KHANAM R,HUSSAIN M.Yolov11:An overview of the key architectural enhancements[J].arXiv:2410.17725,2024. [22] JIA D,YUAN Y,HE H,et al.Detrs with hybrid matching[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2023:19702-19712. [23] ZONG Z,SONG G,LIU Y.Detrs with collaborative hybrid assignments training[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2023:6748-6758. [24] LIU Z,LIN Y,CAOY,et al.Swin transformer:Hierarchical vision transformer using shifted windows[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2021:10012-10022. [25] DOSOVITSKIY A,BEYER L,KOLESNIKOV A,et al.Anlimage is worth 16x16 words:Transformers for image recognition at scale[J].arXiv:2010.11929,2020. [26] ZHANG H,LI F,LIU S,et al.DINO:DETR with Improved DeNoising Anchor Boxes for End-to-End Object Detection[C]//International Conference on Learning Representations(ICLR).2023. [27] KALERVO A,YLIOINAS J,HÄIKIÖ M,et al.CubiCasa5K:A Dataset and an Improved Multi-task Model for Floorplan Image Analysis[C]//Scandinavian Conference on Image Analysis.Springer,2019:28-40. [28] HE K,ZHANG X,REN S,et al.Deep residual learning forimage recognition[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition(CVPR).2016:770-778. [29] ZHU X,SU W,LU L,et al.Deformable DETR:DeformableTransformers for End-to-End Object Detection[C]//International Conference on Learning Representations(ICLR).2021. [30] MENG D,CHEN X,FAN Z,et al.Conditional detr for fast trai-ning convergence[C]//Proceedings of the IEEE International Conference on Computer Vision(ICCV).2021:3651-3660. |
|
||