计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250800006-7.doi: 10.11896/jsjkx.250800006

• 图像处理&多媒体技术 • 上一篇    下一篇

面向弱纹理工件的单目实时6D位姿估计方法

冯迎宾1, 康学仕1, 王天龙2   

  1. 1 沈阳理工大学自动化与电气工程学院 沈阳 110159
    2 中国科学院机器人与智能制造创新研究院 沈阳 110169
  • 出版日期:2026-06-16 发布日期:2026-06-12
  • 通讯作者: 冯迎宾(781787793@qq.com)
  • 基金资助:
    辽宁省科学技术计划项目(2023JH2/10700006)

Monocular Real-time 6D Pose Estimation for Weakly Textured Workpieces

FENG Yingbin1, KANG Xueshi1 , WANG Tianlong2   

  1. 1 School of Automation and Electrical Engineering,Shenyang Ligong University,Shenyang 110159,China
    2 Institute of Robotics and Intelligent Manufacturing Innovation,Chinese Academy of Sciences,Shenyang 110169,China
  • Published:2026-06-16 Online:2026-06-12
  • About author:FENG Yingbin,born in 1986,Ph.D,professor.His main research interests include robot environmental perception and modeling,and unmanned autonomous driving technology.
  • Supported by:
    Liaoning Province Science and Technology Program Project(2023JH2/10700006).

摘要: 针对工业场景中弱纹理工件尺寸不一、遮挡堆叠及光照变化等问题,提出了一种单目6D位姿估计方法RAAS-PVNet。为解决传统卷积结构在建模多尺度信息方面能力不足的问题,设计了分辨率自适应矩形卷积RARConv,动态调整卷积核大小和采样点数量,在PVNet主干网络中引入RARConv,提高了模型的多尺度能力;提出角距协同加权投票策略AS,引入方向向量延长线的垂直距离约束,结合连续权重融合机制,准确衡量每个投票点的可信度,使投票结果聚焦于高质量点,提高了模型的抗遮挡能力;面对位姿估计领域工业零件数据集匮乏的问题,设计了一种真实数据和合成数据按比例结合的数据集制作方法,构建工件数据集6DInd。实验表明,RAAS-PVNet在6DInd上的2D Projection和ADD(-S)分别提升10.22%和10.26%,在遮挡及光照变化下均具有良好的鲁棒性,30fps的处理速度满足实时性需求。

关键词: 6D位姿估计, RAAS-PVNet, 分辨率自适应矩形卷积, 角距协同加权投票策略, 6DInd

Abstract: Aiming at the problems of different sizes,occlusion stacking and lighting changes of weakly textured workpieces in industrial scenes,a monocular 6D pose estimation method RAAS-PVNet is proposed.The design resolution adaptive rectangular convolutional RARConv dynamically adjusts the size of the convolutional kernel and the number of sampling points,which solves the problem of the insufficient ability of the traditional convolutional structure in modeling multi-scale information.The angular distance collaborative weighted voting strategy AS is proposed,the vertical distance constraint of the direction vector extension line is introduced,and the credibility of each voting point is accurately measured by combining the continuous weight fusionmecha-nism,so that the voting results are focused on high-quality points and the anti-occlusion ability of the model is improved.Faced with the problem of lack of industrial part datasets in the field of pose estimation,a dataset production method combining real data and synthetic data in proportion is designed to construct the workpiece dataset 6DInd.Experiments show that the 2D Projection and ADD(-S) of RAAS-PVNet on 6DInd are increased by 10.22% and 10.26%,respectively,and have good robustness under occlusion and lighting changes,and the processing speed of 30 fps meets the real-time requirements.

Key words: 6D pose estimation, RAAS-PVNet, Resolution adaptive rectangular convolution, Angular spacing collaborative weighted voting strategy, 6DInd

中图分类号: 

  • TP183
[1] WEN B,YANG W,KAUTZ J,et al.Foundationpose:Unified 6d pose estimation and tracking of novel objects[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2024:17868-17879.
[2] WAN Q,NING S X,ZHONG H,et al.6D pose estimation and robotic arm grasping method for weakly rextured workpiece[J].Control Theory & Applications,2025,42(7):1443-1452.
[3] LIU J,SUN W,ZENG K,et al.Novel object 6d pose estimation with a single reference view[J].arXiv:2503.05578,2025.
[4] JIN L,ZHOU G,LIU Z,et al.IRPE:Instance-level reconstruction-based 6D pose estimator[J].Image and Vision Computing,2025,154:105340.
[5] WANG S L,YONGY,WU C R.6D Pose Estimation of Low Texture Industrial Parts Based on Pseudo-Siamese Neural Network[J].Acta Electronica Sinica,2023,51(1):192-201.
[6] FAN Z,ZHU Y,HE Y,et al.Deep learning on monocular object pose detection and tracking:A comprehensive overview[J].ACM Computing Surveys,2022,55(4):1-40.
[7] LABBÉ Y,CARPENTIER J,AUBRY M,et al.Cosypose:Con-sistent multi-view multi-object 6d pose estimation[C]//Computer Vision-ECCV 2020:16th European Conference,Part XVII 16.Springer International Publishing,2020:574-591.
[8] LI Z,STAMOS I.Depth-based 6dof object pose estimation using swin transformer[C]//IEEE/RSJ International Conference on Intelligent Robots and Systems(IROS 2023).IEEE,2023:1185-1191.
[9] TEKIN B,SINHA S N,FUA P.Real-time seamless single shot6d object pose prediction[C]//Proceedings of the IEEEConfe-rence on Computer Vision and Pattern Reconition.2018:292-301.
[10] WANG G,MANHARDT F,TOMBARI F,et al.Gdr-net:Ge-ometry-guided direct regression network for monocular 6d object pose estimation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2021:16611-16621.
[11] CAO T,ZHANG W,FU Y,et al.Dgecn++:A depth-guided edge convolutional network for end-to-end 6d pose estimation via attention mechanism[J].IEEE Transactions on Circuits and Systems for Video Technology,2023,34(6):4214-4228.
[12] PENG S,LIU Y,HUANG Q,et al.Pvnet:Pixel-wise votingnetwork for 6dof pose estimation[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2019:4561-4570.
[13] LEPETIT V,MORENO-NOGUER F,FUA P.EPnP:An accurate O(n) solution to the P n P problem[J].International Journal of Computer Vision,2009,81:155-166.
[14] WANG X,ZHENG Z,SHAO J,et al.Adaptive RectangularConvolution for Remote Sensing Pansharpening[J].arXiv:2503.00467,2025.
[15] DENNINGER M,SUNDERMEYER M,WINKELBAUER D,et al.Blenderproc:Reducing the reality gap with photorealistic rendering[C]//16th Robotics:Science and Systems(RSS 2020).Workshops.2020.
[16] XIAO J,HAYS J,EHINGER K A,et al.Sun database:Large-scale scene recognition from abbey to zoo[C]//2010 IEEE Computer Society Conference on Computer Vision and Pattern Re-cognition.IEEE,2010:3485-3492.
[17] SONG C,SONG J,HUANG Q.Hybridpose:6d object pose estimation under hybrid representations[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2020:431-440.
[18] IWASE S,LIU X,KHIRODKAR R,et al.Repose:Fast 6d object pose refinement via deep texture rendering[C]//Procee-dings of the IEEE/CVF International Conference on Computer Vision.2021:3303-3312.
[19] LU Y,PEIS.DFW-PVNet:data field weighting based pixel-wise voting network for effective 6D pose estimation[J].Applied Intelligence,2025,55(4):240.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!