Computer Science ›› 2026, Vol. 53 ›› Issue (9): 209-218.doi: 10.11896/jsjkx.260500015

• Computer Graphics & Multimedia • Previous Articles     Next Articles

Six Degrees of Freedom Pose Estimation Method Fusing Spectral and Geometric Features

HE Yuhuang1,3, WANG Qianlei1,3 DAI Yuntong1,3 DU Wu2,4, QIN Xiaolin1,3   

  1. 1 Chengdu Institute of Computer Application,Chinese Academy of Sciences,Chengdu 610213,China
    2 School of Computer Science,Sichuan University,Chengdu 610065,China
    3 School of Computer Science and Technology,University of Chinese Academy of Sciences,Beijing 100049,China
    4 Chengdu Sanshi Technology Company Limited,Chengdu 610052,China
  • Received:2026-05-06 Revised:2026-07-05 Online:2026-09-15 Published:2026-09-10
  • About author:HE Yuhuang,born in 2001,postgra-duate.His main research interests include pose estimation and computer vision.
    QIN Xiaolin,born in 1980,Ph.D,researcher,is a senior member of CCF(No.12344M).His main research interests include automated reasoning and spatial intelligence,industry-specific vertical large models.
  • Supported by:
    Sichuan Science and Technology Program(2025JDDQ0008,2024NSFJQ0035).

Abstract: To solve the problems of insufficient feature and pose ambiguity caused by texture-less and symmetric objects in complex scenes during 6D pose estimation,a pose estimation network named SGF-Pose integrating frequency-domain feature enhancement and geometric enhancement is proposed.Firstly,a spectral feature enhancement module(SFEM) is introduced in the feature extraction stage.By learning a spectral mask,background high-frequency noise is adaptively suppressed to improve the robustness of the predictions.Furthermore,a geometry enhancement module(GEM) is designed to explicitly measure pixel-level geometric curvature by calculating the spatial gradients of coordinate mappings.Finally,the 6D pose is predicted through a differentiable PnP solver module by combining dense prediction maps and geometric confidence.Experimental results show that,using the ADD-(S) metric,SGF-Pose achieves an accuracy of 96.9% on the LineMOD dataset,which is 1.7 percentage points higher than that of DPOD.On the Occlusion LineMOD dataset,an accuracy of 59.5% is achieved,representing an improvement of 3.4 percentage points over GDR-Net.Using the AR metric,an accuracy of 66.4% is achieved by SGF-Pose on the T-LESS dataset,which is 2.4 percentage points higher than that of CosyPose.These results indicate that SGF-Pose performs well in pose estimation for weak-texture and symmetric scenes.

Key words: Pose estimation, Frequency domain analysis, Feature fusion, Texture-less object, Deep learning

CLC Number: 

  • TP391.41
[1] WANG Y,XIE J,CHENG J,et al.Review of object pose estimation in RGB images based on deep learning[J].Journal of Computer Applications,2023,43(8):2546-2555.
[2] GUO N,LI J Y,REN X.Survey of rigid object pose estimation algorithms based on deep learning[J].Computer Science,2023,50(2):178-189.
[3] DORUK A E,OZKAYA T E,GÜLMEZ F,et al.A comparative study for 6D pose estimation of textureless and symmetric objects used in automotive manufacturing industry[C] //2023 5th International Congress on Human-Computer Interaction,Optimization and Robotic Applications.Piscataway:IEEE,2023:1-7.
[4] WANG Y N,JIANG Y M,JIANG J,et al.Key technologies of robot perception and control and its intelligent manufacturing applications[J].Acta Automatica Sinica,2023,49(3):494-513.
[5] GUAN J,HAO Y M,WU Q X,et al.A survey of 6DoF object pose estimation methods for different application scenarios[J].Sensors,2024,24(4):1076.
[6] DI Y,MANHARDT F,WANG G,et al.SO-Pose:exploitingself-occlusion for direct 6D pose estimation[C] //2021 IEEE/CVF International Conference on Computer Vision.Piscataway:IEEE,2021:12376-12385.
[7] WANG G,MANHARDT F,TOMBARI F,et al.GDR-Net:geometry-guided direct regression network for monocular 6D object pose estimation[C] //2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2021:16606-16616.
[8] PENG S D,ZHOU X W,LIU Y,et al.PVNet:pixel-wise voting network for 6DoF object pose estimation[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2022,44(6):3212-3223.
[9] LIAN R Y,LING H B.CheckerPose:progressive dense keypoint localization for object pose estimation with graph neural network[C] //2023 IEEE/CVF International Conference on Computer Vision.Piscataway:IEEE,2023:13976-13987.
[10] XU L,QU H X,CAI Y J,et al.6D-Diff:a keypoint diffusion framework for 6D object pose estimation[C] //2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2024:9676-9686.
[11] YU L X,CHEN Y B,QU H J,et al.Multi-stage grasping method for unordered mixed objects grasping based on GraspNet[J].Computer Science,2026,53(4):318-325.
[12] XU M Z,SHI X H.6D pose estimation based on multi feature fusion[J].Computer Applications and Software,2025,42(6):193-201.
[13] HODAN T,HALUZA P,OBDRZALEK S,et al.T-LESS:anRGB-D dataset for 6D pose estimation of texture-less objects[J].arXiv:1701.05498,2017.
[14] ZHAO H,WEI S X,SHI D H,et al.Learning symmetry-aware geometry correspondences for 6D object pose estimation[C] //2023 IEEE/CVF International Conference on Computer Vision.Piscataway:IEEE,2023:13999-14008.
[15] BAY H,TUYTELAARS T,VAN G L.SURF:speeded up robust features[C] //Proceedings of the 9th European Conference on Computer Vision.Berlin:Springer,2006:404-417.
[16] CALONDER M,LEPETIT V,STRECHA C,et al.BRIEF:binary robust independent elementary features[C] //Proceedings of the 11th European Conference on Computer Vision.Berlin:Springer,2010:778-792.
[17] RUBLEE E,RABAUD V,KONOLIGE K,et al.ORB:an effi-cient alternative to SIFT or SURF[C] //Proceedings of the 2011 International Conference on Computer Vision.Piscataway:IEEE,2011:2564-2571.
[18] XIANG Y,SCHMIDT T,NARAYANAN V,et al.PoseCNN:a convolution neural network for 6D object pose estimation in cluttered scenes[J].arXiv:1711.00199,2017.
[19] LI Y,WANG G,JI X Y,et al.DeepIM:deep iterative matching for 6D pose estimation[J].International Journal of Computer Vision,2020,128:657-678.
[20] LABBE Y,CARPENTIER J,AUBRY M,et al.CosyPose:consistent multi-view multi-object 6D pose estimation[C] //16th European Conference on Computer Vision.Cham:Springer,2020:574-591.
[21] MOON S P,SON H T,HUR D C,et al.GenFlow:generalizablerecurrent flow for 6D pose refinement of novel objects[C] //2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2024:10039-10049.
[22] PERIYASAMY A S,AMINI A,TSATURYAN V,et al.YOLOPose V2:understanding and improving transformer-based 6D pose estimation[J].arXiv:2307.11550,2023.
[23] WANG C,XU D F,ZHU Y K,et al.DenseFusion:6D object pose estimation by iterative dense fusion[C] //2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2019:3338-3347.
[24] HE Y S,HUANG H B,FAN H Q,et al.FFB6D:a full flow bidirectional fusion network for 6D pose estimation[C] //2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2021:3002-3012.
[25] PARK K,PATTEN T,VINCZE M.Pix2Pose:pixel-wise coordinate regression of objects for 6D pose estimation[C] //2019 IEEE/CVF International Conference on Computer Vision.Piscataway:IEEE,2019:7667-7676.
[26] SU Y Z,SALEH M,FETZER T,et al.ZebraPose:coarse to fine surface encoding for 6DoF object pose estimation[C] //2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2022:6728-6738.
[27] HODAN T,BARATH D,MATAS J.EPOS:estimating 6D pose of objects with symmetries[C] //2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2020:11700-11709.
[28] ORNEK E P,LABBE Y,TEKIN B,et al.FoundPose:unseenobject pose estimation with foundation features[C] //European Conference on Computer Vision.Cham:Springer,2024:163-182.
[29] LIU L H,LIN J H,LIU Z X,et al.PicoPose:progressive pixel-to-pixel correspondence learning for novel object pose estimation[J].arXiv:2504.02617,2025.
[30] WANG H,SRIDHAR S,HUANG J W,et al.Normalized objectcoordinate space for category-level 6D object pose and size estimation[C] //2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2019:2637-2646.
[31] LI Z G,WANG G,JI X Y.CDPN:coordinates-based disentangled pose network for real-time RGB-based 6-DoF object pose estimation[C] //2019 IEEE/CVF International Conference on Computer Vision.Piscataway:IEEE,2019:7677-7686.
[32] TEKIN B,SINHA S N,FUA P.Real-time seamless single shot 6D object pose prediction[C] //2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2018:292-301.
[33] ZAKHAROV S,SHUGUROV I,ILIC S.DPOD:6D pose object detector and refiner[C] //2019 IEEE/CVF International Confe-rence on Computer Vision.Piscataway:IEEE,2019:1941-1950.
[34] TAN T,DONG Q L.SMOC-Net:leveraging camera pose forself-supervised monocular object pose estimation[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2023:21307-21316.
[35] CHEN H Z,MANHARDT F,NAVAB N,et al.TexPose:neural texture learning for self-supervised 6D object pose estimation[C] //2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2023:4841-4852.
[36] WANG G,MANHARDT F,SHAO J Z,et al.Self6D:self-su-pervised monocular 6D object pose estimation[C] //European Conference on Computer Vision.Cham:Springer,2020:108-125.
[37] IWASE S,LIU X Y,KHIRODKAR R,et al.RePOSE:fast 6D object pose refinement via deep texture rendering[C] //2021 IEEE/CVF International Conference on Computer Vision.Piscataway:IEEE,2021:3283-3292.
[38] CHEN B,CHIN T J,KLIMAVICIUS M.Occlusion-robust object pose estimation with holistic representation[C] //2022 IEEE/CVF Winter Conference on Applications of Computer Vision.Piscataway:IEEE,2022:2223-2233.
[39] SHUGUROV I,ZAKHAROV S,ILIC S.DPODv2:dense correspondence-based 6 DoF pose estimation[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2022,44(11):7417-7435.
[40] LABBE Y,MANUELLI L,MOUSAVIAN A,et al.MegaPose:6D pose estimation of novel objects via render & compare[C] //Proceedings of the 6th Conference on Robot Learning.Cambridge,MA:JMLR,2023:715-725.
[41] NGUYEN V N,GROUEIX T,SALZMANN M,et al.Giga-Pose:fast and robust novel object pose estimation via one correspondence[C] //2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway:IEEE,2024:9903-9913.
[1] ZHANG Hu, XU Guolong, WANG Yujie. Knowledge-based Visual Question Answering Method Based on Hypergraph Convolutional Transformer [J]. Computer Science, 2026, 53(9): 333-339.
[2] GUO Peilin, ZOU Zhiyi, WANG Bo, LUO Jiawei. STMVF:Novel Multi-view Graph Convolutional Network for Spatial Transcriptomics Cell Deconvolution with Dual Cross-attention Mechanism [J]. Computer Science, 2026, 53(8): 40-49.
[3] PAN Yuquan, YUAN Deyu, WANG Anran, JIA Yuan. Enhanced GNNs Across Social Networks User Identity Linkage Algorithm Based on HiddenFeatures [J]. Computer Science, 2026, 53(8): 50-60.
[4] WU Shuiqing, QIU Jihao, LIU Xiang, DONG Yuxi, WEN Yimin. Proxy-based Source-free Domain Adaptation for EEG Emotion Recognition Method [J]. Computer Science, 2026, 53(8): 94-102.
[5] ZHOU Lijuan, LIU Zhihuan, LI Xinran, NIU Changyong. Zero-shot Skeleton-based Action Recognition Based on Class Knowledge Fusion and Latent SpaceOptimization [J]. Computer Science, 2026, 53(8): 148-155.
[6] CAI Yi, WANG Xiaobin, CHEN Ruili, XU Jinfeng. Handwriting Gender Recognition Method Based on Multi-scale Directional Attention Transformer [J]. Computer Science, 2026, 53(8): 156-164.
[7] HU Changyu, FAN Xinyu, ZHANG Zhi, DING Zixu, ZHANG Zhengyue, PENG Juhong. Hierarchical Lightweight Micro-expression Recognition Based on Optical Flow Partitioned FeatureFusion [J]. Computer Science, 2026, 53(8): 174-181.
[8] WANG Jingyang, XUE Weimin, HUANG Min, WU Shaoguang. Improved YOLOv11n Model for Small Target Detection in UAV Aerial Images [J]. Computer Science, 2026, 53(8): 201-208.
[9] YANG Chenguang, LU Jicang, GUO Jiaxing. Fake News Detection Model Based on Cross-modal Feature Fusion and Alignment [J]. Computer Science, 2026, 53(8): 257-265.
[10] LI Qingkai, QUN Nuo, NI Shengqiao, YANG Jin. From Bytes to Semantics:New Paradigm for Tibetan Named Entity Recognition [J]. Computer Science, 2026, 53(8): 266-275.
[11] REN Yanzhang, GAO Tai, LI Ying, WANG Bin. Gated Bidirectional Mamba Multimodal Feature Fusion Framework for Drug-Target InteractionPrediction [J]. Computer Science, 2026, 53(8): 326-335.
[12] LI Xiaochao, YUAN Zisu, LI Qianmu, LIU Fan, CHE Xun. Survey on Code Representation Learning for Vulnerability Detection [J]. Computer Science, 2026, 53(8): 388-402.
[13] JIANG Lingla, CHEN Wen, SUN Wei, ZHAO Kui. Research on Deep Learning-based Side-channel Analysis Method with Dynamically ComposableMulti-head Attention [J]. Computer Science, 2026, 53(8): 437-445.
[14] JIAO Hanbing, KANG Junhua, XIAO Teng, DENG Fei. Low-light Image Enhancement Network Based on Multi-level Illumination Excitation and JointLoss Constraint [J]. Computer Science, 2026, 53(7): 62-70.
[15] NING Shiqiang, ZHOU Lianzhen, ZHANG Lifeng. Identification of Authentic and Forged Paper-based Fingerprints Based on LG-GFNet Feature Fusion Network [J]. Computer Science, 2026, 53(7): 91-100.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!