计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250700192-7.doi: 10.11896/jsjkx.250700192

• 图像处理&多媒体技术 • 上一篇    下一篇

用于无人机-卫星跨视角视觉地理定位的金字塔池化视觉状态空间模型

岳文洁1, 蒋杰1, 詹礼新1, 周秉泉2, 周天健1   

  1. 1 国防科技大学系统工程学院 长沙 410000
    2 中国信息通信研究院 北京 100000
  • 出版日期:2026-06-16 发布日期:2026-06-12
  • 通讯作者: 蒋杰(jiejiang@nudt.edu.cn)
  • 作者简介:(yuewenjie19@nudt.edu.cn)

Pyramid Pooling Visual State Space Model for UAV-Satellite Cross-view Geo-localization

YUE Wenjie1, JIANG Jie1, ZHAN Lixin1, ZHOU Bingquan2, ZHOU Tianjian1   

  1. 1 College of System Engineering,National University of Defense Technology,Changsha 410000,China
    2 China Academy of Information and Communications Technology,Beijing 100000,China
  • Published:2026-06-16 Online:2026-06-12
  • About author:YUE Wenjie,born in 2001,postgra-duate.Her main research interest is UAV cross-view geo-localization.
    JIANG Jie,born in 1974,Ph.D,professor,Ph.D supervisor.His main research interests include artificial intelligence and deep learning,visualization and vi-sual analytics,virtual reality and intelligent interaction.

摘要: 无人机与卫星图像之间的跨视角地理定位作为一种替代GNSS-INS的有力手段,在卫星信号弱或受阻的环境下展现出巨大潜力。然而,由于视角、光照和分辨率等差异所带来的巨大视觉变化,会对图像匹配造成严峻挑战。为此,提出了一种新型方法P2VSSM(Pyramid Pooling Visual State Space Model),通过将金字塔池化自注意力机制嵌入Mamba架构,提升了跨视角图像的特征提取能力。设计的PPSA模块融合多尺度上下文信息,增强语义抽象能力和全局建模效果。同时,引入信息噪声对比估计损失替代传统三元组损失,避免了繁重的困难负样本挖掘过程,在训练过程中显著提升了负样本的丰富度与对比学习的稳定性。在两个公开的无人机-卫星数据集University-1652与SUES-200上的实验结果表明,此方法在无人机到卫星以及卫星到无人机的检索任务中均取得了领先性能。大量消融实验也验证了所提方法的有效性与鲁棒性。

关键词: 跨视角地理定位, 无人机-卫星匹配, Mamba, 金字塔池化, 视觉表示学习

Abstract: Cross-view geo-localization between UAV and satellite images has emerged as a promising alternative to GNSS-INS,particularly in environments where satellite signals are weak or obstructed.However,significant visual discrepancies caused by differences in viewpoint,illumination,and resolution pose considerable challenges for image matching.To address this issue,a novel method called P2VSSM(Pyramid Pooling Visual State Space Model) is proposed.By integrating a pyramid pooling self-attention mechanism into the Mamba architecture,the model enhances feature extraction capabilities for cross-view images.The proposed PPSA module aggregates multi-scale contextual information,improving both semantic abstraction and global modeling.Additionally,the InfoNCE loss is introduced to replace the traditional triplet loss,thereby avoiding the heavy burden of hard negative mining and significantly improving the diversity of negative samples and the stability of contrastive learning during training.Experimental results on two public UAV-satellite datasets,University-1652 and SUES-200,demonstrate that the proposed method achieves state-of-the-art performance in both UAV-to-satellite and satellite-to-UAV retrieval tasks.Extensive ablation studies further confirm the effectiveness and robustness of the proposed approach.

Key words: Cross-view geo-localization, UAV-satellite matching, Mamba, Pyramid pooling, Representation learning

中图分类号: 

  • TP391
[1] MOHSAN S A H,OTHMAN N Q H,LI Y,et al.Unmanned aerial vehicles(UAVs):Practical aspects,applications,open challenges,security issues,and future trends[J].Intelligent Service Robotics,2023,16(1):109-137.
[2] ANGRISANO A.GNSS/INS integration methods[D].Naples:Università degli Studi di Napoli “Parthenope”,2010.
[3] GROVES P D,JIANG Z,RUDI M,et al.A portfolio approach to NLOS and multipath mitigation in dense urban areas[C]//The Institute of Navigation.2013.
[4] COUTURIER A,AKHLOUFI M A.A review on absolute vi-sual localization for UAV[J].Robotics and Autonomous Systems,2021,135:103666.
[5] GU A,DAO T.Mamba:Linear-time sequence modeling with selective state spaces[J].arXiv:2312.00752,2023.
[6] ZHU Q,FANG Y,CAI Y,et al.Rethinking scanning strategies with vision mamba in semantic segmentation of remote sensing imagery:an experimental study[J].IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing,2024,17:18223-18234.
[7] ZHENG Z,WEI Y,YANG Y.University-1652:A multi-viewmulti-source benchmark for drone-based geo-localization[C]//Proceedings of the 28th ACM International Conference on Multimedia.2020:1395-1403.
[8] ZHU R,YIN L,YANG M,et al.SUES-200:A multi-heightmulti-scene cross-view image benchmark across drone and satellite[J].IEEE Transactions on Circuits and Systems for Video Technology,2023,33(9):4825-4839.
[9] CASTALDO F,ZAMIR A,ANGST R,et al.Semantic cross-view matching[C]//Proceedings of the IEEE International Conference on Computer Vision Workshops.2015:9-17.
[10] LIN T Y,BELONGIE S,HAYS J.Cross-view image geoloca-lization[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.2013:891-898.
[11] SENLET T,ELGAMMAL A.A framework for global vehicle localization using stereo images and satellite and road maps[C]//2011 IEEE International Conference on Computer Vision Workshops(ICCV Workshops).IEEE,2011:2034-2041.
[12] WORKMAN S,SOUVENIR R,JACOBS N.Wide-area image geolocalization with aerial reference imagery[C]//Proceedings of the IEEE International Conference on Computer Vision.2015:3961-3969.
[13] VO N N,HAYS J.Localizing and orienting street viewsusing overhead imagery[C]//European Conference on Computer Vision.Cham:Springer,2016:494-509.
[14] HU S,FENG M,NGUYEN R M H,et al.Cvm-net:Cross-view matching network for image-based ground-to-aerial geo-localization[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.2018:7258-7267.
[15] SHI Y,LIU L,YU X,et al.Spatial-aware feature aggregation for cross-view image based geo-localization[C]//Advances in Neural Information Processing Systems.2019.
[16] SHI Y,YU X,LIU L,et al.Optimal feature transport for cross-view image geo-localization[C]//Proceedings of the AAAI Conference on Artificial Intelligence.2020:11990-11997.
[17] YANG H,LU X,ZHU Y.Cross-view geo-localization with layer-to-layer transformer[J].Advances in Neural Information Processing Systems,2021,34:29009-29020.
[18] ZHU S,SHAH M,CHEN C.Transgeo:Transformer is all you need for cross-view image geo-localization[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:1162-1171.
[19] ALI-BEY A,CHAIB-DRAA B,GIGUERE P.Mixvpr:Feature mixing for visual place recognition[C]//Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.2023:2998-3007.
[20] YE J,LIN H,OU L,et al.Where am I? Cross-View Geo-localization with Natural Language Descriptions[J].arXiv:2412.17007,2024.
[21] ZHU L,LIAO B,ZHANG Q,et al.Vision mamba:efficient vi-sual representation learning with bidirectional state space model[C]//Proceedings of the 41st International Conference on Machine Learning(ICML'24).JMLR.org,2024:62429-62442.
[22] DOSOVITSKIY A,BEYER L,KOLESNIKOV A,et al.AnImage is Worth 16x16 Words:Transformers for Image Recognition at Scale[J].arXiv:2010.11929,2020.
[23] LIU Y,TIAN Y,ZHAO Y,et al.Vmamba:Visual state space model[J].Advances in Neural Information Processing Systems,2024,37:103031-103063.
[24] YANG C,CHEN Z,ESPINOSA M,et al.Plainmamba:Improving non-hierarchical mamba in visual recognition[J].arXiv:2403.17695,2024.
[25] SCHROFF F,KALENICHENKO D,PHILBIN J.Facenet:Aunified embedding for face recognition and clustering[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.2015:815-823.
[26] OORD A,LI Y,VINYALS O.Representation learning with contrastive predictive coding[J].arXiv:1807.03748,2018.
[27] DENG J,DONG W,SOCHER R,et al.Imagenet:A large-scale hierarchical image database[C]//2009 IEEE Conference on Computer Vision and Pattern Recognition.IEEE,2009:248-255.
[28] LOSHCHILOV I,HUTTER F.Decoupled Weight Decay Regularization[C]//International Conference on Learning Representations.2017.
[29] DING L,ZHOU J,MENG L,et al.A practical cross-view image matching method between UAV and satellite for UAV-based geo-localization[J].Remote Sensing,2020,13(1):47.
[30] LEIBE B,LEONARDIS A,SCHIELE B.Robust object detection with interleaved categorization and segmentation[J].International Journal of Computer Vision,2008,77(1):259-289.
[31] WANG T,ZHENG Z,YAN C,et al.Each part matters:Localpatterns facilitate cross-view geo-localization[J].IEEE Transactions on Circuits and Systems for Video Technology,2021,32(2):867-879.
[32] DAI M,HU J,ZHUANG J,et al.A transformer-based feature segmentation and region alignment method for UAV-view geo-localization[J].IEEE Transactions on Circuits and Systems for Video Technology,2021,32(7):4376-4389.
[33] GE F,ZHANG Y,WANG L,et al.Multilevel feedback jointrepresentation learning network based on adaptive area elimination for cross-view geo-localization[J].IEEE Transactions on Geoscience and Remote Sensing,2024,62:1-15.
[34] CHEN Q,WANG T,YANG Z,et al.Sdpl:Shifting-dense partition learning for uav-view geo-localization[J].IEEE Transactions on Circuits and Systems for Video Technology,2024,34(11):11810-11824.
[35] DU H,HE J,ZHAO Y.CCR:A counterfactual causal reasoning-based method for cross-view geo-localization[J].IEEE Transactions on Circuits and Systems for Video Technology,2024,34(11):11630-11643.
[36] LYU H,ZHU H,ZHU R,et al.Direction-guided multiscale feature fusion network for geo-localization[J].IEEE Transactions on Geoscience and Remote Sensing,2024,62:1-13.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!