计算机科学 ›› 2026, Vol. 53 ›› Issue (8): 148-155.doi: 10.11896/jsjkx.250600091

• 计算机图形学 & 多媒体 • 上一篇    下一篇

基于类别知识融合与隐空间优化的零样本骨架动作识别

周丽娟, 刘治宦, 李欣冉, 牛常勇   

  1. 郑州大学计算机与人工智能学院 郑州 450001
  • 收稿日期:2025-06-13 修回日期:2025-10-13 出版日期:2026-08-15 发布日期:2026-08-17
  • 通讯作者: 牛常勇(iecyniu@zzu.edu.cn)
  • 作者简介:(ieljzhou@zzu.edu.cn)
  • 基金资助:
    国家自然科学基金(62006211);郑州大学青年骨干教师培养计划

Zero-shot Skeleton-based Action Recognition Based on Class Knowledge Fusion and Latent SpaceOptimization

ZHOU Lijuan, LIU Zhihuan, LI Xinran, NIU Changyong   

  1. School of Computer Science and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China
  • Received:2025-06-13 Revised:2025-10-13 Published:2026-08-15 Online:2026-08-17
  • About author:ZHOU Lijuan,born in 1987,Ph.D,associate professor,is a member of CCF(No.C1554M).Her main research interests include computer vision and multimodal processing.
    NIU Changyong,born in 1975,Ph.D,associate professor.His main research interests include large-scale data management and analysis based on cloud platforms,deep learning algorithm research and application,and mobile terminal system development.
  • Supported by:
    National Natural Science Foundation of China(62006211)and Young Backbone Teacher Training Program of Zhengzhou University.

摘要: 考虑到不同多源类别知识的重要性和不同肢体动作的独立性,提出一种基于类别知识融合与隐空间优化的零样本骨架动作识别方法。该方法通过重要性采样的稀疏注意力机制融合多源类别知识,并引入总相关性损失降低隐空间各独立变量间的冗余相关性。具体地,首先采用预训练模型提取骨架和多源文本特征,通过基于重要性采样的特征融合方法学习文本融合表示;然后,采用总相关性约束的生成式跨模态对齐方法学习骨架与文本特征的语义关联;最后,使用未见类类别表示生成的隐空间特征训练未见类分类器,实现对未见类动作的识别。在NTU RGB+D,NTU RGB+D 120和PKU-MMD数据集上的实验结果表明,所提方法显著优于现有主流方法。

关键词: 零样本学习, 骨架动作识别, 跨模态对齐, 注意力, 文本特征融合

Abstract: Considering the importance of different multi-source class knowledge and the independence of various body movements,this paper proposes a zero-shot skeleton-based action recognition method based on category knowledge fusion and latent space optimization.The method integrates multi-source category knowledge through a sparse attention mechanism with importance sampling,and introduces a total correlation loss to reduce redundant correlations among independent variables in the latent space.Specifically,it firstly employs pre-trained models to extract skeleton features and multi-source textual features,then learns text fusion representations through an importance sampling-based feature fusion approach.Subsequently,a total correlation-constrained generative cross-modal alignment method establishes semantic associations between skeleton and textual features.Fina-lly,latent space features generated from unseen class representations are utilized to train classifiers for recognizing unseen actions.Experiments on NTU RGB+D,NTU RGB+D 120,and PKU-MMD datasets demonstrate that the proposed method significantly outperforms all existing mainstream approaches.

Key words: Zero-shot learning, Skeleton action recognition, Cross-model alignment, Attention, Text feature fusion

中图分类号: 

  • TP183
[1] YAN W,YIN Y.Human Action Recognition Algorithm Based on Adaptive Shifted Graph Convolutional Neural Network with 3D Skeleton Similarity[J].Computer Science,2024,51(4):236-242.
[2] HUANG H,WAG Y,CAI M.Bottleneck Multi-scale GraphConvolutional Network for Skeleton-based Action Recognition[J].Computer Science,2024,51(S2):344-348.
[3] ZHOU L,LI W,OGUNBONA P,et al.Jointly Learning Visual Poses and Pose Lexicon for Semantic Action Recognition[J].IEEE Transactions on Circuits and Systems for Video Technology,2020,30(2):457-467.
[4] ZHOU L,LI W,OGUNBONA P,et al.Semantic Action Recognition by Learning a Pose Lexicon[J].Pattern Recognition,2017,72:548-562.
[5] ZHOU L,JIANG T.Learning Body Part-based Pose Lexiconsfor Semantic Action Recognition[J].IET Computer Vision,2023,17(2):135-155.
[6] LI S W,WEI Z X,CHEN W J,et al.Sa-dvae:Improving zero-shot skeleton-based action recognition by disentangled variational autoencoders[C]//European Conference on Computer Vision.Cham:Springer Nature Switzerland,2024:447-462.
[7] ZHOU Y,QIANG W,RAO A,et al.Zero-shot skeleton-based action recognition via mutual information estimation and maximization[C]//Proceedings of ACM International Conference on Multimedia.New York:ACM Press,2023:5302-5310.
[8] LI M Z,JIA Z,ZHANG Z,et al.Multi-semantic fusion model for generalized zero-shot skeleton-based action recognition[C]//International Conference on Image and Graphics.Cham:Springer Nature Switzerland,2023:68-80.
[9] ZHOU L,JIAO X.Multi-modal and Multi-part with Skeletons and Texts for Action Recognition[J].Expert Systems With Applications,2025,272,126646.
[10] CHEN S,HUANG D.Elaborative rehearsal for zero-shot action recognition[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.Los Alamitos:IEEE Computer Society,2021:13638-13647.
[11] ZHOU L,MAO J.Improving Class Representation for Zero-Shot Action Recognition[C]//Proceedings of the ACM International Conference on Multimedia in Asia.New York:ACM Press,2023:1-7.
[12] FROME A,CORRADO G S,SHLENS J,et al.Devise:A deep visual-semantic embedding model[C]//Advances in Neural Information Processing Systems 26.San Diego,CA:Neural Information Processing Systems Foundation,2013:2121-2129.
[13] JASANI B,MAZAGONWALLA A.Skeleton based zero shotaction recognition in joint pose-language semantic space[J].ar-Xiv:1911.11344,2019.
[14] CHEN Y,GUO J,HE T,et al.Fine-grained side information guided dual-prompts for zero-shot skeleton action recognition[C]//Proceedings of ACM International Conference on Multimedia.New York:ACM Press,2024:778-786.
[15] GAO J,WEI W,YUE Q,et al.Embedded Zero-shot Learning Algorithm Based on Swin Transformer[J].Journal of Chinese Computer Systems,2024,45(4):784-791.
[16] SCHONFELD E,EBRAHIMI S,SINHA S,et al.Generalized zero-and few-shot learning via aligned variational autoencoders[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2019:8247-8255.
[17] FENG Y,YU J,SANG J,et al.Survey on Knowledge-based Zero-shot Visual Recognition[J].Journal of Software,2021,32(2):370-405.
[18] GAO J,ZHANG T,XU C.I know the relationships:Zero-shot action recognition via two-stream graph convolutional networks and knowledge graphs[C]//Proceedings of the AAAIConfe-rence on Artificial Intelligence.Menlo Park,CA:AAAI Press,2019:8303-8311.
[19] LAMPERT C H,NICKISCH H,HARMELING S.Attribute-based classification for zero-shot visual object categorization[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2014,36(3):453-465.
[20] ZHOU L,MAO J,XU X,et al.Generative external knowledge for zero-shot action recognition[J].Expert Systems with Applications,2025,279:127420.
[21] GUPTA P,SHARMA D,SARVADEVABHATLA R K.Syntactically guided generative embeddings for zero-shot skeleton action recognition[C]//Proceedings of the 2021 IEEE International Conference on Image Processing.Piscataway,NJ:IEEE,2021:439-443.
[22] SHAHROUDY A,LIU J,NG T T,et al.Ntu rgb+d:A large scale dataset for 3d human activity analysis[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2016:1010-1019.
[23] LIU J,SHAHROUDY A,PEREZ M,et al.Ntu rgb+d 120:A large-scale benchmark for 3d human activity understanding[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2019,42(10):2684-2701.
[24] LIU C,HU Y,LI Y,et al.Pku-mmd:A large scale benchmark for continuous multi-modal human action understanding[J].arXiv:1703.07475,2017.
[25] CHENG K,ZHANG Y,HE X,et al.Skeleton-based action re-cognition with shift graph convolutional network[C]//Procee-dings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2020:183-192.
[26] YAN S,XIONG Y,LIN D.Spatial temporal graph convolutional networks for skeleton-based action recognition[C]//Procee-dings of the AAAI Conference on Artificial Intelligence.Menlo Park,CA:AAAI Press,2018,.
[27] DOSOVITSKIY A,BEYER L,KOLESNIKOV A,et al.Animage is worth 16×16 words:Transformers for image recognition at scale[J].arXiv:2010.11929,2020.
[28] HUBERT TSAI Y H,HUANG L K,SALAKHUTDINOV R.Learning robust visual-semantic embeddings[C]//Proceedings of the IEEE International Conference on Computer Vision.Piscataway,NJ:IEEE,2017:3571-3580.
[29] WRAY M,LARLUS D,CSURKA G,et al.Fine-grained action re-trieval through multiple parts-of-speech embeddings[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.Piscataway,NJ:IEEE,2019:450-459.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!