Computer Science ›› 2026, Vol. 53 ›› Issue (8): 148-155.doi: 10.11896/jsjkx.250600091

• Computer Graphics & Multimedia • Previous Articles     Next Articles

Zero-shot Skeleton-based Action Recognition Based on Class Knowledge Fusion and Latent SpaceOptimization

ZHOU Lijuan, LIU Zhihuan, LI Xinran, NIU Changyong   

  1. School of Computer Science and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China
  • Received:2025-06-13 Revised:2025-10-13 Online:2026-08-15 Published:2026-08-17
  • About author:ZHOU Lijuan,born in 1987,Ph.D,associate professor,is a member of CCF(No.C1554M).Her main research interests include computer vision and multimodal processing.
    NIU Changyong,born in 1975,Ph.D,associate professor.His main research interests include large-scale data management and analysis based on cloud platforms,deep learning algorithm research and application,and mobile terminal system development.
  • Supported by:
    National Natural Science Foundation of China(62006211)and Young Backbone Teacher Training Program of Zhengzhou University.

Abstract: Considering the importance of different multi-source class knowledge and the independence of various body movements,this paper proposes a zero-shot skeleton-based action recognition method based on category knowledge fusion and latent space optimization.The method integrates multi-source category knowledge through a sparse attention mechanism with importance sampling,and introduces a total correlation loss to reduce redundant correlations among independent variables in the latent space.Specifically,it firstly employs pre-trained models to extract skeleton features and multi-source textual features,then learns text fusion representations through an importance sampling-based feature fusion approach.Subsequently,a total correlation-constrained generative cross-modal alignment method establishes semantic associations between skeleton and textual features.Fina-lly,latent space features generated from unseen class representations are utilized to train classifiers for recognizing unseen actions.Experiments on NTU RGB+D,NTU RGB+D 120,and PKU-MMD datasets demonstrate that the proposed method significantly outperforms all existing mainstream approaches.

Key words: Zero-shot learning, Skeleton action recognition, Cross-model alignment, Attention, Text feature fusion

CLC Number: 

  • TP183
[1] YAN W,YIN Y.Human Action Recognition Algorithm Based on Adaptive Shifted Graph Convolutional Neural Network with 3D Skeleton Similarity[J].Computer Science,2024,51(4):236-242.
[2] HUANG H,WAG Y,CAI M.Bottleneck Multi-scale GraphConvolutional Network for Skeleton-based Action Recognition[J].Computer Science,2024,51(S2):344-348.
[3] ZHOU L,LI W,OGUNBONA P,et al.Jointly Learning Visual Poses and Pose Lexicon for Semantic Action Recognition[J].IEEE Transactions on Circuits and Systems for Video Technology,2020,30(2):457-467.
[4] ZHOU L,LI W,OGUNBONA P,et al.Semantic Action Recognition by Learning a Pose Lexicon[J].Pattern Recognition,2017,72:548-562.
[5] ZHOU L,JIANG T.Learning Body Part-based Pose Lexiconsfor Semantic Action Recognition[J].IET Computer Vision,2023,17(2):135-155.
[6] LI S W,WEI Z X,CHEN W J,et al.Sa-dvae:Improving zero-shot skeleton-based action recognition by disentangled variational autoencoders[C]//European Conference on Computer Vision.Cham:Springer Nature Switzerland,2024:447-462.
[7] ZHOU Y,QIANG W,RAO A,et al.Zero-shot skeleton-based action recognition via mutual information estimation and maximization[C]//Proceedings of ACM International Conference on Multimedia.New York:ACM Press,2023:5302-5310.
[8] LI M Z,JIA Z,ZHANG Z,et al.Multi-semantic fusion model for generalized zero-shot skeleton-based action recognition[C]//International Conference on Image and Graphics.Cham:Springer Nature Switzerland,2023:68-80.
[9] ZHOU L,JIAO X.Multi-modal and Multi-part with Skeletons and Texts for Action Recognition[J].Expert Systems With Applications,2025,272,126646.
[10] CHEN S,HUANG D.Elaborative rehearsal for zero-shot action recognition[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.Los Alamitos:IEEE Computer Society,2021:13638-13647.
[11] ZHOU L,MAO J.Improving Class Representation for Zero-Shot Action Recognition[C]//Proceedings of the ACM International Conference on Multimedia in Asia.New York:ACM Press,2023:1-7.
[12] FROME A,CORRADO G S,SHLENS J,et al.Devise:A deep visual-semantic embedding model[C]//Advances in Neural Information Processing Systems 26.San Diego,CA:Neural Information Processing Systems Foundation,2013:2121-2129.
[13] JASANI B,MAZAGONWALLA A.Skeleton based zero shotaction recognition in joint pose-language semantic space[J].ar-Xiv:1911.11344,2019.
[14] CHEN Y,GUO J,HE T,et al.Fine-grained side information guided dual-prompts for zero-shot skeleton action recognition[C]//Proceedings of ACM International Conference on Multimedia.New York:ACM Press,2024:778-786.
[15] GAO J,WEI W,YUE Q,et al.Embedded Zero-shot Learning Algorithm Based on Swin Transformer[J].Journal of Chinese Computer Systems,2024,45(4):784-791.
[16] SCHONFELD E,EBRAHIMI S,SINHA S,et al.Generalized zero-and few-shot learning via aligned variational autoencoders[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2019:8247-8255.
[17] FENG Y,YU J,SANG J,et al.Survey on Knowledge-based Zero-shot Visual Recognition[J].Journal of Software,2021,32(2):370-405.
[18] GAO J,ZHANG T,XU C.I know the relationships:Zero-shot action recognition via two-stream graph convolutional networks and knowledge graphs[C]//Proceedings of the AAAIConfe-rence on Artificial Intelligence.Menlo Park,CA:AAAI Press,2019:8303-8311.
[19] LAMPERT C H,NICKISCH H,HARMELING S.Attribute-based classification for zero-shot visual object categorization[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2014,36(3):453-465.
[20] ZHOU L,MAO J,XU X,et al.Generative external knowledge for zero-shot action recognition[J].Expert Systems with Applications,2025,279:127420.
[21] GUPTA P,SHARMA D,SARVADEVABHATLA R K.Syntactically guided generative embeddings for zero-shot skeleton action recognition[C]//Proceedings of the 2021 IEEE International Conference on Image Processing.Piscataway,NJ:IEEE,2021:439-443.
[22] SHAHROUDY A,LIU J,NG T T,et al.Ntu rgb+d:A large scale dataset for 3d human activity analysis[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2016:1010-1019.
[23] LIU J,SHAHROUDY A,PEREZ M,et al.Ntu rgb+d 120:A large-scale benchmark for 3d human activity understanding[J].IEEE Transactions on Pattern Analysis and Machine Intelligence,2019,42(10):2684-2701.
[24] LIU C,HU Y,LI Y,et al.Pku-mmd:A large scale benchmark for continuous multi-modal human action understanding[J].arXiv:1703.07475,2017.
[25] CHENG K,ZHANG Y,HE X,et al.Skeleton-based action re-cognition with shift graph convolutional network[C]//Procee-dings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.Piscataway,NJ:IEEE,2020:183-192.
[26] YAN S,XIONG Y,LIN D.Spatial temporal graph convolutional networks for skeleton-based action recognition[C]//Procee-dings of the AAAI Conference on Artificial Intelligence.Menlo Park,CA:AAAI Press,2018,.
[27] DOSOVITSKIY A,BEYER L,KOLESNIKOV A,et al.Animage is worth 16×16 words:Transformers for image recognition at scale[J].arXiv:2010.11929,2020.
[28] HUBERT TSAI Y H,HUANG L K,SALAKHUTDINOV R.Learning robust visual-semantic embeddings[C]//Proceedings of the IEEE International Conference on Computer Vision.Piscataway,NJ:IEEE,2017:3571-3580.
[29] WRAY M,LARLUS D,CSURKA G,et al.Fine-grained action re-trieval through multiple parts-of-speech embeddings[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.Piscataway,NJ:IEEE,2019:450-459.
[1] CAI Yi, WANG Xiaobin, CHEN Ruili, XU Jinfeng. Handwriting Gender Recognition Method Based on Multi-scale Directional Attention Transformer [J]. Computer Science, 2026, 53(8): 156-164.
[2] YANG Chenguang, LU Jicang, GUO Jiaxing. Fake News Detection Model Based on Cross-modal Feature Fusion and Alignment [J]. Computer Science, 2026, 53(8): 257-265.
[3] FENG Haoyu, ZHANG Yuxuan, LIU Zixuan, MENG Hua. Network Anomaly Traffic Detection Based on Deep Multi-instance Learning [J]. Computer Science, 2026, 53(8): 426-436.
[4] JIANG Lingla, CHEN Wen, SUN Wei, ZHAO Kui. Research on Deep Learning-based Side-channel Analysis Method with Dynamically ComposableMulti-head Attention [J]. Computer Science, 2026, 53(8): 437-445.
[5] FU Le, HUANG Xiaofang, LIAO Min, SONG Luhua. Dynamic Adversarial Detection Framework Based on Multimodal Uncertainty Fusion [J]. Computer Science, 2026, 53(8): 469-477.
[6] LU Shan, LIU Yuelong, ZHAO Zhiqi, GU Jie. Social Networks and Stock Price Synchronicity:Explainable Predictive Model Based on GraphAttention Network [J]. Computer Science, 2026, 53(8): 85-93.
[7] LIAO Xuechao, ZOU Hang, LYU Peidong, ZENG Zhiqiang. Lithium-ion Battery Capacity Data Augmentation and Prediction Based on Diffusion,Denoise and Coding-Decoding Attention [J]. Computer Science, 2026, 53(8): 103-116.
[8] DENG Jiayan, TIAN Shirui, LIU Hou, ZHU Ningbo, DUAN Mingxing. Zero-shot Pedestrian Trajectory Prediction Method Based on Compositional Motion [J]. Computer Science, 2026, 53(8): 117-126.
[9] ZHU Bin, LI Xiaobin. AETC:Image Classification Model via Attention-based Topological Features Fusion [J]. Computer Science, 2026, 53(7): 24-33.
[10] CHEN Yifan, DING Cong, CAO Min. Image Anomaly Detection Based on Masked Convolutional Kernel [J]. Computer Science, 2026, 53(7): 45-53.
[11] ZHANG Yan, ZHOU Jian, HAN Lei, CHENG Chunling. Gate-controlled Agent Attention Mechanism-based Soft Prompt Transfer Method [J]. Computer Science, 2026, 53(7): 139-145.
[12] DING Zhijun, AI Fangju, LIU Aihuan. Document-level Event Argument Extraction Model Based on Hierarchical Dependency Aggregation and Event Enhancement [J]. Computer Science, 2026, 53(7): 156-167.
[13] LI Kaiju, YIN Chenyang, CHENG Zhangtao, LIU Xueting, ZHOU Fan. Causal Subgraph Learning for Cascade Popularity Prediction [J]. Computer Science, 2026, 53(7): 213-221.
[14] XU Jian, CHEN Shijie, FENG Jiancong, YANG Geng. CBT-AD:CNN-BiLSTM-Transformer Hybrid Model for Time Series Anomaly Detection [J]. Computer Science, 2026, 53(7): 230-241.
[15] WU Kai, SUN Zhe, ZHANG Xu, CAO Yadong, SUN Zhixin. Research on Modeling and Scheduling Methods for Intra-city Delivery Based on Heterogeneous Graph Neural Networks [J]. Computer Science, 2026, 53(7): 242-250.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!