计算机科学 ›› 2026, Vol. 53 ›› Issue (6A): 250600147-5.doi: 10.11896/jsjkx.250600147

• 图像处理&多媒体技术 • 上一篇    下一篇

基于拓扑信息的多层图卷积动作识别方法

黄海新, 何添禹, 侯广帅   

  1. 沈阳理工大学自动化与电气工程学院 沈阳 110159
  • 出版日期:2026-06-16 发布日期:2026-06-12
  • 通讯作者: 黄海新(huanghaixin@sylu.edu.cn)
  • 基金资助:
    国家重点研发计划(2022YFC3302500)

Multi-layer Graph Convolutional Action Recognition Method Based on Topological Information

HUANG Haixin, HE Tianyu, HOU Guangshuai   

  1. School of Automation and Electrical Engineering,Shenyang Ligong University,Shenyang 110159,China
  • Published:2026-06-16 Online:2026-06-12
  • About author:HUANG Haixin,born in 1973,Ph.D,associate professor.Her main research interests include machine learning,artificial intelligence and intelligent grid.
  • Supported by:
    Key R&D Program of China(2022YFC3302500).

摘要: 人体动作识别通过对视频中的时空特征进行分析,实现了对人体行为的识别。作为计算机视觉领域的重要研究课题之一,其高效准确的识别性能,已在人机交互、智能安防等多个应用场景中展现出广泛的应用价值。图卷积网络(Graph Convolutional Networks,GCNs)凭借在人体骨骼拓扑结构建模方面的显著优势,已成为动作识别任务中的主流方法。然而,现有方法通常对整体骨架结构进行统一建模,忽略了人体由多个功能性区域组成的层次化特征,导致模型在复杂行为识别任务中的表现受限。为此,提出了一种基于拓扑信息的多层图卷积网络(TM-GCN)。模型采用多分支架构,通过对人体骨架进行分区建模,有效捕捉骨骼节点间的空间依赖关系。同时引入拓扑感知单元,在图卷积过程中提取并融合拓扑特征,增强模型对骨骼拓扑信息的表达能力。基于 NTU-RGB+D 数据集的实验结果表明,TM-GCN在人体骨骼动作识别任务中取得了较为出色的性能,有效提升了动作识别的准确率。

关键词: 动作识别, 骨架模态, 图卷积网络, 拓扑感知, 计算机视觉

Abstract: Human action recognition achieves the identification of human behaviors by analyzing spatiotemporal features in vi-deos.As one of the important research topics in the field of computer vision,its efficient and accurate recognition performance has demonstrated wide application value in various scenarios such as human-computer interaction and intelligent security.Graph Convolutional Networks(GCNs),owing to their significant advantages in modeling human skeletal topology,have become a mainstream method for action recognition tasks.However,existing approaches generally adopt a unified modeling of the entire skeleton structure,overlooking the hierarchical characteristics of the human body composed of multiple functional regions.This limitation restricts model performance in complex action recognition tasks.To address these,this paper proposes a Topology-informed Multi-layer Graph Convolutional Network(TMGCN).The model employs a multi-branch architecture to partition and model the human skeleton,effectively capturing spatial dependencies between skeletal nodes.Additionally,it introduces a Topology Perception Unit(TPU) to extract and integrate topological features during graph convolution,enhancing the model's representation capability for skeletal topology.Experimental results based on NTU-RGB+D dataset show that TM-GCN has achieved excellent performance in human skeletal action recognition tasks,and effectively improved the accuracy of action recognition.

Key words: Action recognition, Skeleton modality, Graph convolutional network, Topology-aware, Computer vision

中图分类号: 

  • TP183
[1] PARK J Y,KIM J H.Online incremental classification reso-nance network and its application to human-robot interaction[J].IEEE Transactions on Neural Networks and Learning Systems,2019,31(5):1426-1436.
[2] CAO Z,SIMON T,WEI S E,et al.Realtime multi-person 2Dpose estimation using part affinityfields[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.2017:7291-7299.
[3] YAN S,XIONG Y,LIN D.Spatial temporal graph convolutional networks for skeleton-based action recognition[C]//Procee-dings of the AAAI Conference on Artificial Intelligence.2018.
[4] SHI L,ZHANG Y,CHENG J,et al.Two-stream adaptive graph convolutional networks for skeleton-based action recognition[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2019:12026-12035.
[5] TANG Y,TIAN Y,LU J,et al.Deep progressive reinforcement learning for skeleton-based action recognition[C]//Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition.2018:5323-5332.
[6] LEE J,LEE M,LEE D,et al.Hierarchically decomposed graph convolutional networks for skeleton-based action recognition[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2023:10444-10453.
[7] CHENG K,ZHANG Y,HE X,et al.Skeleton-based action recognition with shift graph convolutional network[C]//Procee-dings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2020:183-192.
[8] HEDEGAARD L,HEIDARI N,IOSIFIDIS A.Online skeleton-based action recognition with continual spatio-temporal graph convolutional networks[J].arXiv:2203.11009,2022.
[9] VELICˇKOVIĆ P,CUCURULL G,CASANOVA A,et al.Graph attention networks[J].arXiv:1710.10903,2017.
[10] YING C,CAI T,LUO S,et al.Do transformers really perform badly for graph representation?[J].Advances in Neural Information Processing Systems,2021,34:28877-28888.
[11] CHENG K,ZHANG Y,CAO C,et al.Decoupling GCN withdropgraph module for skeleton-based action recognition[C]//European Conference on Computer Vision.2020:536-553.
[12] ZHOU Y,YAN X,CHENG Z Q,et al.BlockGCN:Redefine Topology Awareness for Skeleton-Based Action Recognition[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).Seattle,WA,USA:IEEE,2024:2049-2058.
[13] TIAN Q,YU J J,ZHANG Z.Skeleton-Based Action Recogni-tion Combining Adaptive Local Graph Convolution and Multi-Scale Temporal Modeling[J].Computer Applications and Research,2025,42(7):2199-2205.
[14] CHEN H,SHEN Y,ZHANG Y,et al.Skeleton-Based Action Recognition through Dual-Granularity Feature Fusion with Self-Adapting Graph Convolution and Multi-Scale Temporal Convolution[J].Neurocomputing,2025,639:130261.
[15] YAN S J,XIONG Y J,LIN D H.Spatial temporal graph convolutional networks for skeleton-based action recognition[C]//Proceedings of the AAAI Conference on Artificial Intelligence.2018:7444-7452.
[16] LIN L,ZHANG J,LIU J.Actionlet-dependent contrastive learning skeleton-based action for unsupervised recognition[C]//Proceedings of the 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).Vancouver,BC,Canada,2023:2363-2372.
[17] HUA Y,WU W,ZHENG C,et al.Part aware contrastive learning for self-supervised action recognition[C]//Proceedings of the Thirty Second International Joint Conference on Artificial Intelligence.2023:855-863.
[18] ZHU Y S,HAN H,YU Z T,et al.Modeling the relative visual tempo for self-supervised skeleton-based action recognition[C]//2023 IEEE/CVF International Conference on Computer Vision(ICCV).2023:13867-13876.
[19] SHI L,ZHANG Y F,CHENG J,et al.Two stream adaptivegraph convolutional networks for skeleton-based action recognition[C]//Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition(CVPR).Long Beach,2019:12018-12027.
[20] CHI H G,HA M H,CHI S G,et al.InfoGCN:Representation learning for human skeleton-based action recognition[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2022:20186-20196.
[21] LEE J,LEE M,LEE D,et al.Hierarchically decomposed graph convolutional networks for skeleton-based action recognition[C]//Proceedings of the IEEE/CVF International Conference on Computer Vision.2023:10444-10453.
[22] ZHOU H Y,LIU Q J,WANG Y H.Learning discriminative representations for skeleton-based action recognition[C]//Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.2023:10608-10617.
Viewed
Full text


Abstract

Cited

  Shared   
  Discussed   
No Suggested Reading articles found!