The skeleton-based methods have received extensive attention in human action recognition. However, the simple skeleton sequence representation is insufficient for extracting the rich spatio-temporal information, as many two-branch networks typically incur high computational costs. This study introduces a spatio-temporal representation method for a Large Scale Joint Matrix (LSJM), which can capture rich spatial semantic information and improve the recognition performance in deep convolutional neural networks. Furthermore, this paper proposes a Dual Long-Short Cascade Network (DLSCNet), which can fully exploit the temporal information related to Long-Short time scales and which features an increased computational efficiency. This approach is evaluated and validated on two benchmark datasets and one self-collected dataset, namely Florence-3D, UTKinect-3D, and HanYue-3D. The experimental results demonstrate that this approach can effectively support large neural networks and improve the computational efficiency while maintaining the recognition accuracy.
human skeleton, action recognition, skeleton sequence representation, spatio-temporal information, Dual Long- Short Cascade Network.
Yitong ZHOU, Leiyue YAO, Chao ZENG, Qing YE, "Skeleton-based Large-Scale Joint Matrix and Dual Long- Short Cascade Network for Human Action Recognition", Studies in Informatics and Control, ISSN 1220-1766, vol. 35(3), pp. 15-28, 2026. https://doi.org/10.24846/v35i3y202602