Jing Wang1, Haojuan Wang2, and Jinglong Geng3
1Shanxi Vocational University of Engineering Science and Technology, Physical Education College, Shanxi Jinzhong 030619, China
2Shanxi University of Chinese Medicine, Shanxi Jinzhong 030619, China
3Taiyuan Normal University, Shanxi Jinzhong 030619, China
Received: May 12, 2026
Accepted: July 07, 2026
Publication Date: August 22, 2026
CNN-based spatial feature extraction pipeline
Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.
Download Citation: BibTeX | http://dx.doi.org/10.6180/jase.202611_34.060
To address the challenges of significant fine-grained action differences, complex multi-agent interactions, and strong temporal dynamic dependencies in volleyball scenarios, this paper proposes an action recognition method based on multimodal spatiotemporal fusion. First, convolutional neural networks are used to extract spatial features from images, and human pose information is combined to enhance the representation of fine-grained structural differences. Second, a hierarchical temporal modeling framework is constructed, using individual level LSTM and group-level LSTM to characterize the action evolution of individual players and multi-agent interaction relationships, respectively. Simultaneously, a People Pooling mechanism is introduced to adaptively fuse features from different subjects, highlighting key roles and modeling “ball-player-player” interactions. Finally, RGB and Skeleton multimodal information are fused to improve model robustness. Experimental results on the publicly available Volleyball dataset and a self-built dataset demonstrate that the proposed method outperforms several baseline methods in terms of accuracy, F1-score, and mAP, validating its effectiveness and generalization ability in fine-grained action recognition and complex scene modeling.
Keywords: Volleyball motion recognition; fine-grained motion modeling; multi-agent interaction; temporal modeling; multimodal fusion
- [1] I. Zahra, Y. Wu, H. F. Alhasson, S. S. Alharbi, H. Aljuaid, A. Jalal, and H. Liu, (2025) “Dynamic graph neural networks for UAV-based group activity recognition in structured team sports” Frontiers in Neurorobotics 19: 1631998. DOI: 10.3389/fnbot.2025.1631998.
- [2] A. Alqarafi and B. Almogadwy, (2025) “STRIKE-net: An explainable dynamic spatiotemporal graph-transformer network for fine-grained soccer action recognition” Applied Soft Computing: 114224. DOI: 10.1016/j.asoc.2025.114224.
- [3] K. Seweryn, A. Wróblewska, and S. Łukasik, (2026) “Survey of action recognition, spotting, and spatiotemporal localization in soccer—current trends and research perspectives” ACM Transactions on Intelligent Systems and Technology 17(2): 1–37. DOI: 10.1145/3776541.
- [4] C. Zhang, G. Shan, J. Lim, and B. H. Roh, (2024) “Dynamic reinforcement learning for optimal go AI training: adaptive adjustment and optimization” IEEE Transactions on Consumer Electronics 71(1): 292–302. DOI: 10.1109/TCE.2024.3487141.
- [5] M. A. Hossen and P. E. Abas, (2025) “Machine learning for human activity recognition: State-of-the-art techniques and emerging trends” Journal of Imaging 11(3): 91. DOI: 10.3390/jimaging11030091.
- [6] N. V. S. R. Chappa and K. Luu. “LiGAR: LiDAR-guided hierarchical transformer for multi-modal group activity recognition”. In: 2025 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV). 2025, 3035–3044. DOI: 10.1109/WACV61041.2025.00300.
- [7] D. Karki, M. Ramazanova, A. Cioppa, S. Giancola, and B. Ghanem. “Pixels or positions? benchmarking modalities in group activity recognition”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2026, 9998–10008. DOI: 10.48550/arXiv.2511.12606.
- [8] L. Wang, P. Koniusz, and Y. Gao, (2025) “Video Understanding by Design: How Datasets Shape Architectures and Insights” arXiv preprint arXiv:2509.09151: DOI: 10.48550/arXiv.2509.09151.
- [9] R. Zhang, Y. Wang, X. Hu, C. Mai, W. Liu, D. Xu, et al. “Beyond the Individual: Introducing Group Intention Forecasting with SHOT Dataset”. In: Proceedings of the 33rd ACM International Conference on Multimedia. 2025, 13002–13008. DOI: 10.1145/3746027.3758248.
- [10] Y. Liu, F. Liu, L. Jiao, Q. Bao, L. Li, Y. Guo, and P. Chen, (2024) “A knowledge-based hierarchical causal inference network for video action recognition” IEEE Transactions on Multimedia 26: 9135–9149. DOI: 10.1109/TMM.2024.3386339.
- [11] Z. Chen, W. Sun, Y. Tian, J. Jia, Z. Zhang, J. Wang, et al. “Gaia: Rethinking action quality assessment for AI-generated videos”. In: Advances in Neural Information Processing Systems. 37. 2024, 40111–40144. DOI: 10.52202/079017-1267.
- [12] X. Wang, X. Lan, and W. Zhu. Video Grounding and Its Generalization: From ID and Task-specific Models to OOD and Large Foundation Models. Springer International Publishing AG, 2025. DOI: 10.1007/978-3-031-94837-4.
- [13] Y. Wen, M. Liu, S. Wu, and B. Ding. “Chase: Learning convex hull adaptive shift for skeleton-based multi-entity action recognition”. In: Advances in Neural Information Processing Systems. 37. 2024, 9388–9420. DOI: 10.52202/079017-0298.
- [14] B. Amirgaliyev, M. Mussabek, T. Rakhimzhanova, and A. Zhumadillayeva, (2025) “A review of machine learning and deep learning methods for person detection, tracking and identification, and face recognition with applications” Sensors 25(5): 1410. DOI: 10.3390/s25051410.
- [15] K. Ashutosh and K. Grauman. “Learning skill-attributes for transferable assessment in video”. In: Advances in Neural Information Processing Systems. 38. 2026, 160403–160430. DOI: 10.48550 / arXiv. 2511 . 13993.
