{"id":10606,"date":"2026-08-22T16:56:15","date_gmt":"2026-08-22T08:56:15","guid":{"rendered":"\/jase\/?post_type=tkuisotope&#038;p=10606"},"modified":"2026-08-22T18:00:06","modified_gmt":"2026-08-22T10:00:06","slug":"jase-202611-34-060","status":"publish","type":"tkuisotope","link":"\/jase\/?tkuisotope=jase-202611-34-060","title":{"rendered":"Interaction-Aware Multimodal Spatiotemporal Learning for Fine-Grained Volleyball Training Action Recognition"},"content":{"rendered":"\n<div class=\"wp-block-tkuwpbs5-bs5-row row article-info\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=807\" data-type=\"page\" data-id=\"807\">2026<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder-open\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=9439\" data-type=\"page\" data-id=\"9439\">Volume 34<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-6 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div dv_publish\" data-aos=\"normal\"><div class=\"wp-block-post-date\"><time datetime=\"2026-08-22T16:56:15+08:00\">2026-08-22<\/time><\/div><\/div>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-row row\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-5 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div au-ol\" data-aos=\"normal\">\n<p>Jing Wang<sup>1<\/sup>, Haojuan Wang<sup>2<\/sup><a href=\"mailto:wang22007@outlook.com\"><i class=\"fa fa-envelope\"><\/i><\/a>, and Jinglong Geng<sup>3<\/sup><\/p>\n\n\n\n<p style=\"font-size:14px\"><sup>1<\/sup>Shanxi Vocational University of Engineering Science and Technology, Physical Education College, Shanxi Jinzhong 030619, China <\/p>\n\n\n\n<p style=\"font-size:14px\"><sup>2<\/sup>Shanxi University of Chinese Medicine, Shanxi Jinzhong 030619, China <\/p>\n\n\n\n<p style=\"font-size:14px\"><sup>3<\/sup>Taiyuan Normal University, Shanxi Jinzhong 030619, China<\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div\" style=\"margin-top:var(--wp--preset--spacing--40)\" data-aos=\"normal\">\n<p>Received: May 12, 2026<br>Accepted: July 07, 2026<br>Publication Date:&nbsp;August 22, 2026<\/p>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-7 align-self-start clk=\u5716\u7247\"><img decoding=\"async\" src=\"\/jase\/wp-content\/uploads\/2026\/08\/34_060.jpg\" class=\"img-fluid img-fluid mx-auto d-block\" alt=\"\u4e0a\u50b3\u5716\u7247\">\n\n\n<p class=\"has-text-align-center\">CNN-based spatial feature extraction pipeline&nbsp;<\/p>\n<\/div>\n<\/div>\n\n\n\n<p class=\"has-small-font-size\"><i class=\"fab fa-creative-commons\"><\/i>&nbsp;<strong>Copyright&nbsp;<\/strong>The Author(s). This is an open access article distributed under the terms of the&nbsp;<a rel=\"noreferrer noopener\" href=\"https:\/\/creativecommons.org\/licenses\/by\/4.0\/\" target=\"_blank\">Creative Commons Attribution&nbsp;License (CC BY 4.0)<\/a>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.<\/p>\n\n\n\n<p>Download Citation:\u00a0 <a href=\"\/jase\/wp-content\/uploads\/2026\/08\/V34.0060.txt\" data-type=\"attachment\" data-id=\"9812\" target=\"_blank\" rel=\"noreferrer noopener\">BibTeX <\/a>| <a rel=\"noreferrer noopener\" href=\"http:\/\/dx.doi.org\/10.6180\/jase.202611_34.060\" target=\"_blank\">http:\/\/dx.doi.org\/10.6180\/jase.202611_34.060<\/a>\u00a0\u00a0<\/p>\n\n\n\n<p class=\"btn btn-primary article-btn\"><a href=\"\/jase\/wp-content\/uploads\/2026\/08\/060_2026_1211_V34.pdf\" data-type=\"attachment\" data-id=\"10600\" target=\"_blank\" rel=\"noreferrer noopener\">Download PDF<\/a><\/p>\n\n\n\n<div style=\"height:24px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<p>To address the challenges of significant fine-grained action differences, complex multi-agent interactions, and strong temporal dynamic dependencies in volleyball scenarios, this paper proposes an action recognition method based on multimodal spatiotemporal fusion. First, convolutional neural networks are used to extract spatial features from images, and human pose information is combined to enhance the representation of fine-grained structural differences. Second, a hierarchical temporal modeling framework is constructed, using individual level LSTM and group-level LSTM to characterize the action evolution of individual players and multi-agent interaction relationships, respectively. Simultaneously, a People Pooling mechanism is introduced to adaptively fuse features from different subjects, highlighting key roles and modeling &#8220;ball-player-player&#8221; interactions. Finally, RGB and Skeleton multimodal information are fused to improve model robustness. Experimental results on the publicly available Volleyball dataset and a self-built dataset demonstrate that the proposed method outperforms several baseline methods in terms of accuracy, F1-score, and mAP, validating its effectiveness and generalization ability in fine-grained action recognition and complex scene modeling.<\/p>\n\n\n\n<p><em>Keywords:&nbsp;Volleyball motion recognition; fine-grained motion modeling; multi-agent interaction; temporal modeling; multimodal fusion<\/em><\/p>\n\n\n\n<div style=\"height:2rem\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div ref_ol\" data-aos=\"normal\">\n<div class=\"container\">\n<div id=\"model-response-message-contentr_53180f64958e6cc2\" class=\"markdown markdown-main-panel md-content enable-luminous-fast-follows enable-updated-hr-color stronger\" dir=\"ltr\" aria-busy=\"false\" aria-live=\"polite\">\n<ol>\n<li data-path-to-node=\"0\">[1] I. Zahra, Y. Wu, H. F. Alhasson, S. S. Alharbi, H. Aljuaid, A. Jalal, and H. Liu, (2025) &#8220;Dynamic graph neural networks for UAV-based group activity recognition in structured team sports&#8221; Frontiers in Neurorobotics 19: 1631998. DOI: 10.3389\/fnbot.2025.1631998.<\/li>\n<li data-path-to-node=\"0\">[2] A. Alqarafi and B. Almogadwy, (2025) &#8220;STRIKE-net: An explainable dynamic spatiotemporal graph-transformer network for fine-grained soccer action recognition&#8221; Applied Soft Computing: 114224. DOI: 10.1016\/j.asoc.2025.114224.<\/li>\n<li data-path-to-node=\"0\">[3] K. Seweryn, A. Wr\u00f3blewska, and S. \u0141ukasik, (2026) &#8220;Survey of action recognition, spotting, and spatiotemporal localization in soccer\u2014current trends and research perspectives&#8221; ACM Transactions on Intelligent Systems and Technology 17(2): 1\u201337. DOI: 10.1145\/3776541.<\/li>\n<li data-path-to-node=\"0\">[4] C. Zhang, G. Shan, J. Lim, and B. H. Roh, (2024) &#8220;Dynamic reinforcement learning for optimal go AI training: adaptive adjustment and optimization&#8221; IEEE Transactions on Consumer Electronics 71(1): 292\u2013302. DOI: 10.1109\/TCE.2024.3487141.<\/li>\n<li data-path-to-node=\"0\">[5] M. A. Hossen and P. E. Abas, (2025) &#8220;Machine learning for human activity recognition: State-of-the-art techniques and emerging trends&#8221; Journal of Imaging 11(3): 91. DOI: 10.3390\/jimaging11030091.<\/li>\n<li data-path-to-node=\"0\">[6] N. V. S. R. Chappa and K. Luu. &#8220;LiGAR: LiDAR-guided hierarchical transformer for multi-modal group activity recognition&#8221;. In: 2025 IEEE\/CVF Winter Conference on Applications of Computer Vision (WACV). 2025, 3035\u20133044. DOI: 10.1109\/WACV61041.2025.00300.<\/li>\n<li data-path-to-node=\"0\">[7] D. Karki, M. Ramazanova, A. Cioppa, S. Giancola, and B. Ghanem. &#8220;Pixels or positions? benchmarking modalities in group activity recognition&#8221;. In: Proceedings of the IEEE\/CVF Conference on Computer Vision and Pattern Recognition. 2026, 9998\u201310008. DOI: 10.48550\/arXiv.2511.12606.<\/li>\n<li data-path-to-node=\"0\">[8] L. Wang, P. Koniusz, and Y. Gao, (2025) &#8220;Video Understanding by Design: How Datasets Shape Architectures and Insights&#8221; arXiv preprint arXiv:2509.09151: DOI: 10.48550\/arXiv.2509.09151.<\/li>\n<li data-path-to-node=\"0\">[9] R. Zhang, Y. Wang, X. Hu, C. Mai, W. Liu, D. Xu, et al. &#8220;Beyond the Individual: Introducing Group Intention Forecasting with SHOT Dataset&#8221;. In: Proceedings of the 33rd ACM International Conference on Multimedia. 2025, 13002\u201313008. DOI: 10.1145\/3746027.3758248.<\/li>\n<li data-path-to-node=\"0\">[10] Y. Liu, F. Liu, L. Jiao, Q. Bao, L. Li, Y. Guo, and P. Chen, (2024) &#8220;A knowledge-based hierarchical causal inference network for video action recognition&#8221; IEEE Transactions on Multimedia 26: 9135\u20139149. DOI: 10.1109\/TMM.2024.3386339.<\/li>\n<li data-path-to-node=\"0\">[11] Z. Chen, W. Sun, Y. Tian, J. Jia, Z. Zhang, J. Wang, et al. &#8220;Gaia: Rethinking action quality assessment for AI-generated videos&#8221;. In: Advances in Neural Information Processing Systems. 37. 2024, 40111\u201340144. DOI: 10.52202\/079017-1267.<\/li>\n<li data-path-to-node=\"0\">[12] X. Wang, X. Lan, and W. Zhu. Video Grounding and Its Generalization: From ID and Task-specific Models to OOD and Large Foundation Models. Springer International Publishing AG, 2025. DOI: 10.1007\/978-3-031-94837-4.<\/li>\n<li data-path-to-node=\"0\">[13] Y. Wen, M. Liu, S. Wu, and B. Ding. &#8220;Chase: Learning convex hull adaptive shift for skeleton-based multi-entity action recognition&#8221;. In: Advances in Neural Information Processing Systems. 37. 2024, 9388\u20139420. DOI: 10.52202\/079017-0298.<\/li>\n<li data-path-to-node=\"0\">[14] B. Amirgaliyev, M. Mussabek, T. Rakhimzhanova, and A. Zhumadillayeva, (2025) &#8220;A review of machine learning and deep learning methods for person detection, tracking and identification, and face recognition with applications&#8221; Sensors 25(5): 1410. DOI: 10.3390\/s25051410.<\/li>\n<li data-path-to-node=\"0\">[15] K. Ashutosh and K. Grauman. &#8220;Learning skill-attributes for transferable assessment in video&#8221;. In: Advances in Neural Information Processing Systems. 38. 2026, 160403\u2013160430. DOI: 10.48550 \/ arXiv. 2511 . 13993.<\/li>\n<\/ol>\n<\/div>\n<\/div>\n<\/div>\n\n\n\n<p><\/p>\n","protected":false},"author":3,"template":"wp-custom-template-detail-4-aricles","meta":{"_uag_custom_page_level_css":""},"categories":[12,1682,6],"tags":[1844],"acf":[],"uagb_featured_image_src":[],"uagb_author_info":{"display_name":"\u6797\u923a\u6db5","author_link":"\/jase\/?author=3"},"uagb_comment_info":0,"uagb_excerpt":"&nbsp;Copyright&nbsp;The Author(s). This is an open access article distributed under the terms of the&nbsp;Creative Commons Attribution&nbsp;License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited. Download Citation:\u00a0 BibTeX | http:\/\/dx.doi.org\/10.6180\/jase.202611_34.060\u00a0\u00a0 Download PDF To address the challenges of significant fine-grained action differences, complex multi-agent&hellip;","_links":{"self":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope\/10606"}],"collection":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope"}],"about":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/types\/tkuisotope"}],"author":[{"embeddable":true,"href":"\/jase\/index.php?rest_route=\/wp\/v2\/users\/3"}],"wp:attachment":[{"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=10606"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=10606"},{"taxonomy":"post_tag","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=10606"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}