Jing Yu1, Hang Li2, Shou-Lin Yin2, Qingwu Shi4 and Shahid Karim3
1Luxun Academy of Fine Arts, Shenyang 110034, P.R. China
2Software College, Shenyang Normal University, Shenyang 110034, P.R. China
3Institute of Image and Information Technology, Harbin Institute of Technology, Harbin 150000, P.R. China
4College of Information Science & Electronic Technique, Jiamusi University
Received: July 22, 2019
Accepted: October 19, 2019
Publication Date: May 10, 2026
Part of the results: left segment result, right recognition result.
Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.
Download Citation: BibTeX | http://dx.doi.org/10.6180/jase.202003_23(1).0004
Currently, gesture recognition provides a faster, simpler, convenient, effective and more natural way for human-computer interaction, which has been widely concerned. Gesture recognition plays an important role in real life. The manual feature extraction in traditional gesture recognition methods is time-consuming and strenuous. Moreover, in order to improve the accuracy of recognition, the quantity and quality of features to be extracted are required to be very high, which is a bottleneck for traditional gesture recognition methods. Therefore, we propose a deep learning method for dynamic gesture recognition in Human-to-Computer interfaces. An improved inverted residual network architecture is utilized as the basis of SSD (Single Shot MultiBox Detector) network for feature extraction. And the convolution structure of the auxiliary layer is predicted by using the inverse residual structure combining the cavity convolution. It uses multi-scale information, which can reduce the amount of calculation and parameters number. Transfer learning is used to optimize the trained network model so as to reduce the training time and make the model more convergent. Finally, experimental results show that the proposed method can recognize different gestures quickly and effectively.
Keywords: Gesture Recognition, Deep Learning, Human-to-Computer Interfaces, Feature Extraction
- [1] Yang, C., J. Long, M. A. Urbin, et al. (2018) Real-time myocontrol of a human–computer interface by paretic muscles after stroke, IEEE Transactions on Cognitive & Developmental Systems 10(4), 1126–1132. doi: 10. 1109/TCDS.2018.2830388
- [2] Mert, A., and A. Akan (2018) Emotion recognition from EEG signals by using multivariate empirical mode decomposition, Pattern Analysis & Applications 21(1), 81–89. doi: 10.1007/s10044-016-0567-6
- [3] Yu, J., H. Li, and S. L. Yin (2019) New intelligent interface study based on K-means gaze tracking, International Journal of Computational Science and Engineering 18(1), 12–20. doi: 10.1504/IJCSE.2019. 096971
- [4] Molchanov, P., S. Gupta, K. Kim, et al. (2015) Hand gesture recognition with 3D convolutional neural networks, Computer Vision & Pattern Recognition Workshops. doi: 10.1109/CVPRW.2015.7301342
- [5] Wilson, A. D., and A. F. Bobick (2016) Parametric hidden Markov models for gesture recognition, IEEE Trans. pattern Anal. & Mach. intell 21(9), 884 900. doi: 10.1109/34.790429
- [6] Caramiaux, B., N. Montecchio, and A. Tanaka (2014) Adaptive gesture recognition with variation estimation for interactive systems, Acm Transactions on Interactive Intelligent Systems 4(4), 1–34. doi: 10.1145/ 2643204
- [7] Gao, J., P. Li, and Z. K. Chen (2019) A canonical polyadic deep convolutional computation model for big data feature learning in Internet of Things, Future Generation Computer Systems. doi: 10.1016/j.future. 2019.04.048
- [8] Lin, T., H. Li, and S. L. Yin (2018) Modified pyramid dual tree direction filter-based image de-noising via curvature scale and non-local mean multi-grade remnant multi-grade remnant filter, International Journal of Communication Systems 31(16). doi: 10.1002/dac. 3486
- [9] Yin, S. L., and J. Bi (2019) Medical image annotation based on deep transfer learning, Journal of Applied Science and Engineering 22(2), 385–390. doi: 10. 6180/jase.201906_22(2).0020
- [10] Yin, S. L., Y. Zhang, and S. Karim (2018) Large scale remote sensing image segmentation based on fuzzy region competition and Gaussian mixture model, IEEE Access 6, 26069–26080. doi: 10.1109/ACCESS.2018. 2834960
- [11] Yin, S. L., Y. Zhang, and S. Karim (2019) Region search based on hybrid CNN in optical remote sensing images under cloud computing environment, International Journal of Distributed Sensor Networks 15(5). doi: 10.1177/1550147719852036
- [12] Ren, S., K. He, R. Girshick, et al. (2017) Faster RCNN: towards real-time object detection with region proposal networks, IEEE Transactions on Pattern Analysis & Machine Intelligence 39(6), 1137–1149. doi: 10.1109/TPAMI.2016.2577031
- [13] Li, J., H. C. Wong, S. L. Lo, et al. (2018) Multiple object detection by deformable part-based model and RCNN, IEEE Signal Processing Letters PP(99):1-1. doi: 10.1109/LSP.2017.2789325
- [14] Shen, J., J. Bu, B. Ju, et al. (2012) Refining Gaussian mixture model based on enhanced manifold learning, Neurocomputing 87(1), 19–25. doi: 10.1016/j.neucom. 2012.01.029
- [15] Bambach, S., S. Lee, D. J. Crandall, et al. (2015) Lending A hand: detecting hands and recognizing activities in complex egocentric interactions, 2015 IEEE International Conference on Computer Vision (ICCV). IEEE Computer Society. doi: 10.1109/ICCV.2015.226
- [16] Lepetit-Aimon, G., R. Duval, and F. Cheriet (2018) Large receptive field fully convolutional network for semantic segmentation of retinal vasculature in fundus images, International Workshop on Computational Pathology 201 209. doi: 10.1007/978-3-030-009496_24
- [17] Liu, W., D. Anguelov, D. Erhan, et al. (2016) SSD: single shot MultiBox detector, European Conference on Computer Vision. ECCV, 21 37. doi: 10.1007/9783-319-46448-0_2
- [18] Bambach, S., S. Lee, D. J. Crandall, et al. (2015) Lending A hand: detecting hands and recognizing activities in complex egocentric interactions, 2015 IEEE International Conference on Computer Vision (ICCV). IEEE Computer Society. doi: 10.1109/ICCV.2015.226
- [19] Zhou, Z., Z. Cao, and Y. Pi (2018) Dynamic gesture recognition with a Terahertz Radar based on range profile sequences and Doppler signatures, Sensors 18(1), 10. doi: 10.3390/s18010010
- [20] Verma, B., and A. Choudhary (2018) Framework for dynamic hand gesture recognition using Grassmann manifold for intelligent vehicles, Iet Intelligent Transport Systems 12(7), 721–729. doi: 10.1049/iet-its.2017. 0331
- [21] Zhang, Z., Z. Tian, and Z. Mu (2018) Latern: dynamic continuous hand gesture recognition using FMCW radar sensor, IEEE Sensors Journal 18(8), 1 1. doi: 10. 1109/JSEN.2018.2808688
- [22] Nguyen, X. S., L. Brun, O. Lezoray, et al. (2019) Skeleton-based hand gesture recognition by learning SPD matrices with neural networks, IEEE International Conference on Automatic Face & Gesture Recognition (FG). IEEE. doi: 10.1109/FG.2019.8756512
