Journal of Applied Science and Engineering

Published by Tamkang University Press

ESCI jase impact factor scopus logo open access rate of Scopus journal

Dynamic Gesture Recognition Based on Deep Learning in Human-to-Computer Interfaces

Jing Yu1, Hang Li2, Shou-Lin Yin2, Qingwu Shi4 and Shahid Karim3

1Luxun Academy of Fine Arts, Shenyang 110034, P.R. China

2Software College, Shenyang Normal University, Shenyang 110034, P.R. China

3Institute of Image and Information Technology, Harbin Institute of Technology, Harbin 150000, P.R. China

4College of Information Science & Electronic Technique, Jiamusi University

Received: July 22, 2019
Accepted: October 19, 2019
Publication Date: May 10, 2026

上傳圖片

Part of the results: left segment result, right recognition result.

 Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.

Download Citation:  BibTeX | http://dx.doi.org/10.6180/jase.202003_23(1).0004  

Download PDF

Currently, gesture recognition provides a faster, simpler, convenient, effective and more natural way for human-computer interaction, which has been widely concerned. Gesture recognition plays an important role in real life. The manual feature extraction in traditional gesture recognition methods is time-consuming and strenuous. Moreover, in order to improve the accuracy of recognition, the quantity and quality of features to be extracted are required to be very high, which is a bottleneck for traditional gesture recognition methods. Therefore, we propose a deep learning method for dynamic gesture recognition in Human-to-Computer interfaces. An improved inverted residual network architecture is utilized as the basis of SSD (Single Shot MultiBox Detector) network for feature extraction. And the convolution structure of the auxiliary layer is predicted by using the inverse residual structure combining the cavity convolution. It uses multi-scale information, which can reduce the amount of calculation and parameters number. Transfer learning is used to optimize the trained network model so as to reduce the training time and make the model more convergent. Finally, experimental results show that the proposed method can recognize different gestures quickly and effectively.

Keywords: Gesture Recognition, Deep Learning, Human-to-Computer Interfaces, Feature Extraction

  1. [1] Yang, C., J. Long, M. A. Urbin, et al. (2018) Real-time myocontrol of a human–computer interface by paretic muscles after stroke, IEEE Transactions on Cognitive & Developmental Systems 10(4), 1126–1132. doi: 10. 1109/TCDS.2018.2830388
  2. [2] Mert, A., and A. Akan (2018) Emotion recognition from EEG signals by using multivariate empirical mode decomposition, Pattern Analysis & Applications 21(1), 81–89. doi: 10.1007/s10044-016-0567-6
  3. [3] Yu, J., H. Li, and S. L. Yin (2019) New intelligent interface study based on K-means gaze tracking, International Journal of Computational Science and Engineering 18(1), 12–20. doi: 10.1504/IJCSE.2019. 096971
  4. [4] Molchanov, P., S. Gupta, K. Kim, et al. (2015) Hand gesture recognition with 3D convolutional neural networks, Computer Vision & Pattern Recognition Workshops. doi: 10.1109/CVPRW.2015.7301342
  5. [5] Wilson, A. D., and A. F. Bobick (2016) Parametric hidden Markov models for gesture recognition, IEEE Trans. pattern Anal. & Mach. intell 21(9), 884 900. doi: 10.1109/34.790429
  6. [6] Caramiaux, B., N. Montecchio, and A. Tanaka (2014) Adaptive gesture recognition with variation estimation for interactive systems, Acm Transactions on Interactive Intelligent Systems 4(4), 1–34. doi: 10.1145/ 2643204
  7. [7] Gao, J., P. Li, and Z. K. Chen (2019) A canonical polyadic deep convolutional computation model for big data feature learning in Internet of Things, Future Generation Computer Systems. doi: 10.1016/j.future. 2019.04.048
  8. [8] Lin, T., H. Li, and S. L. Yin (2018) Modified pyramid dual tree direction filter-based image de-noising via curvature scale and non-local mean multi-grade remnant multi-grade remnant filter, International Journal of Communication Systems 31(16). doi: 10.1002/dac. 3486
  9. [9] Yin, S. L., and J. Bi (2019) Medical image annotation based on deep transfer learning, Journal of Applied Science and Engineering 22(2), 385–390. doi: 10. 6180/jase.201906_22(2).0020
  10. [10] Yin, S. L., Y. Zhang, and S. Karim (2018) Large scale remote sensing image segmentation based on fuzzy region competition and Gaussian mixture model, IEEE Access 6, 26069–26080. doi: 10.1109/ACCESS.2018. 2834960
  11. [11] Yin, S. L., Y. Zhang, and S. Karim (2019) Region search based on hybrid CNN in optical remote sensing images under cloud computing environment, International Journal of Distributed Sensor Networks 15(5). doi: 10.1177/1550147719852036
  12. [12] Ren, S., K. He, R. Girshick, et al. (2017) Faster RCNN: towards real-time object detection with region proposal networks, IEEE Transactions on Pattern Analysis & Machine Intelligence 39(6), 1137–1149. doi: 10.1109/TPAMI.2016.2577031
  13. [13] Li, J., H. C. Wong, S. L. Lo, et al. (2018) Multiple object detection by deformable part-based model and RCNN, IEEE Signal Processing Letters PP(99):1-1. doi: 10.1109/LSP.2017.2789325
  14. [14] Shen, J., J. Bu, B. Ju, et al. (2012) Refining Gaussian mixture model based on enhanced manifold learning, Neurocomputing 87(1), 19–25. doi: 10.1016/j.neucom. 2012.01.029
  15. [15] Bambach, S., S. Lee, D. J. Crandall, et al. (2015) Lending A hand: detecting hands and recognizing activities in complex egocentric interactions, 2015 IEEE International Conference on Computer Vision (ICCV). IEEE Computer Society. doi: 10.1109/ICCV.2015.226
  16. [16] Lepetit-Aimon, G., R. Duval, and F. Cheriet (2018) Large receptive field fully convolutional network for semantic segmentation of retinal vasculature in fundus images, International Workshop on Computational Pathology 201 209. doi: 10.1007/978-3-030-009496_24
  17. [17] Liu, W., D. Anguelov, D. Erhan, et al. (2016) SSD: single shot MultiBox detector, European Conference on Computer Vision. ECCV, 21 37. doi: 10.1007/9783-319-46448-0_2
  18. [18] Bambach, S., S. Lee, D. J. Crandall, et al. (2015) Lending A hand: detecting hands and recognizing activities in complex egocentric interactions, 2015 IEEE International Conference on Computer Vision (ICCV). IEEE Computer Society. doi: 10.1109/ICCV.2015.226
  19. [19] Zhou, Z., Z. Cao, and Y. Pi (2018) Dynamic gesture recognition with a Terahertz Radar based on range profile sequences and Doppler signatures, Sensors 18(1), 10. doi: 10.3390/s18010010
  20. [20] Verma, B., and A. Choudhary (2018) Framework for dynamic hand gesture recognition using Grassmann manifold for intelligent vehicles, Iet Intelligent Transport Systems 12(7), 721–729. doi: 10.1049/iet-its.2017. 0331
  21. [21] Zhang, Z., Z. Tian, and Z. Mu (2018) Latern: dynamic continuous hand gesture recognition using FMCW radar sensor, IEEE Sensors Journal 18(8), 1 1. doi: 10. 1109/JSEN.2018.2808688
  22. [22] Nguyen, X. S., L. Brun, O. Lezoray, et al. (2019) Skeleton-based hand gesture recognition by learning SPD matrices with neural networks, IEEE International Conference on Automatic Face & Gesture Recognition (FG). IEEE. doi: 10.1109/FG.2019.8756512