Journal of Applied Science and Engineering

Published by Tamkang University Press

ESCI jase impact factor scopus logo open access rate of Scopus journal

Improving the Embedded AI Enhanced Learning Environment for College English Teaching

Yuanyang Lei

The Department of General Education, Henan Medical College, Zhengzhou, 450000, China.

Received: April 29, 2026
Accepted: June 02, 2026
Publication Date: August 12, 2026

上傳圖片

SS-DRNN model’s Overview

 Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.

Download Citation:  BibTeX | http://dx.doi.org/10.6180/jase.202611_34.035  

Download PDF

To enhance the English learning environment for college students, this study explores the integration of Artificial Intelligence (AI), focusing on pronunciation improvement through an AI-powered system. Traditional teaching methods face challenges such as large class sizes and limited personalized feedback, which hinder effective pronunciation development. To address these limitations, an AI-driven framework is proposed to provide real-time, targeted feedback for correcting pronunciation errors. The study utilizes English speech data collected from university students with varying proficiency levels, including beginner, intermediate, and advanced learners. The dataset comprises monologues, dialogues, and pronunciation drills, capturing common errors such as misarticulations, vowel and consonant inconsistencies, and stress-related issues. The data are annotated with word-level transcriptions and error labels, along with acoustic features such as pitch, duration, formant frequencies, and Mel-frequency cepstral coefficients (MFCCs). For error detection and classification, a Salp Swarm Integrated Dense Recurrent Neural Network (SS-DRNN) model is employed. The model enables automatic identification and correction of pronunciation errors while delivering adaptive feedback. Experimental results demonstrate high performance, achieving 98.25% accuracy and strong F1-score, precision, and recall. The proposed system enhances learning efficiency, reduces teacher workload, and supports scalable, personalized English instruction in higher education environments.

Keywords: Learning Environment, College English Teaching, Mel-Frequency Cepstral Coefficients (MFCCs), English Speech Data, Feedback.

  1. [1] R. Prabhavalkar, T. Hori, T. N. Sainath, R. Schluter, and S. Watanabe, (2024) “End-to-End Speech Recognition: A Survey” IEEE/ACM Transactions on Audio, Speech, and Language Processing 32: 325–351. DOI: 10.1109/TASLP.2023.3328283.
  2. [2] X. Li and X. Huang, (2024) “Improvement and Optimization Method of College English Teaching Level Based on Convolutional Neural Network Model in an Embedded Systems Context” Computer-Aided Design and Applications 21(S8): 212–227. DOI: 10.14733/cadaps.2024.S8.212-227.
  3. [3] F. Wu, Y. Chen, and D. Han, (2022) “Development countermeasures of college English education based on deep learning and artificial intelligence” Mobile Information Systems 2022: 1–10. DOI: 10.1155/2022/8389800.
  4. [4] L. Huang, (2022) “An empirical study of integrating information technology in English teaching in artificial intelligence era” Scientific Programming 2022: 1–12. DOI: 10.1155/2022/6775097.
  5. [5] L. Geng, (2021) “Evaluation model of college English multimedia teaching effect based on deep convolutional neural networks” Mobile Information Systems 2021: 1–11. DOI: 10.1155/2021/1874584.
  6. [6] M. N. Daoud and M. A. Ben Messaoud. “Phoneme-level mispronunciation detection in Quranic recitation using Shallow Transformer”. In: Proceedings of the Third Arabic Natural Language Processing Conference. 2025, 457–463. DOI: 10.18653/v1/2025.arabicnlp-sharedtasks.63.
  7. [7] H. Li, C. Tang, X. Yue, and X. Li, (2025) “Sentence-level consistency of conformer-based pre-training distillation for Chinese speech recognition” Frontiers in Communications and Networks 6(1): 1662788–1662796. DOI: 10.3389/frcmn.2025.1662788.
  8. [8] K. Li, X. Qian, and H. Meng, (2016) “Mispronunciation detection and diagnosis in L2 English speech using multi-distribution deep neural networks” IEEE/ACM Transactions on Audio, Speech, and Language Processing 24(12): 2528–2539. DOI: 10.1109/TASLP.2016.2621675.
  9. [9] B. C. Yan, H. W. Wang, S. W. F. Jiang, F. A. Chao, and B. Chen. “Maximum F1-score training for end-to-end mispronunciation detection and diagnosis of L2 English speech”. In: IEEE International Conference on Multimedia and Expo (ICME). 2022, 1–6. DOI: 10.1109/ICME52920.2022.9858931.
  10. [10] H. Liu, (2021) “College oral English teaching reform driven by big data and deep neural network technology” Wireless Communications and Mobile Computing 2021: 1–10. DOI: 10.1155/2021/8389469.
  11. [11] L. Peng, Y. Gao, R. Bao, Y. Li, and J. Zhang, (2023) “End-to-End Mispronunciation Detection and Diagnosis Using Transfer Learning” Applied Sciences 13(11): 6793. DOI: 10.3390/app13116793.
  12. [12] Ş. S. Çalık, A. Küçükmanisa, and Z. H. Kilimci, (2024) “A novel framework for mispronunciation detection of Arabic phonemes using audio-oriented transformer models” Applied Acoustics 215: 109711. DOI: 10.1016/j.apacoust.2023.109711.
  13. [13] Z. Fan, X. Zhang, M. Huang, and Z. Bu, (2024) “Sampleformer: An efficient conformer-based neural network for automatic speech recognition” Intelligent Data Analysis 28(6): 1647–1659. DOI: 10.3233/IDA-230612.
  14. [14] H. Kheddar, M. Hemis, and Y. Himeur, (2024) “Automatic speech recognition using advanced deep learning approaches: A survey” Information Fusion 109: 102422. DOI: 10.1016/j.inffus.2024.102422.
  15. [15] M. A. H. Wadud, M. Alatiyyah, and M. F. Mridha, (2023) “Non-Autoregressive End-to-End Neural Modeling for Automatic Pronunciation Error Detection” Applied Sciences 13(1): 109. DOI: 10.3390/app13010109.