Journal of Applied Science and Engineering

Published by Tamkang University Press

ESCI jase impact factor scopus logo open access rate of Scopus journal

Optimizing a Reinforcement Learning Recommendation Algorithm for Personalized English Vocabulary Learning

Xin GUO1 and Si CHEN2

1Qinhuangdao Open University, Qinhuangdao City, Hebei Province, 066000, China

2Xingtai Open University, Xingtai City, Hebei Province, 054000, China

Received: September 21, 2025
Accepted: May 16, 2026
Publication Date: June 15, 2026

上傳圖片

Strategy oscillation amplitude time series curve

 Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.

Download Citation:  BibTeX | http://dx.doi.org/10.6180/jase.202610_33.004  

Download PDF

To address the challenges of personalized English vocabulary learning, this paper proposes a reinforcement learning recommendation algorithm optimization method that aims to balance dynamic user feature modeling with long-term policy stability. Traditional methods often face the problem of recommendations favoring short-term goals while neglecting long term learning effects. This is particularly true for English vocabulary learning, as the influence of the Ebbinghaus forgetting curve further illustrates this problem. To this end, this paper designs a dual-module collaborative framework that combines a dynamic feature encoder with a multi-time-scale policy network to optimize learning policies. User modeling uses a Transformer and GRU (Gate Recurrent Unit) architecture to construct long-term states, and the policy is optimized using the Deep Deterministic Policy Gradient (DDPG) framework. In terms of reward mechanism design, a multi-time-scale reward function is introduced, using a time decay factor to adjust the ratio of immediate and delayed rewards, balancing short-term memory and long-term knowledge accumulation. Experimental results demonstrate that the optimization method performs exceptionally well in the following areas: Top-10 accuracy reaches 88%; 7-day memory retention reaches 85% on the first day and remains stable; and in cold-start scenarios, the average first-time interaction coverage reaches 60%. Furthermore, the policy stability evaluation is only 0.02, validating the algorithm’s efficiency and stability. This paper not only demonstrates algorithmic innovation but also has
significant application value in the engineering implementation of personalized learning recommendation systems.

Keywords: reinforcement learning; personalized recommendation; English learning; time- scaled reward; deep deterministic policy gradient

  1. [1] X. Chen, D. Zou, H. Xie, and G. Cheng, (2021) “Twenty years of personalized language learning” Educational Technology & Society 24(1): 205–222.
  2. [2] J. Lahiassi, S. Aammou, and O. Warraki, (2023) “Enhancing personalized learning with a recommendation system in private online courses” Conhecimento & Diversidade 15(3): 176–189. DOI: 10.18316/rcd.v15i3.11144.
  3. [3] J.-W. Tzeng, N.-F. Huang, A.-C. Chuang, T.-W. Huang, and H.-Y. Chang, (2025) “Massive open online course recommendation system based on a reinforcement learning algorithm” Neural Computing and Applications 37(18): 11607–11618. DOI: 10.1007/s00521-023-08686-8.
  4. [4] Y. Deng, Y. Li, B. Ding, and W. Lam, (2022) “Leveraging long short-term user preference in conversational recommendation via multi-agent reinforcement learning” IEEE Transactions on Knowledge and Data Engineering 35(11): 11541–11555. DOI: 10.1109/TKDE.2022.3225109.
  5. [5] T. Liu, Q. Wu, L. Chang, and T. Gu, (2022) “A review of deep learning-based recommender system in e-learning environments” Artificial Intelligence Review 55(8): 5953–5980. DOI: 10.1007/s10462-022-10135-2.
  6. [6] J. Lahiassi, S. Aammou, and Y. Jdidou, (2025) “PERSONALIZED LEARNER RECOMMENDATIONS: ENHANCING GROUP DYNAMICS IN COLLABORATIVE LEARNING” Conhecimento & Diversidade 17(45): 591–607. DOI: 10.18316/rcd.v17i45.12504.
  7. [7] R. Fariani, K. Junus, and H. Santoso, (2023) “A systematic literature review on personalised learning in the higher education context” Technology, Knowledge and Learning 28(2): 449–476. DOI: 10.1007/s10758-022-09628-4.
  8. [8] J. Luo, F. Li, and J. Jiao, (2025) “A dynamic multiobjective recommendation method based on soft actor-critic with discrete actions” Journal of King Saud University – Computer and Information Sciences 37(1): 1. DOI: 10.1007/s44443-025-00016-3.
  9. [9] S. Cheng, Z. Wu, M. Qian, and W. Huang, (2024) “Point-of-interest recommendation based on bidirectional self-attention mechanism by fusing spatio-temporal preference” Multimedia Tools and Applications 83(9): 26333–26347. DOI: 10.1007/s11042-023-16542-z.
  10. [10] H. Wang and W. Fu, (2021) “Personalized learning resource recommendation method based on dynamic collaborative filtering” Mobile Networks and Applications 26(1): 473–487. DOI: 10.1007/s11036-020-01673-6.
  11. [11] X. Shi, Q. Liu, H. Xie, Y. Bai, and M. Shang, (2024) “Maximum entropy policy for long-term fairness in interactive recommender systems” IEEE Transactions on Services Computing 17(3): 1029–1043. DOI: 10.1109/TSC.2024.3349636.
  12. [12] N. Vedavathi and K. Anil Kumar, (2021) “An efficient e-learning recommendation system for user preferences using hybrid optimization algorithm” Soft Computing 25(14): 9377–9388. DOI: 10.1007/s00500-021-05753-x.
  13. [13] Z. Li and H. Wang, (2023) “Study on recommendation of personalized learning resources based on deep reinforcement learning” International Journal of Information and Communication Technology 23(4): 299–313. DOI: 10.1504/IJICT.2023.134832.
  14. [14] Y. Lin, Y. Liu, F. Lin, L. Zuo, P. Wu, W. Zeng, H. Chen, and C. Miao, (2023) “A survey on reinforcement learning for recommender systems” IEEE Transactions on Neural Networks and Learning Systems 35(10): 13164–13184. DOI: 10.1109/TNNLS.2023.3280161.
  15. [15] M. Mu and M. Yuan, (2024) “Research on a personalized learning path recommendation system based on cognitive graph with a cognitive graph” Interactive Learning Environments 32(8): 4237–4255. DOI: 10.1080/10494820.2023.2195446.
  16. [16] R. Guan, H. Pang, F. Giunchiglia, Y. Liang, and X. Feng, (2022) “Cross-domain meta-learner for cold-start recommendation” IEEE Transactions on Knowledge and Data Engineering 35(8): 7829–7843. DOI: 10.1109/TKDE.2022.3208005.
  17. [17] J. Meng, (2025) “AVAR-RL: Adaptive reinforcement learning approach for personalized English vocabulary acquisition” Discover Artificial Intelligence 5(1): 1–17. DOI: 10.1007/s44163-025-00584-3.
  18. [18] D. Chen, S. Yongchareon, E.-K. Lai, J. Yu, Q. Sheng, and Y. Li, (2022) “Transformer with bidirectional GRU for nonintrusive, sensor-based activity recognition in a multiresident environment” IEEE Internet of Things Journal 9(23): 23716–23727. DOI: 10.1109/JIOT.2022.3190307.
  19. [19] M. Afsar, T. Crump, and B. Far, (2022) “Reinforcement learning based recommender systems: A survey” ACM Computing Surveys 55(7): 1–38. DOI: 10.1145/3543846.
  20. [20] L. Xia, C. Huang, Y. Xu, and J. Pei, (2022) “Multi-behavior sequential recommendation with temporal graph transformer” IEEE Transactions on Knowledge and Data Engineering 35(6): 6099–6112. DOI: 10.1109/TKDE.2022.3175094.
  21. [21] R. Zhang, D. Zou, and H. Xie, (2022) “Spaced repetition for authentic mobile-assisted word learning: Nature, learner perceptions, and factors leading to positive perceptions” Computer Assisted Language Learning 35(9): 2593–2626. DOI: 10.1080/09588221.2021.1888752.
  22. [22] M. Walsh, M. Krusmark, T. Jastrembski, D. Hansen, K. Honn, and G. Gunzelmann, (2023) “Enhancing learning and retention through the distribution of practice repetitions across multiple sessions” Memory & Cognition 51(2): 455–472. DOI: 10.3758/s13421-022-01361-8.
  23. [23] X. Liu, M. Yu, C. Yang, L. Zhou, H. Wang, and H. Zhou, (2024) “Value distribution DDPG with dual-prioritized experience replay for coordinated control of coal-fired power generation systems” IEEE Transactions on Industrial Informatics 20(6): 8181–8194. DOI: 10.1109/TII.2024.3369712.
  24. [24] M. Daniel, A. Magassouba, M. Aranda, J. Ramón, R. Rodriguez, and Y. Mezouar, (2023) “Multi actor-critic DDPG for robot action space decomposition: A framework to control large 3D deformation of soft linear objects” IEEE Robotics and Automation Letters 9(2): 1318–1325. DOI: 10.1109/LRA.2023.3342672.
  25. [25] D. Dutta and S. Upreti, (2022) “A survey and comparative evaluation of actor-critic methods in process control” The Canadian Journal of Chemical Engineering 100(9): 2028–2056. DOI: 10.1002/cjce.24508.
  26. [26] A. Fakhrezi, G. Budiman, and D. Perdana, (2025) “Enhancing Soybean Fertilization Optimization with Prioritized Experience Replay and Noisy Networks in Deep Q-Networks” Jurnal Ilmiah Teknik Elektro Komputer Dan Informatika (JITEKI) 11(2): 154–168. DOI: 10.26555/jiteki.v11i2.30690.
  27. [27] D. Song, T. Ma, J. Shen, and F. Xu, (2024) “Incremental learning-based quantitative crack detection using prioritized experience replaying and layered importance sampling” IEEE Sensors Journal 24(15): 25132–25140. DOI: 10.1109/JSEN.2024.3415810.
  28. [28] Z. Lou, Y. Wang, S. Shan, K. Zhang, and H. Wei, (2024) “Balanced prioritized experience replay in off-policy reinforcement learning” Neural Computing and Applications 36(25): 15721–15737. DOI: 10.1007/s00521-024-09913-6.
  29. [29] V. Coscrato and D. Bridge, (2023) “Estimating and evaluating the uncertainty of rating predictions and top-n recommendations in recommender systems” ACM Transactions on Recommender Systems 1(2): 1–34. DOI: 10.1145/3584021.
  30. [30] W. Zhao, Z. Lin, Z. Feng, P. Wang, and J.-R. Wen, (2022) “A revisiting study of appropriate offline evaluation for top-N recommendation algorithms” ACM Transactions on Information Systems 41(2): 1–41. DOI: 10.1145/3545796.
  31. [31] F. Nielsen, (2022) “The Kullback–Leibler divergence between lattice Gaussian distributions” Journal of the Indian Institute of Science 102(4): 1177–1188. DOI: 10.1007/s41745-021-00279-5.
  32. [32] M. Rahad, R. Shabab, M. Ahammad, M. Reza, A. Karmaker, and M. Hossain, (2025) “KL-FedDis: A federated learning approach with distribution information sharing using Kullback-Leibler divergence for non-IID data” Neuroscience Informatics 5(1): 100182. DOI: 10.1016/j.neuri.2024.100182.