{"id":7893,"date":"2026-06-15T10:46:21","date_gmt":"2026-06-15T02:46:21","guid":{"rendered":"\/jase\/?post_type=tkuisotope&#038;p=7893"},"modified":"2026-06-19T19:20:12","modified_gmt":"2026-06-19T11:20:12","slug":"jase-202610-33-004","status":"publish","type":"tkuisotope","link":"\/jase\/?tkuisotope=jase-202610-33-004","title":{"rendered":"Optimizing a Reinforcement Learning Recommendation Algorithm for Personalized English Vocabulary Learning"},"content":{"rendered":"\n<div class=\"wp-block-tkuwpbs5-bs5-row row article-info\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=807\" data-type=\"page\" data-id=\"807\">2026<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder-open\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=7886\" data-type=\"page\" data-id=\"7886\">Volume 33<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-6 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div dv_publish\" data-aos=\"normal\"><div class=\"wp-block-post-date\"><time datetime=\"2026-06-15T10:46:21+08:00\">2026-06-15<\/time><\/div><\/div>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-row row\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-5 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div au-ol\" data-aos=\"normal\">\n<p>Xin GUO<sup>1<\/sup> and Si CHEN<sup>2<\/sup><a href=\"mailto:chensi202571@163.com\"><i class=\"fa fa-envelope\"><\/i><\/a><\/p>\n\n\n\n<p style=\"font-size:14px\"><sup>1<\/sup>Qinhuangdao Open University, Qinhuangdao City, Hebei Province, 066000, China<\/p>\n\n\n\n<p style=\"font-size:14px\"><sup>2<\/sup>Xingtai Open University, Xingtai City, Hebei Province, 054000, China<\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div\" style=\"margin-top:var(--wp--preset--spacing--40)\" data-aos=\"normal\">\n<p>Received: September 21, 2025<br>Accepted:&nbsp;May 16, 2026<br>Publication Date:&nbsp;June 15, 2026<\/p>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-7 align-self-start clk=\u5716\u7247\"><img decoding=\"async\" src=\"\/jase\/wp-content\/uploads\/2026\/06\/33_004.jpg\" class=\"img-fluid img-fluid mx-auto d-block\" alt=\"\u4e0a\u50b3\u5716\u7247\">\n\n\n<p class=\"has-text-align-center\">Strategy oscillation amplitude time series curve<\/p>\n<\/div>\n<\/div>\n\n\n\n<p class=\"has-small-font-size\"><i class=\"fab fa-creative-commons\"><\/i>&nbsp;<strong>Copyright&nbsp;<\/strong>The Author(s). This is an open access article distributed under the terms of the&nbsp;<a rel=\"noreferrer noopener\" href=\"https:\/\/creativecommons.org\/licenses\/by\/4.0\/\" target=\"_blank\">Creative Commons Attribution&nbsp;License (CC BY 4.0)<\/a>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.<\/p>\n\n\n\n<p>Download Citation:\u00a0 <a href=\"\/jase\/wp-content\/uploads\/2026\/06\/V33.0004.txt\" data-type=\"attachment\" data-id=\"8033\" target=\"_blank\" rel=\"noreferrer noopener\">BibTeX <\/a>| <a rel=\"noreferrer noopener\" href=\"http:\/\/dx.doi.org\/10.6180\/jase.202610_33.004\" target=\"_blank\">http:\/\/dx.doi.org\/10.6180\/jase.202610_33.004<\/a>\u00a0\u00a0<\/p>\n\n\n\n<p class=\"btn btn-primary article-btn\"><a href=\"\/jase\/wp-content\/uploads\/2026\/06\/004_2025_1262_V33.pdf\" data-type=\"attachment\" data-id=\"7907\" target=\"_blank\" rel=\"noreferrer noopener\">Download PDF<\/a><\/p>\n\n\n\n<div style=\"height:24px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<p>To address the challenges of personalized English vocabulary learning, this paper proposes a reinforcement learning recommendation algorithm optimization method that aims to balance dynamic user feature modeling with long-term policy stability. Traditional methods often face the problem of recommendations favoring short-term goals while neglecting long term learning effects. This is particularly true for English vocabulary learning, as the influence of the Ebbinghaus forgetting curve further illustrates this problem. To this end, this paper designs a dual-module collaborative framework that combines a dynamic feature encoder with a multi-time-scale policy network to optimize learning policies. User modeling uses a Transformer and GRU (Gate Recurrent Unit) architecture to construct long-term states, and the policy is optimized using the Deep Deterministic Policy Gradient (DDPG) framework. In terms of reward mechanism design, a multi-time-scale reward function is introduced, using a time decay factor to adjust the ratio of immediate and delayed rewards, balancing short-term memory and long-term knowledge accumulation. Experimental results demonstrate that the optimization method performs exceptionally well in the following areas: Top-10 accuracy reaches 88%; 7-day memory retention reaches 85% on the first day and remains stable; and in cold-start scenarios, the average first-time interaction coverage reaches 60%. Furthermore, the policy stability evaluation is only 0.02, validating the algorithm\u2019s efficiency and stability. This paper not only demonstrates algorithmic innovation but also has<br>significant application value in the engineering implementation of personalized learning recommendation systems.<\/p>\n\n\n\n<p><em>Keywords:&nbsp;<\/em>r<em>einforcement learning; personalized recommendation; English learning; time- scaled reward; deep deterministic policy gradient<\/em><\/p>\n\n\n\n<div style=\"height:2rem\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div ref_ol\" data-aos=\"normal\">\n<div class=\"container\">\n<div id=\"model-response-message-contentr_442dd220420d5a90\" class=\"markdown markdown-main-panel stronger enable-updated-hr-color\" dir=\"ltr\" aria-live=\"polite\" aria-busy=\"false\">\n<ol>\n<li data-path-to-node=\"0\">[1] X. Chen, D. Zou, H. Xie, and G. Cheng, (2021) \u201cTwenty years of personalized language learning\u201d <i data-path-to-node=\"1\" data-index-in-node=\"99\">Educational Technology &amp; Society<\/i> 24(1): 205\u2013222.<\/li>\n<li data-path-to-node=\"0\">[2] J. Lahiassi, S. Aammou, and O. Warraki, (2023) \u201cEnhancing personalized learning with a recommendation system in private online courses\u201d <i data-path-to-node=\"1\" data-index-in-node=\"288\">Conhecimento &amp; Diversidade<\/i> 15(3): 176\u2013189. DOI: 10.18316\/rcd.v15i3.11144.<\/li>\n<li data-path-to-node=\"0\">[3] J.-W. Tzeng, N.-F. Huang, A.-C. Chuang, T.-W. Huang, and H.-Y. Chang, (2025) \u201cMassive open online course recommendation system based on a reinforcement learning algorithm\u201d <i data-path-to-node=\"1\" data-index-in-node=\"538\">Neural Computing and Applications<\/i> 37(18): 11607\u201311618. DOI: 10.1007\/s00521-023-08686-8.<\/li>\n<li data-path-to-node=\"0\">[4] Y. Deng, Y. Li, B. Ding, and W. Lam, (2022) \u201cLeveraging long short-term user preference in conversational recommendation via multi-agent reinforcement learning\u201d <i data-path-to-node=\"1\" data-index-in-node=\"165\">IEEE Transactions on Knowledge and Data Engineering<\/i> 35(11): 11541\u201311555. DOI: 10.1109\/TKDE.2022.3225109.<\/li>\n<li data-path-to-node=\"0\">[5] T. Liu, Q. Wu, L. Chang, and T. Gu, (2022) \u201cA review of deep learning-based recommender system in e-learning environments\u201d <i data-path-to-node=\"1\" data-index-in-node=\"397\">Artificial Intelligence Review<\/i> 55(8): 5953\u20135980. DOI: 10.1007\/s10462-022-10135-2.<\/li>\n<li data-path-to-node=\"0\">[6] J. Lahiassi, S. Aammou, and Y. Jdidou, (2025) \u201cPERSONALIZED LEARNER RECOMMENDATIONS: ENHANCING GROUP DYNAMICS IN COLLABORATIVE LEARNING\u201d <i data-path-to-node=\"1\" data-index-in-node=\"620\">Conhecimento &amp; Diversidade<\/i> 17(45): 591\u2013607. DOI: 10.18316\/rcd.v17i45.12504.<\/li>\n<li data-path-to-node=\"0\">[7] R. Fariani, K. Junus, and H. Santoso, (2023) \u201cA systematic literature review on personalised learning in the higher education context\u201d <i data-path-to-node=\"1\" data-index-in-node=\"835\">Technology, Knowledge and Learning<\/i> 28(2): 449\u2013476. DOI: 10.1007\/s10758-022-09628-4.<\/li>\n<li data-path-to-node=\"0\">[8] J. Luo, F. Li, and J. Jiao, (2025) \u201cA dynamic multiobjective recommendation method based on soft actor-critic with discrete actions\u201d <i data-path-to-node=\"1\" data-index-in-node=\"1056\">Journal of King Saud University &#8211; Computer and Information Sciences<\/i> 37(1): 1. DOI: 10.1007\/s44443-025-00016-3.<\/li>\n<li data-path-to-node=\"0\">[9] S. Cheng, Z. Wu, M. Qian, and W. Huang, (2024) \u201cPoint-of-interest recommendation based on bidirectional self-attention mechanism by fusing spatio-temporal preference\u201d <i data-path-to-node=\"1\" data-index-in-node=\"1338\">Multimedia Tools and Applications<\/i> 83(9): 26333\u201326347. DOI: 10.1007\/s11042-023-16542-z.<\/li>\n<li data-path-to-node=\"0\">[10] H. Wang and W. Fu, (2021) \u201cPersonalized learning resource recommendation method based on dynamic collaborative filtering\u201d <i data-path-to-node=\"1\" data-index-in-node=\"1552\">Mobile Networks and Applications<\/i> 26(1): 473\u2013487. DOI: 10.1007\/s11036-020-01673-6.<\/li>\n<li data-path-to-node=\"0\">[11] X. Shi, Q. Liu, H. Xie, Y. Bai, and M. Shang, (2024) \u201cMaximum entropy policy for long-term fairness in interactive recommender systems\u201d <i data-path-to-node=\"1\" data-index-in-node=\"1775\">IEEE Transactions on Services Computing<\/i> 17(3): 1029\u20131043. DOI: 10.1109\/TSC.2024.3349636.<\/li>\n<li data-path-to-node=\"0\">[12] N. Vedavathi and K. Anil Kumar, (2021) \u201cAn efficient e-learning recommendation system for user preferences using hybrid optimization algorithm\u201d <i data-path-to-node=\"1\" data-index-in-node=\"2013\">Soft Computing<\/i> 25(14): 9377\u20139388. DOI: 10.1007\/s00500-021-05753-x.<\/li>\n<li data-path-to-node=\"0\">[13] Z. Li and H. Wang, (2023) \u201cStudy on recommendation of personalized learning resources based on deep reinforcement learning\u201d <i data-path-to-node=\"1\" data-index-in-node=\"2209\">International Journal of Information and Communication Technology<\/i> 23(4): 299\u2013313. DOI: 10.1504\/IJICT.2023.134832.<\/li>\n<li data-path-to-node=\"0\">[14] Y. Lin, Y. Liu, F. Lin, L. Zuo, P. Wu, W. Zeng, H. Chen, and C. Miao, (2023) \u201cA survey on reinforcement learning for recommender systems\u201d <i data-path-to-node=\"1\" data-index-in-node=\"2466\">IEEE Transactions on Neural Networks and Learning Systems<\/i> 35(10): 13164\u201313184. DOI: 10.1109\/TNNLS.2023.3280161.<\/li>\n<li data-path-to-node=\"0\">[15] M. Mu and M. Yuan, (2024) \u201cResearch on a personalized learning path recommendation system based on cognitive graph with a cognitive graph\u201d <i data-path-to-node=\"1\" data-index-in-node=\"2722\">Interactive Learning Environments<\/i> 32(8): 4237\u20134255. DOI: 10.1080\/10494820.2023.2195446.<\/li>\n<li data-path-to-node=\"0\">[16] R. Guan, H. Pang, F. Giunchiglia, Y. Liang, and X. Feng, (2022) \u201cCross-domain meta-learner for cold-start recommendation\u201d <i data-path-to-node=\"1\" data-index-in-node=\"2937\">IEEE Transactions on Knowledge and Data Engineering<\/i> 35(8): 7829\u20137843. DOI: 10.1109\/TKDE.2022.3208005.<\/li>\n<li data-path-to-node=\"0\">[17] J. Meng, (2025) \u201cAVAR-RL: Adaptive reinforcement learning approach for personalized English vocabulary acquisition\u201d <i data-path-to-node=\"1\" data-index-in-node=\"3160\">Discover Artificial Intelligence<\/i> 5(1): 1\u201317. DOI: 10.1007\/s44163-025-00584-3.<\/li>\n<li data-path-to-node=\"0\">[18] D. Chen, S. Yongchareon, E.-K. Lai, J. Yu, Q. Sheng, and Y. Li, (2022) \u201cTransformer with bidirectional GRU for nonintrusive, sensor-based activity recognition in a multiresident environment\u201d <i data-path-to-node=\"1\" data-index-in-node=\"3434\">IEEE Internet of Things Journal<\/i> 9(23): 23716\u201323727. DOI: 10.1109\/JIOT.2022.3190307.<\/li>\n<li data-path-to-node=\"0\">[19] M. Afsar, T. Crump, and B. Far, (2022) \u201cReinforcement learning based recommender systems: A survey\u201d <i data-path-to-node=\"1\" data-index-in-node=\"3623\">ACM Computing Surveys<\/i> 55(7): 1\u201338. DOI: 10.1145\/3543846.<\/li>\n<li data-path-to-node=\"0\">[20] L. Xia, C. Huang, Y. Xu, and J. Pei, (2022) \u201cMulti-behavior sequential recommendation with temporal graph transformer\u201d <i data-path-to-node=\"1\" data-index-in-node=\"3804\">IEEE Transactions on Knowledge and Data Engineering<\/i> 35(6): 6099\u20136112. DOI: 10.1109\/TKDE.2022.3175094.<\/li>\n<li data-path-to-node=\"0\">[21] R. Zhang, D. Zou, and H. Xie, (2022) \u201cSpaced repetition for authentic mobile-assisted word learning: Nature, learner perceptions, and factors leading to positive perceptions\u201d <i data-path-to-node=\"1\" data-index-in-node=\"4086\">Computer Assisted Language Learning<\/i> 35(9): 2593\u20132626. DOI: 10.1080\/09588221.2021.1888752.<\/li>\n<li data-path-to-node=\"0\">[22] M. Walsh, M. Krusmark, T. Jastrembski, D. Hansen, K. Honn, and G. Gunzelmann, (2023) \u201cEnhancing learning and retention through the distribution of practice repetitions across multiple sessions\u201d <i data-path-to-node=\"1\" data-index-in-node=\"4375\">Memory &amp; Cognition<\/i> 51(2): 455\u2013472. DOI: 10.3758\/s13421-022-01361-8.<\/li>\n<li data-path-to-node=\"0\">[23] X. Liu, M. Yu, C. Yang, L. Zhou, H. Wang, and H. Zhou, (2024) \u201cValue distribution DDPG with dual-prioritized experience replay for coordinated control of coal-fired power generation systems\u201d <i data-path-to-node=\"1\" data-index-in-node=\"196\">IEEE Transactions on Industrial Informatics<\/i> 20(6): 8181\u20138194. DOI: 10.1109\/TII.2024.3369712.<\/li>\n<li data-path-to-node=\"0\">[24] M. Daniel, A. Magassouba, M. Aranda, J. Ram\u00f3n, R. Rodriguez, and Y. Mezouar, (2023) \u201cMulti actor-critic DDPG for robot action space decomposition: A framework to control large 3D deformation of soft linear objects\u201d <i data-path-to-node=\"1\" data-index-in-node=\"509\">IEEE Robotics and Automation Letters<\/i> 9(2): 1318\u20131325. DOI: 10.1109\/LRA.2023.3342672.<\/li>\n<li data-path-to-node=\"0\">[25] D. Dutta and S. Upreti, (2022) \u201cA survey and comparative evaluation of actor-critic methods in process control\u201d <i data-path-to-node=\"1\" data-index-in-node=\"711\">The Canadian Journal of Chemical Engineering<\/i> 100(9): 2028\u20132056. DOI: 10.1002\/cjce.24508.<\/li>\n<li data-path-to-node=\"0\">[26] A. Fakhrezi, G. Budiman, and D. Perdana, (2025) \u201cEnhancing Soybean Fertilization Optimization with Prioritized Experience Replay and Noisy Networks in Deep Q-Networks\u201d <i data-path-to-node=\"1\" data-index-in-node=\"973\">Jurnal Ilmiah Teknik Elektro Komputer Dan Informatika (JITEKI)<\/i> 11(2): 154\u2013168. DOI: 10.26555\/jiteki.v11i2.30690.<\/li>\n<li data-path-to-node=\"0\">[27] D. Song, T. Ma, J. Shen, and F. Xu, (2024) \u201cIncremental learning-based quantitative crack detection using prioritized experience replaying and layered importance sampling\u201d <i data-path-to-node=\"1\" data-index-in-node=\"1263\">IEEE Sensors Journal<\/i> 24(15): 25132\u201325140. DOI: 10.1109\/JSEN.2024.3415810.<\/li>\n<li data-path-to-node=\"0\">[28] Z. Lou, Y. Wang, S. Shan, K. Zhang, and H. Wei, (2024) \u201cBalanced prioritized experience replay in off-policy reinforcement learning\u201d <i data-path-to-node=\"1\" data-index-in-node=\"1475\">Neural Computing and Applications<\/i> 36(25): 15721\u201315737. DOI: 10.1007\/s00521-024-09913-6.<\/li>\n<li data-path-to-node=\"0\">[29] V. Coscrato and D. Bridge, (2023) \u201cEstimating and evaluating the uncertainty of rating predictions and top-n recommendations in recommender systems\u201d <i data-path-to-node=\"1\" data-index-in-node=\"1717\">ACM Transactions on Recommender Systems<\/i> 1(2): 1\u201334. DOI: 10.1145\/3584021.<\/li>\n<li data-path-to-node=\"0\">[30] W. Zhao, Z. Lin, Z. Feng, P. Wang, and J.-R. Wen, (2022) \u201cA revisiting study of appropriate offline evaluation for top-N recommendation algorithms\u201d <i data-path-to-node=\"1\" data-index-in-node=\"1944\">ACM Transactions on Information Systems<\/i> 41(2): 1\u201341. DOI: 10.1145\/3545796.<\/li>\n<li data-path-to-node=\"0\">[31] F. Nielsen, (2022) \u201cThe Kullback\u2013Leibler divergence between lattice Gaussian distributions\u201d <i data-path-to-node=\"1\" data-index-in-node=\"2116\">Journal of the Indian Institute of Science<\/i> 102(4): 1177\u20131188. DOI: 10.1007\/s41745-021-00279-5.<\/li>\n<li data-path-to-node=\"0\">[32] M. Rahad, R. Shabab, M. Ahammad, M. Reza, A. Karmaker, and M. Hossain, (2025) \u201cKL-FedDis: A federated learning approach with distribution information sharing using Kullback-Leibler divergence for non-IID data\u201d <i data-path-to-node=\"1\" data-index-in-node=\"2426\">Neuroscience Informatics<\/i> 5(1): 100182. DOI: 10.1016\/j.neuri.2024.100182.<\/li>\n<\/ol>\n<\/div>\n<\/div>\n<\/div>\n\n\n\n<p><\/p>\n","protected":false},"author":3,"template":"wp-custom-template-detail-4-aricles","meta":{"_uag_custom_page_level_css":""},"categories":[12,1483,6],"tags":[1487],"acf":[],"uagb_featured_image_src":[],"uagb_author_info":{"display_name":"\u6797\u923a\u6db5","author_link":"\/jase\/?author=3"},"uagb_comment_info":0,"uagb_excerpt":"&nbsp;Copyright&nbsp;The Author(s). This is an open access article distributed under the terms of the&nbsp;Creative Commons Attribution&nbsp;License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited. Download Citation:\u00a0 BibTeX | http:\/\/dx.doi.org\/10.6180\/jase.202610_33.004\u00a0\u00a0 Download PDF To address the challenges of personalized English vocabulary learning, this paper&hellip;","_links":{"self":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope\/7893"}],"collection":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope"}],"about":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/types\/tkuisotope"}],"author":[{"embeddable":true,"href":"\/jase\/index.php?rest_route=\/wp\/v2\/users\/3"}],"wp:attachment":[{"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=7893"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=7893"},{"taxonomy":"post_tag","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=7893"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}