{"id":9680,"date":"2026-08-05T21:52:34","date_gmt":"2026-08-05T13:52:34","guid":{"rendered":"\/jase\/?post_type=tkuisotope&#038;p=9680"},"modified":"2026-08-06T23:01:28","modified_gmt":"2026-08-06T15:01:28","slug":"jase-202611-34-010","status":"publish","type":"tkuisotope","link":"\/jase\/?tkuisotope=jase-202611-34-010","title":{"rendered":"Design and Empirical Study of a Personalized Training Path Optimization Algorithm Based on Artificial Intelligence Reinforcement Learning"},"content":{"rendered":"\n<div class=\"wp-block-tkuwpbs5-bs5-row row article-info\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=807\" data-type=\"page\" data-id=\"807\">2026<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder-open\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=9439\" data-type=\"page\" data-id=\"9439\">Volume 34<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-6 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div dv_publish\" data-aos=\"normal\"><div class=\"wp-block-post-date\"><time datetime=\"2026-08-05T21:52:34+08:00\">2026-08-05<\/time><\/div><\/div>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-row row\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-5 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div au-ol\" data-aos=\"normal\">\n<p>Yi Zhexue<sup>1<\/sup>, Lin Ze<sup>2<\/sup>, and Chen Feng<sup>2<\/sup><a href=\"mailto:19066196@zjcst.edu.cn\"><i class=\"fa fa-envelope\"><\/i><\/a><\/p>\n\n\n\n<p style=\"font-size:14px\"><sup>1<\/sup>Continuing Education College, Zhejiang College of Security Technology; Wenzhou, Zhejiang Province 325000, China<\/p>\n\n\n\n<p style=\"font-size:14px\"><sup>2<\/sup>College of AI, Zhejiang College of Security Technology; Wenzhou, Zhejiang Province 325000, China<\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div\" style=\"margin-top:var(--wp--preset--spacing--40)\" data-aos=\"normal\">\n<p>Received: April 13, 2026<br>Accepted:&nbsp;June 12, 2026<br>Publication Date:&nbsp;August 05, 2026<\/p>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-7 align-self-start clk=\u5716\u7247\"><img decoding=\"async\" src=\"\/jase\/wp-content\/uploads\/2026\/08\/34_010.jpg\" class=\"img-fluid img-fluid mx-auto d-block\" alt=\"\u4e0a\u50b3\u5716\u7247\">\n\n\n<p class=\"has-text-align-center\">Comprehensive&nbsp;Reinforcement&nbsp;Learning Performance&nbsp;Metrics<\/p>\n<\/div>\n<\/div>\n\n\n\n<p class=\"has-small-font-size\"><i class=\"fab fa-creative-commons\"><\/i>&nbsp;<strong>Copyright&nbsp;<\/strong>The Author(s). This is an open access article distributed under the terms of the&nbsp;<a rel=\"noreferrer noopener\" href=\"https:\/\/creativecommons.org\/licenses\/by\/4.0\/\" target=\"_blank\">Creative Commons Attribution&nbsp;License (CC BY 4.0)<\/a>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.<\/p>\n\n\n\n<p>Download Citation:\u00a0 <a href=\"\/jase\/wp-content\/uploads\/2026\/08\/V34.0010.txt\" data-type=\"attachment\" data-id=\"9764\" target=\"_blank\" rel=\"noreferrer noopener\">BibTeX <\/a>| <a rel=\"noreferrer noopener\" href=\"http:\/\/dx.doi.org\/10.6180\/jase.202611_34.010\" target=\"_blank\">http:\/\/dx.doi.org\/10.6180\/jase.202611_34.010<\/a>\u00a0\u00a0<\/p>\n\n\n\n<p class=\"btn btn-primary article-btn\"><a href=\"\/jase\/wp-content\/uploads\/2026\/08\/010_2026_0849_V34.pdf\" data-type=\"attachment\" data-id=\"9628\" target=\"_blank\" rel=\"noreferrer noopener\">Download PDF<\/a><\/p>\n\n\n\n<div style=\"height:24px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<p>Intelligent personalized learning systems must continuously adapt instructional strategies based on students\u2019 progress and engagement levels. However, designing adaptive curricula remains challenging, as traditional supervised learning and standard reinforcement learning approaches often fail to effectively integrate mastery assessment with long-term policy optimization. This study proposes a computational framework for personalized training path optimization that integrates mastery modelling with deep reinforcement learning to improve algorithmic efficiency, adaptive policy convergence, and long-term educational decision-making under dynamic learner-state transitions. The optimization problem is formulated as a Markov Decision Process and optimized using a Deep Q-Network (DQN) with experience replay and target-network stabilization to enhance computational efficiency, policy stability, and convergence performance across varying experimental learning conditions. (\u03b1 = 0.01, \u03b3 = 0.95, batch size = 64, replay buffer = 50,000,\u03b5 decayed from 1.0 to 0.01 over 1000 episodes). Experimental results demonstrate that the proposed framework achieves an accuracy of 91.38% and a success rate of 89.56%, outperforming Q-Learning (78.42%) and Policy Gradient (84.67%) methods. The model exhibits stable convergence, with episode rewards increasing from 43.97 to 89.54 and Q-loss decreasing from 15.1682 to 0.2533. These results confirm that integrating probabilistic learner-state modelling with deep reinforcement learning provides a computationally efficient and analytically robust solution for scalable personalized education optimization.<\/p>\n\n\n\n<p><em>Keywords:&nbsp;Personalized Learning, Deep Reinforcement Learning, Bayesian Knowledge Tracing, Deep Q-Network, Adaptive Education<\/em><\/p>\n\n\n\n<div style=\"height:2rem\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div ref_ol\" data-aos=\"normal\">\n<ol>\n<li>[1] L.-Y. Zhou and Y.-Y. Wang, (2025) \u201cSimulation of Personalized English Learning Path Recommendation System Based on Knowledge Graph and Deep Reinforcement Learning\u201d Scientific Reports 15(1): 34554. DOI: 10.1038\/s41598-025-17918-x.<\/li>\n<li>[2] S. Amin, M. I. Uddin, A. A. Alarood, W. K. Mashwani, A. Alzahrani, and A. O. Alzahrani, (2023) \u201cSmart E-Learning Framework for Personalized Adaptive Learning and Sequential Path Recommendations Using Reinforcement Learning\u201d IEEE Access 11: 89769\u201389790. DOI: 10.1109\/ACCESS.2023.3305584.<\/li>\n<li>[3] Y. Ma, L. Wang, J. Zhang, F. Liu, and Q. Jiang, (2023) \u201cA Personalized Learning Path Recommendation Method Incorporating Multi-Algorithm\u201d Applied Sciences 13(10): 5946. DOI: 10.3390\/app13105946.<\/li>\n<li>[4] B. Jiang et al., (2022) \u201cData-Driven Personalized Learning Path Planning Based on Cognitive Diagnostic Assessments in MOOCs\u201d Applied Sciences 12(8): 3982. DOI: 10.3390\/app12083982.<\/li>\n<li>[5] Q. Yang and C. Liang, (2025) \u201cA Second-Classroom Personalized Learning Path Recommendation System Based on Large Language Model Technology\u201d Applied Sciences 15(14): 7655. DOI: 10.3390\/app15147655.<\/li>\n<li>[6] O. Bulut, J. Shin, S. N. Yildirim-Erbasli, G. Gorgun, and Z. A. Pardos, (2023) \u201cAn Introduction to Bayesian Knowledge Tracing with pyBKT\u201d Psych 5(3): 770\u2013786. DOI: 10.3390\/psych5030050.<\/li>\n<li>[7] S.-Y. Jeong and Y.-K. Kim, (2023) \u201cDeep Learning-Based Context-Aware Recommender System Considering Change in Preference\u201d Electronics 12(10): 2337. DOI: 10.3390\/electronics12102337.<\/li>\n<li>[8] M. Barone, M. Naeem, M. Ciaschi, G. Tretola, and A. Coronato, (2025) \u201cAI-Based Intelligent System for Personalized Examination Scheduling\u201d Technologies 13(11): 518. DOI: 10.3390\/technologies13110518.<\/li>\n<li>[9] J. P. L\u00f3pez-Goyez, A. Gonz\u00e1lez-Briones, and Y. Demazeau, (2026) \u201cAn Adaptive Multi-Agent Architecture with Reinforcement Learning and Generative AI for Intelligent Tutoring Systems: A Moodle-Based Case Study\u201d Applied Sciences 16(3): 1323. DOI: 10.3390\/app16031323.<\/li>\n<li>[10] A. Darvish and M. Sepehri, (2025) \u201cArtificial Intelligence-Driven Project Portfolio Optimization Under Deep Uncertainty Using Adaptive Reinforcement Learning\u201d Applied Sciences 15(23): 12713. DOI: 10.3390\/app152312713.<\/li>\n<li>[11] S. Saleem, M. U. Aziz, M. J. Iqbal, and S. Abbas, (2025) \u201cAI in Education: Personalized Learning Systems and Their Impact on Student Performance and Engagement\u201d Critical Review of Social Science Studies 3(1): 2445\u20132459. DOI: 10.59075\/c35qa453.<\/li>\n<li>[12] C. Y. Zaharuddin and G. Yao, (2024) \u201cEnhancing Student Engagement with AI-Driven Personalized Learning Systems\u201d International Transactions on Education Technology 3(1): 1\u20138. DOI: 10.33050\/itee.v3i1.662.<\/li>\n<li>[13] C. F. Mahmoud and J. T. S\u00f8rensen, (2024) \u201cArtificial Intelligence in Personalized Learning with a Focus on Current Developments and Future Prospects\u201d Research Advances in Education 3(8): 25\u201331. DOI: 10.56397\/RAE.2024.08.04.<\/li>\n<li>[14] C. C. Y. Yang and H. Ogata, (2023) \u201cPersonalized Learning Analytics Intervention Approach for Enhancing Student Learning Achievement and Behavioral Engagement in Blended Learning\u201d Education and Information Technologies 28(3): 2509\u20132528. DOI: 10.1007\/s10639-022-11291-2.<\/li>\n<li>[15] S. K. Sahu, A. Mokhade, and N. D. Bokde, (2023) \u201cAn Overview of Machine Learning, Deep Learning, and Reinforcement Learning-Based Techniques in Quantitative Finance: Recent Progress and Challenges\u201d Applied Sciences 13(3): 1956. DOI: 10.3390\/app13031956.<\/li>\n<\/ol>\n<\/div>\n\n\n\n<p><\/p>\n","protected":false},"author":3,"template":"wp-custom-template-detail-4-aricles","meta":{"_uag_custom_page_level_css":""},"categories":[12,1682,6],"tags":[1692],"acf":[],"uagb_featured_image_src":[],"uagb_author_info":{"display_name":"\u6797\u923a\u6db5","author_link":"\/jase\/?author=3"},"uagb_comment_info":0,"uagb_excerpt":"&nbsp;Copyright&nbsp;The Author(s). This is an open access article distributed under the terms of the&nbsp;Creative Commons Attribution&nbsp;License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited. Download Citation:\u00a0 BibTeX | http:\/\/dx.doi.org\/10.6180\/jase.202611_34.010\u00a0\u00a0 Download PDF Intelligent personalized learning systems must continuously adapt instructional strategies based on&hellip;","_links":{"self":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope\/9680"}],"collection":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope"}],"about":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/types\/tkuisotope"}],"author":[{"embeddable":true,"href":"\/jase\/index.php?rest_route=\/wp\/v2\/users\/3"}],"wp:attachment":[{"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=9680"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=9680"},{"taxonomy":"post_tag","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=9680"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}