{"id":3031,"date":"2026-04-09T23:23:52","date_gmt":"2026-04-09T15:23:52","guid":{"rendered":"https:\/\/iweb20wp-b205b.url.tku.edu.tw\/jase\/?post_type=tkuisotope&#038;p=3031"},"modified":"2026-06-09T21:36:23","modified_gmt":"2026-06-09T13:36:23","slug":"efficient-hindsight-experience-replay-with-transformed-data-augmentation","status":"publish","type":"tkuisotope","link":"\/jase\/?tkuisotope=efficient-hindsight-experience-replay-with-transformed-data-augmentation","title":{"rendered":"Efficient Hindsight Experience Replay with Transformed Data Augmentation"},"content":{"rendered":"\n<div class=\"wp-block-tkuwpbs5-bs5-row row article-info\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=2961\" data-type=\"page\" data-id=\"807\">2024<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-3 align-self-start\">\n<p><i class=\"fa fa-folder-open\" aria-hidden=\"true\"><\/i>&nbsp;<a href=\"\/jase\/?page_id=3017\" data-type=\"page\" data-id=\"1055\">Volume 27, Issue 2<\/a><\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-6 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div dv_publish\" data-aos=\"normal\"><div class=\"wp-block-post-date\"><time datetime=\"2026-04-09T23:23:52+08:00\">2026-04-09<\/time><\/div><\/div>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-row row\">\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-5 align-self-start\">\n<div class=\"wp-block-tkuwpbs5-bs5-div au-ol\" data-aos=\"normal\">\n<p>Jiazheng Sun and Weiguang Li<a href=\"mailto:1035805780@qq.com\"><i class=\"fa fa-envelope\"><\/i><\/a><\/p>\n\n\n\n<p style=\"font-size:14px\">School of Mechanical and Automotive Engineering, South China University of Technology Guangzhou, Guangdong, China<\/p>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div\" style=\"margin-top:var(--wp--preset--spacing--40)\" data-aos=\"normal\">\n<p>Received:\u00a0February 27, 2022<br>Accepted:\u00a0April 6, 2023<br>Publication Date:\u00a0April 9, 2026<\/p>\n<\/div>\n<\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-column col-md-7 align-self-start clk=\u5716\u7247\"><img decoding=\"async\" src=\"\/jase\/wp-content\/uploads\/2026\/04\/27_02_11.jpg\" class=\"img-fluid img-fluid mx-auto d-block\" alt=\"\u4e0a\u50b3\u5716\u7247\">\n\n\n<p class=\"has-text-align-center img_caption\">FetchReach result<\/p>\n<\/div>\n<\/div>\n\n\n\n<p class=\"has-small-font-size\"><i class=\"fab fa-creative-commons\"><\/i>&nbsp;<strong>Copyright&nbsp;<\/strong>The Author(s). This is an open access article distributed under the terms of the&nbsp;<a rel=\"noreferrer noopener\" href=\"https:\/\/creativecommons.org\/licenses\/by\/4.0\/\" target=\"_blank\">Creative Commons Attribution&nbsp;License (CC BY 4.0)<\/a>, which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.<\/p>\n\n\n\n<p>Download Citation:\u00a0 <a rel=\"noreferrer noopener\" href=\"\/jase\/wp-content\/uploads\/2026\/01\/jase-202509-28-09-0006.pdf\" data-type=\"link\" data-id=\"\/jase\/wp-content\/uploads\/2026\/01\/jase-202509-28-09-0006.pdf\" target=\"_blank\">BibTeX <\/a>| <a href=\"http:\/\/dx.doi.org\/10.6180\/jase.202402_27(2).0011\" target=\"_blank\" rel=\"noreferrer noopener\">http:\/\/dx.doi.org\/10.6180\/jase.202402_27(2).0011<\/a>\u00a0\u00a0<\/p>\n\n\n\n<p class=\"btn btn-primary article-btn\"><a href=\"\/jase\/wp-content\/uploads\/2026\/04\/11_2023_0210_V27i2.pdf\" data-type=\"attachment\" data-id=\"3046\" target=\"_blank\" rel=\"noreferrer noopener\">Download PDF<\/a><\/p>\n\n\n\n<div style=\"height:24px\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<p>Motion control of robots is a high-dimensional, nonlinear control problem that is often difficult to handle using traditional dynamical path planning means. Reinforcement learning is currently an effective means to solve robot motion control problems, but reinforcement learning has disadvantages such as high number of trials and errors and sparse rewards, which restrict the application efficiency of reinforcement learning. The Hindsight Experience Replay(HER) algorithm is a reinforcement learning algorithm that solves the reward sparsity problem by constructing virtual target values. However, the HER algorithm still suffers from the problem of long time in the early stage of training, and there is still room for improving its sample utilization efficiency. Augmentation by existing data to improve training efficiency has been widely used in supervised learning, but is less applied in the field of reinforcement learning. In this paper, we propose the Hindsight Experience Replay with Transformed Data Augmentation (TDAHER) algorithm by constructing a transformed data augmentation method for reinforcement learning samples, combined with the HER algorithm. And in order to solve the problem of the accuracy of the augmented samples in the later stage of training, the decaying participation factor method is introduced. After the comparison of four simulated robot control tasks, it is proved that the algorithm can effectively improve the training efficiency of reinforcement learning.<\/p>\n\n\n\n<p><em>Keywords:\u00a0Reinforcement learning; Machine learning; Motion control; Data augmentation; component;<\/em><\/p>\n\n\n\n<div style=\"height:2rem\" aria-hidden=\"true\" class=\"wp-block-spacer\"><\/div>\n\n\n\n<div class=\"wp-block-tkuwpbs5-bs5-div ref_ol\" data-aos=\"normal\">\n<ol>\n<li>[1] N. Kohl and P. Stone. \u201cPolicy gradient reinforcement learning for fast quadrupedal locomotion\u201d. In: IEEE International Conference on Robotics and Automation, 2004. Proceedings. ICRA\u201904. 2004. 3. IEEE. 2004, 2619\u20132624. DOI: 10.1109\/robot.2004.1307456.<\/li>\n<li>[2] E. Theodorou, J. Buchli, and S. Schaal. \u201cReinforcement learning of motor skills in high dimensions: A path integral approach\u201d. In: 2010 IEEE International Conference on Robotics and Automation. IEEE. 2010, 2397\u20132403. DOI: 10.1109\/ROBOT.2010.5509336.<\/li>\n<li>[3] D.-H. Chun, M.-I. Roh, H.-W. Lee, J. Ha, and D. Yu, (2021) \u201cDeep reinforcement learning-based collision avoidance for an autonomous ship&#8221; Ocean Engineering 234: 109216. DOI: 10.1016\/j.oceaneng.2021.109216.<\/li>\n<li>[4] P. Rauber, A. Ummadisingu, F. Mutz, and J. Schmidhuber, (2021) \u201cReinforcement Learning in SparseReward Environments With Hindsight Policy Gradients&#8221; Neural Computation 33(6): 1498\u20131553. DOI: 10.1162\/neco_a_01387.<\/li>\n<li>[5] M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. Tobin, O. Pieter Abbeel, and W. Zaremba, (2017) \u201cHindsight experience replay&#8221; Advances in neural information processing systems 30:<\/li>\n<li>[6] R. S. Sutton, A. G. Barto, et al., (1999) \u201cReinforcement learning&#8221; Journal of Cognitive Neuroscience 11(1): 126\u2013134.<\/li>\n<li>[7] V. Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. Riedmiller, A. K. Fidjeland, G. Ostrovski, et al., (2015) \u201cHumanlevel control through deep reinforcement learning&#8221; nature 518(7540): 529\u2013533. DOI: 10.1038\/nature14236.<\/li>\n<li>[8] T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y. Tassa, D. Silver, and D. Wierstra, (2015) \u201cContinuous control with deep reinforcement learning&#8221; arXiv preprint arXiv:1509.02971:<\/li>\n<li>[9] R. Munos, T. Stepleton, A. Harutyunyan, and M. Bellemare, (2016) \u201cSafe and efficient off-policy reinforcement learning&#8221; Advances in neural information processing systems 29:<\/li>\n<li>[10] D. A. Van Dyk and X.-L. Meng, (2001) \u201cThe art of data augmentation&#8221; Journal of Computational and Graphical Statistics 10(1): 1\u201350. DOI: 10.1198\/10618600152418584.<\/li>\n<li>[11] C. Shorten and T. M. Khoshgoftaar, (2019) \u201cA survey on image data augmentation for deep learning&#8221; Journal of big data 6(1): 1\u201348. DOI: 10.1186\/s40537-019-0197-0.<\/li>\n<li>[12] M. Bayer, M.-A. Kaufhold, B. Buchhold, M. Keller, J. Dallmeyer, and C. Reuter, (2023) \u201cData augmentation in natural language processing: a novel text generation approach for long and short text classifiers&#8221; International journal of machine learning and cybernetics 14(1): 135\u2013150. DOI: 10.1007\/s13042-022-01553-3.<\/li>\n<li>[13] M. Laskin, A. Srinivas, and P. Abbeel. \u201cCurl: Contrastive unsupervised representations for reinforcement learning\u201d. In: International Conference on Machine Learning. PMLR. 2020, 5639\u20135650.<\/li>\n<li>[14] M. Laskin, K. Lee, A. Stooke, L. Pinto, P. Abbeel, and A. Srinivas, (2020) \u201cReinforcement learning with augmented data&#8221; Advances in neural information processing systems 33: 19884\u201319895.<\/li>\n<li>[15] I. Kostrikov, D. Yarats, and R. Fergus, (2020) \u201cImage augmentation is all you need: Regularizing deep reinforcement learning from pixels&#8221; arXiv preprint arXiv:2004.13649:<\/li>\n<li>[16] Y. Matsuo, Y. LeCun, M. Sahani, D. Precup, D. Silver, M. Sugiyama, E. Uchibe, and J. Morimoto, (2022) \u201cDeep learning, reinforcement learning, and world models&#8221; Neural Networks: DOI: 10.1016\/j.neunet.2022.03.037.<\/li>\n<li>[17] G. Brockman, V. Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba, (2016) \u201cOpenai gym&#8221; arXiv preprint arXiv:1606.01540:<\/li>\n<li>[18] M. Plappert, M. Andrychowicz, A. Ray, B. McGrew, B. Baker, G. Powell, J. Schneider, J. Tobin, M. Chociej, P. Welinder, et al., (2018) \u201cMulti-goal reinforcement learning: Challenging robotics environments and request for research&#8221; arXiv preprint arXiv:1802.09464:<\/li>\n<li>[19] E. Todorov, T. Erez, and Y. Tassa. \u201cMujoco: A physics engine for model-based control\u201d. In: 2012 IEEE\/RSJ international conference on intelligent robots and systems. IEEE. 2012, 5026\u20135033. DOI: 10.1109\/IROS.2012.6386109.<\/li>\n<li>[20] A. Raffin, A. Hill, M. Ernestus, A. Gleave, A. Kanervisto, and N. Dormann. Stable baselines3. 2019.<\/li>\n<\/ol>\n<\/div>\n\n\n\n<p><\/p>\n","protected":false},"author":3,"template":"wp-custom-template-detail-4-aricles","meta":{"_uag_custom_page_level_css":""},"categories":[10,6,516],"tags":[552],"acf":[],"uagb_featured_image_src":[],"uagb_author_info":{"display_name":"\u6797\u923a\u6db5","author_link":"\/jase\/?author=3"},"uagb_comment_info":0,"uagb_excerpt":"&nbsp;Copyright&nbsp;The Author(s). This is an open access article distributed under the terms of the&nbsp;Creative Commons Attribution&nbsp;License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited. Download Citation:\u00a0 BibTeX | http:\/\/dx.doi.org\/10.6180\/jase.202402_27(2).0011\u00a0\u00a0 Download PDF Motion control of robots is a high-dimensional, nonlinear control problem that&hellip;","_links":{"self":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope\/3031"}],"collection":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/tkuisotope"}],"about":[{"href":"\/jase\/index.php?rest_route=\/wp\/v2\/types\/tkuisotope"}],"author":[{"embeddable":true,"href":"\/jase\/index.php?rest_route=\/wp\/v2\/users\/3"}],"wp:attachment":[{"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=3031"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=3031"},{"taxonomy":"post_tag","embeddable":true,"href":"\/jase\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=3031"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}