Journal of Applied Science and Engineering

Published by Tamkang University Press

ESCI jase impact factor scopus logo open access rate of Scopus journal

Deep Reinforcement Learning For Adaptive Multi Axis CNC Tool Path Optimization Under Dynamic Constraints

Xiaorong Zhou, Lidong Huang, and Xiaoping Liu

School of Mechanical Engineering, Hunan Mechanical & Electrical Polytechnic, Changsha, Hunan, 410151, China

Received: June 05, 2026
Accepted: July 18, 2026
Publication Date: August 05, 2026

上傳圖片

Overall Flow of Multi Axis CNC Tool Path Optimization 

 Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.

Download Citation:  BibTeX | http://dx.doi.org/10.6180/jase.202611_34.015  

Download PDF

In modern production, multi-axis Computer Numerical Control (CNC) machines are utilized to cut complex pieces with extreme precision. However, during the cutting process, real-world problems such as tool wear, vibration, heat, and material changes make it difficult to adhere to predetermined tool paths. Traditional methods cannot react in real time, resulting in low efficiency and increased expenses. The goal of this research is to develop an intelligent and adaptive system for CNC tool path planning that uses deep reinforcement learning (DRL) to respond to dynamic machining environments and increase cutting performance. The dataset includes 3D CAD models, sensor data, tool positions, feed rates, and material parameters, which are utilized to train and validate the DRL agent. The proposed method utilizes fixed grid voxelization (FGV) to convert 3D CAD part representations into voxel grids, enabling the algorithm to analyze the machining region as discrete geometry blocks. Next, environment modeling is structured using a Markov Decision Process (MDP), considering tool conditions like location, spindle load, and vibration as states and tool movements as actions. Dijkstra’s algorithm (DA) is used to find the shortest baseline path across the voxel grid, which improves the DRL agent’s decision-making during toolpath generation. The learning agent uses Proximal Policy Optimizer Driven Double Deep Q-Network (PPO-DDQNet) for optimal tool pathways, ensuring precise value prediction and stabilizing the learning process. Geometry Planning under Varying Time Constraints based on angular loss, vibration loss 1450-2000 seconds.

Keywords: CNC environments, tool path planning, Proximal Policy Optimizer Driven Double Deep Q-Network (PPO-DDQNet), cutting performance, Computer Numerical Control (CNC)

  1. [1] O. Tuyboyov, A. Baydullayev, A. Jeltuxin, and Z. Muxiddinov, (2024) “Enhancing CNC machining tool path planning through reinforcement learning and optimization techniques” Applied Mechanics and Materials 923: 49–58. DOI: 10.4028/p-cU4ff5.
  2. [2] L. Zhang, H. Yu, C. Wang, Y. Hu, W. He, and D. Yu, (2025) “A digital solution for CPS-based machining path optimization for CNC systems” Journal of Intelligent Manufacturing 36(2): 1261–1290. DOI: 10.1007/s10845-023-02289-9.
  3. [3] Y. F. Feng, H. Y. Ma, L. Y. Shen, C. M. Yuan, and X. Jiang, (2023) “Real-time tool-path planning using deep learning for subtractive manufacturing” IEEE Transactions on Industrial Informatics 20(4): 5979–5988. DOI: 10.1109/TII.2023.3342474.
  4. [4] X. Zhao, C. Li, Y. Tang, X. Li, and X. Chen, (2024) “Reinforcement learning-based cutting parameter dynamic decision method considering tool wear for a turning machining process” International Journal of Precision Engineering and Manufacturing-Green Technology 11(4): 1053–1070. DOI: 10.1007/s40684-023-00582-9.
  5. [5] D. Kalandyk, B. Kwiatkowski, and D. Mazur, (2024) “CNC machine control using deep reinforcement learning” Bulletin of the Polish Academy of Sciences Technical Sciences: e148940. DOI: 10.24425/bpasts.2024.148940.
  6. [6] A. Pajaziti, O. Tafilaj, A. Gjelaj, and B. Berisha, (2025) “Optimization of toolpath planning and CNC machine performance in time-efficient machining” Machines 13(1): 65. DOI: 10.3390/machines13010065.
  7. [7] K. Wang, S. Zhang, Y. Wu, and F. Jiang, (2025) “Cutting path planning using reinforcement learning with adaptive sequence adjustment and attention mechanisms” The International Journal of Advanced Manufacturing Technology 136(11): 5599–5612. DOI: 10.1007/s00170-025-15200-y.
  8. [8] J. Dornheim, L. Morand, S. Zeitvogel, T. Iraki, N. Link, and D. Helm, (2022) “Deep reinforcement learning methods for structure-guided processing path optimization” Journal of Intelligent Manufacturing 33(1): 333–352. DOI: 10.1007/s10845-021-01805-z.
  9. [9] Y. Zhang, Y. Li, and K. Xu, (2022) “Reinforcement learning–based tool orientation optimization for five-axis machining” The International Journal of Advanced Manufacturing Technology 119(11): 7311–7326. DOI: 10.1007/s00170-022-08668-5.
  10. [10] F. Lu, G. Zhou, C. Zhang, Y. Liu, F. Chang, Q. Lu, and Z. Xiao, (2025) “Energy-efficient tool path generation and expansion optimization for five-axis flank milling with meta-reinforcement learning” Journal of Intelligent Manufacturing: 1–25. DOI: 10.1007 / s10845 – 024 – 02412-4.
  11. [11] Y. Jiang, J. Chen, H. Zhou, J. Yang, P. Hu, and J. Wang, (2022) “Contour error modeling and compensation of CNC machining based on deep learning and reinforcement learning” The International Journal of Advanced Manufacturing Technology 118(1): 551–570. DOI: 10.1007/s00170-021-07895-6.
  12. [12] P. Li, M. Chen, C. Ji, Z. Zhou, X. Lin, and D. Yu, (2024) “An agent-based method for feature recognition and path optimization of computer numerical control machining trajectories” Sensors 24(17): 5720. DOI: 10.3390/s24175720.
  13. [13] Y. Wan, W. Xu, and T. Y. Zuo, (2023) “Tool path optimization for complex cavity milling based on reinforcement learning approach” IEEE Access 11: 66793–66807. DOI: 10.1109/ACCESS.2023.3262169.
  14. [14] S. Liu, Z. Shi, J. Lin, and H. Yu, (2025) “A generalisable tool path planning strategy for free-form sheet metal stamping through deep reinforcement and supervised learning” Journal of Intelligent Manufacturing 36(4): 2601–2627. DOI: 10.1007/s10845-024-02371-w.
  15. [15] C. Guo, Z. Wang, K. Li, L. Ye, N. Yu, and F. Gong, (2025) “Domain knowledge integrated CAM system based on multi-objective path optimal planning and a deep convolutional neural network” Expert Systems with Applications: 127788. DOI: 10.1016/j.eswa.2025.127788.