Journal of Applied Science and Engineering

Published by Tamkang University Press

ESCI jase impact factor scopus logo open access rate of Scopus journal

Robust and Explainable Deep Learning Framework Against Adversarial Attacks for Network Threat Classification

Ruili Wang

Puyang Petrochemical Vocational and Technical College, 457000, Puyang China

Received: April 19, 2026
Accepted: May 18, 2026
Publication Date: June 15, 2026

上傳圖片

Workflow of the proposed REDLF 

 Copyright The Author(s). This is an open access article distributed under the terms of the Creative Commons Attribution License (CC BY 4.0), which permits unrestricted use, distribution, and reproduction in any medium, provided the original author and source are cited.

Download Citation:  BibTeX | http://dx.doi.org/10.6180/jase.202610_33.011  

Download PDF

With the rapid digitization of critical infrastructure and the proliferation of complex cyber threats, deep learning (DL) has become a core technology for network threat classification, enabling accurate identification of malicious activities such as DDoS attacks, malware infections, and phishing attempts. However, DL models are inherently vulnerable to adversarial attacks, subtle, human-imperceptible perturbations added to input network traffic features that can misleadingly alter model predictions, posing severe risks to network security. Additionally, the black-box nature of most advanced DL models (e.g., Transformers, deep neural networks) hinders their practical deployment in security-critical scenarios, as security analysts cannot interpret the
reasoning behind classification decisions. To address these two critical challenges with robustness against adversarial attacks and model explainability, this paper proposes a novel Robust and Explainable Deep Learning Framework (REDLF) for network threat classification. The framework integrates three core components: (1) an Adversarial Training Module (ATM) based on projected gradient descent (PGD) and physical environment modeling to enhance model robustness; (2) a Feature Enhancement Module (FEM) that combines attention mechanisms and variational autoencoders (VAEs) to extract discriminative and robust traffic features; (3) an Explainable Interpretation Module (EIM) that fuses SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) to generate both global and local explanations for classification decisions. Theoretical analysis proves the convergence and robustness of the proposed framework, and extensive experiments are conducted on three benchmark network threat datasets (CSE-CIC-IDS2018, NSL-KDD, and CTU-13) under six typical adversarial attacks (FGSM, PGD, C&W, JSMA, BIM, and DeepFool). Experimental
results demonstrate that REDLF outperforms state-of-the-art methods in both classification accuracy and adversarial robustness achieving an average accuracy of 98.72% on clean data and maintaining an accuracy of over 92.35%understrongadversarial attacks, which is 7.89% higher than SOTA models on average. Furthermore, the EIM module provides intuitive, human-interpretable explanations that clarify the key traffic features (e.g., flow duration, packet length, and protocol type) influencing classification decisions, addressing the black-box problem of DL models. This work contributes to bridging the gap between robustness and explainability in DL-based network threat classification, providing a reliable and interpretable solution for real-world network security defense.

Keywords: Deep Learning; Network Threat Classification; Adversarial Attacks; SHAP; Adversarial Training

  1. [1] S. Yin, H. Li, A. A. Laghari, T. R. Gadekallu, G. A. Sampedro, and A. Almadhor, (2024) “An anomaly detection model based on deep auto-encoder and capsule graph convolution via sparrow search algorithm in 6G Internet of Everything” IEEE Internet of Things Journal 11(18): 29402–29411. DOI: 10.1109/JIOT.2024.3353337.
  2. [2] N. Jhanjhi, M. Humayun, and S. N. Almuayqil, (2021) “Cyber security and privacy issues in industrial internet of things.” Computer Systems Science & Engineering 37(3): DOI: 10.32604/csse.2021.015206.
  3. [3] A. Mallik, (2019) “Man-in-the-middle-attack: Understanding in simple words” International Journal of Data and Network Science 2(2): 109–134. DOI: 10.5267/j.ijdns.2019.1.001.
  4. [4] I. A. Elshaer, A. M. Azazz, S. Fayyad, C. Kooli, A. M. Fouad, A. Hamdy, and E. A. Fathy, (2025) “Consumer boycotts and fast-food chains: economic consequences and reputational damage” Societies 15(5): 114. DOI: 10.3390/soc15050114.
  5. [5] S. Yin, H. Li, A. A. Laghari, L. Teng, T. R. Gadekallu, and A. Almadhor, (2024) “FLSN-MVO: edge computing and privacy protection based on federated learning Siamese network with multi-verse optimization algorithm for industry 5.0” IEEE Open Journal of the Communications Society 6: 3443–3458. DOI: 10.1109/OJCOMS.2024.3520562.
  6. [6] A. A. Laghari, H. Li, Y. Shoulin, S. Karim, A. A. Khan, and M. Ibrar, (2023) “Blockchain applications for Internet of Things (IoT): A review” Multiagent and Grid Systems 19(4): 363–379. DOI: 10.3233/MGS-230074.
  7. [7] S. Ness, (2024) “Adversarial attack detection in smart grids using deep learning architectures” IEEE Access 13: 16314–16323. DOI: 10.1109/ACCESS.2024.3523409.
  8. [8] K. Roshan, A. Zafar, and S. B. U. Haque, (2024) “Untargeted white-box adversarial attack with heuristic defence methods in real-time deep learning based network intrusion detection system” Computer Communications 218: 97–113. DOI: 10.1016/j.comcom.2023.09.030.
  9. [9] Y. K. Saheed and J. E. Chukwuere, (2026) “XAI-Enhanced adversarial resilient deep learning framework for transparent and secure edge deployment in consumer Internet of Things/Industrial Internet of Things environments” Computers and Electrical Engineering 133: 111080. DOI: 10.1016/j.compeleceng.2026.111080.
  10. [10] W. Sun, N. Huang, M. Yan, L. Huang, Z. Liu, X. Liu, and D. Lo, (2026) “Cost-Effective Adversarial Attacks Against Code LLM with Model Attention” IEEE Transactions on Software Engineering 52(4): 1371–1390. DOI: 10.1109/TSE.2026.3663143.
  11. [11] J. Yu, J. Du, and M. Liang. “A Fusion Model to Cognize the Structure of Heterogeneous Graph for Public Safety Scenarios”. In: 2026 IEEE International Conference on Big Data and Smart Computing (BigComp). IEEE. 2026, 88–95. DOI: 10.1109/BigComp68355.2026.00024.
  12. [12] S. Ghosh, R. K. Goyal, and K. Chowdhury, (2026) “Explainable AI-Driven Intrusion Detection System for DoS Attack Classification Using Deep Learning and Optimization Techniques” IEEE Access 14: 5618–5642. DOI: 10.1109/ACCESS.2026.3651187.
  13. [13] M. A. Yagiz and P. Goktas, (2025) “LENS-XAI: redefining lightweight and explainable network security through knowledge distillation and variational autoencoders for scalable intrusion detection in cybersecurity” arXiv preprint arXiv:2501.00790: DOI: 10.48550/arXiv.2501.00790.
  14. [14] N. Papernot, P. McDaniel, X. Wu, S. Jha, and A. Swami. “Distillation as a defense to adversarial perturbations against deep neural networks”. In: 2016 IEEE symposium on security and privacy (SP). IEEE. 2016, 582–597. DOI: 10.1109/SP.2016.41.
  15. [15] Z. Allen-Zhu and Y. Li. “Feature purification: How adversarial training performs robust deep learning”. In: 2021 IEEE 62nd annual symposium on foundations of computer science (FOCS). IEEE. 2022, 977–988. DOI: 10.1109/FOCS52979.2021.00098.
  16. [16] A. Bashaiwth, H. Binsalleeh, and B. AsSadhan, (2023) “An explanation of the LSTM model used for DDoS attacks classification” Applied Sciences 13(15): 8820. DOI: 10.3390/app13158820.
  17. [17] C. Duan, Y. Wang, W. Zhang, Z. Yu, Y. Pei, M. Zhang, and Q. Huang, (2026) “TVAE-GAN: A Generative Model for Providing Early Warnings to High-Risk Students in Basic Education and Its Explanation” Information 17(4): 356. DOI: 10.3390/info17040356.
  18. [18] J. L. Leevy and T. M. Khoshgoftaar, (2020) “A survey and analysis of intrusion detection models based on csecic-ids2018 big data” Journal of Big Data 7(1): 104. DOI: 10.1186/s40537-020-00382-x.
  19. [19] G. Meena and R. R. Choudhary. “A review paper on IDS classification using KDD 99 and NSL KDD dataset in WEKA”. In: 2017 International Conference on Computer, Communications and Electronics (Comptelix). IEEE. 2017, 553–558. DOI: 10.1109/COMPTELIX.2017.8004032.
  20. [20] D. Gaspar, P. Silva, and C. Silva, (2024) “Explainable AI for intrusion detection systems: LIME and SHAP applicability on multi-layer perceptron” IEEE Access 12: 30164–30175. DOI: 10.1109/ACCESS.2024.3368377.