Beyond DQN: A Comparative Analysis of Advanced Deep Reinforcement Learning Algorithms for Cost-Aware Intrusion Detection

Authors

  • Hina Javed epartment of Computer Science, Preston University

Abstract

It is a comparative study of six Deep Reinforcement Learning (DRL) algorithms namely Deep Q-Network (DQN), Double Deep Q-Network (DDQN), Dueling Double Deep Q-Network (D3QN), REINFORCE, Advantage Actor-Critic (A2C), Proximal Policy Optimization (PPO), with statistical validation, multi-dataset evaluation and extensive performance analysis in the context of cost-aware intrusion detection in extreme class imbalance. All DRL agents employed a fixed CNN – LSTM feature extractor and prioritized experience replay (PER) in the controlled experimental setup. They tested the three benchmark datasets, CSE-CIC-IDS2018, CIC-IDS2017 and UNSW-NB15. Asymmetric reward function was used to train each DRL agent. There were 10 runs of the simulation with different random seeds, for which Wilcoxon signed-rank tests and McNemar's tests were performed to test the statistical significance. Key performance measures included recall, alerts per million flows (ARMF), Precision-Recall AUC (PR-AUC), training convergence, inference latency, and computational complexity. At 1,247 ARMF, PPO is statistically more successful (p=0.002, Wilcoxon) than baseline DQN+PER with a mean recall of 93.55% (sigma=0.82). The most balanced D3QN solution was found to be 92.80% recall (sigma=0.76) at 1,105 ARMF. For all DRL methods with PER, the number of alerts was reduced by 28-32x compared to supervised baselines. The generalization was confirmed by cross-dataset validation, where PPO achieved 90.20% recall on CIC-IDS2017 and 87.80% on UNSW-NB15. All value-based methods were shown to be statistically significant (p<0.05). The average of the inference latency was still under 2ms per sample. The results provide guidelines for algorithm selection in SOC deployments. The multi-class and zero-day detection scenarios have not been studied. Only 3 data sets and fixed feature backbone are used in the study. This study statistically compares and contrasts six of the most prominent DRL algorithms for cost-aware intrusion detection with confidence intervals, effect sizes and also reproducibility analysis.

Downloads

Download data is not yet available.

References

1.Alemayehu, M., Ghanem, M. C., Kheddar, H., Dunsin, D., Kerrache, C. A., & Rathee, G. (2026). Low-latency DDoS detection for IIoT and SCADA networks using proximal policy optimisation and deep reinforcement learning. Information, 17(5), 412.

2.Al-Haija, Q. A., & Tamimi, S. A. (2026). A state-of-the-art survey of adversarial reinforcement learning for IoT intrusion detection. Computers, Materials & Continua, 87(1), 1-30.

3.Cevallos, J. F., Rizzardi, A., Sicari, S., & Porisini, A. C. (2023). Deep reinforcement learning for intrusion detection in Internet of Things: Best practices, lessons learnt, and open challenges. Computer Networks, 236, 110016.

4.Chen, Z., Li, H., Wang, R., & Cui, D. (2024). Progressive prioritized experience replay for multi-agent reinforcement learning. In 2024 43rd Chinese Control Conference (CCC) (pp. 8292-8296). IEEE.

5.Fahrmann, D., Jorek, N., Damer, N., Kirchbuchner, F., & Kuijper, A. (2022). Double deep Q-learning with prioritized experience replay for anomaly detection in smart environments. IEEE Access, 10, 60836-60848.

6.Feng, Y. (2026). PER-AE-DRL: A malicious traffic detection model based on prioritized experience replay and adversarial mechanism. Journal of Information Security and Applications, 96, 104298.

7.Janati, M., & Messaoudi, F. (2025). Intrusion detection system-based network behavior analysis: A systemic literature review. International Journal of Advanced Computer Science and Applications, 16(3), 793-802.

8.Kara, A. O., & Boyaci, A. (2026). Transformer-DDQN-based explainable and active intrusion detection architecture for network traffic analysis. Applied Sciences, 16(12), 5912.

9.Kumar, J. (2025). Deep reinforcement learning-based intrusion detection system for next-generation wireless networks. In 2025 6th International Conference on Inventive Research in Computing Applications (ICIRCA) (pp. 283-291). IEEE.

10.Li, H. (2026). Distributed network intrusion detection model using blockchain and deep reinforcement learning. Engineering Research Express, 8, 055225.

11.Louati, F., Ktata, F. B., & Amous, I. (2026). Boosting multi-agent reinforcement learning with pruned prioritized experience replay and collaborative learning. Journal of Ambient Intelligence and Humanized Computing.

12.Mesadieu, F., Torre, D., & Chennameneni, A. (2024). Leveraging deep reinforcement learning technique for intrusion detection in SCADA infrastructure. IEEE Access, 12, 63381-63399.

13.Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., ... & Hassabis, D. (2015). Human-level control through deep reinforcement learning. Nature, 518(7540), 529-533. [Foundational reference]

14.Mpoporo, L. J., Owolawi, P. A., & Tu, C. (2026). Deep reinforcement learning algorithms for intrusion detection: A bibliometric analysis and systematic review. Applied Sciences, 16(2), 1048.

15.Odeh, A. (2025). Application of deep reinforcement learning for intrusion detection in Internet of Things: A systematic review. Internet of Things, 31, 101531.

16.Rushenda, Ramli, K., & Purnamasari, P. D. (2026a). A CNN-LSTM-DQN policy with prioritized experience replay for cost-aware intrusion detection on CSE-CIC-IDS2018. Journal of Embedded System Security and Intelligent Systems, 7(2), 158-173.

17.Rushenda, Ramli, K., & Purnamasari, P. D. (2026b). Stability-aware evaluation of a CNN-LSTM-DQN intrusion detection system for zero-day and drifted network traffic. IIUM Engineering Journal, 27(2).

18.Sangoleye, F., Johnson, J., & Tsiropoulou, E. E. (2024). Intrusion detection in industrial control systems based on deep reinforcement learning. IEEE Access, 12, 151444-151459.

19.Schaul, T., Quan, J., Antonoglou, I., & Silver, D. (2016). Prioritized experience replay. International Conference on Learning Representations (ICLR).

20.Schulman, J., Wolski, F., Dhariwal, P., Radford, A., & Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347.

21.Sharma, V., & Kumar, M. (2025). Rainbow DQN for intrusion detection: A unified deep reinforcement learning approach across benchmark datasets. International Journal of Advanced Management, 38(5S).

22.Shen, Y., Liu, L., & Zhang, W. (2024). Deep Q-network-based heuristic intrusion detection against edge-based SIoT zero-day attacks. Computers & Security, 142, 103875.

23.Tellache, A., Amara Korba, A., Mokhtari, A., & Ghamri-Doudane, Y. (2026). Collaborative privacy-preserving network intrusion detection: A federated multi-agent reinforcement learning approach. Computer Communications, 256, 108565.

24.Yang, W., Acuto, A., Zhou, Y., & Wojtczak, D. (2026). A survey for deep reinforcement learning based network intrusion detection. Applied AI Letters, 7(2).

25.Yu, S., Wang, X., Shen, Y., Wu, G., Yu, S., & Shen, S. (2024). Novel intrusion detection strategies with optimal hyper parameters for industrial Internet of Things based on stochastic games and double deep Q-networks. IEEE Internet of Things Journal, 11(18), 29876-29889.

26.Zhang, H., & Maple, C. (2023). Deep reinforcement learning-based intrusion detection in IoT system: A review. IET Conference Proceedings, 2023(14).

27.Zhang, X., & Xu, Y. (2026). Adaptive real-time intrusion prevention framework using double DQN and prioritized experience replay on live network traffic streams. In 2026 4th International Conference on Intelligent Systems and Computing (ICISC). IEEE.

Zhao, Y., Ma, D., & Liu, W. (2024). Efficient detection of malicious traffic using a decision tree-based proximal policy optimisation algorithm: A deep reinforcement learning malicious traffic detection model incorporating entropy. Entropy, 26(8), 648.

Downloads

Published

2026-09-22