| 注册
首页|期刊导航|网络安全与数据治理|基于DPPO的双智能体协同渗透测试方法研究

基于DPPO的双智能体协同渗透测试方法研究

李科 霍朝宾 贺敏超 杨继

网络安全与数据治理2026,Vol.45Issue(6):1-8,8.
网络安全与数据治理2026,Vol.45Issue(6):1-8,8.DOI:10.19358/j.issn.2097-1788.2026.06.001

基于DPPO的双智能体协同渗透测试方法研究

Research on dual-agent collaborative penetration testing method based on DPPO

李科 1霍朝宾 1贺敏超 1杨继1

作者信息

  • 1. 华北计算机系统工程研究所,北京 100083
  • 折叠

摘要

Abstract

The automated network penetration testing method based on reinforcement learning has received widespread attention in recent years.However,the existing research is mostly limited to small and medium-sized network environments,making it difficult to address the challenges posed by the vast number of hosts and complex topology in real-world networks.This study proposes a dual-agent collaborative pene-tration testing framework based on DPPO(Dual-agent Proximal Policy Optimization),which enhances the efficiency of attack decision-making in large-scale network environments through a division of labor and collaboration mechanism.The framework comprises a target decision-mak-ing agent(TargetDecider)and an attack execution agent(AttackExecutor),responsible for target node selection and specific vulnerability ex-ploitation execution,respectively.The two agents collaborate through shared environmental states,guided by a collaborative reward mecha-nism,to achieve division of labor and cooperation.Experiments were conducted in a simulated network environment built based on CyberBat-tleSim,and compared with traditional agent methods such as DQN,PPO,and SAC.The results show that the DPPO framework requires fe-wer steps and exhibits less fluctuation to compromise key targets,demonstrating higher attack efficiency and strategic stability.At the same time,its sensitive host compromise rate remains at a high level,indicating that the framework can reliably complete core penetration tasks.In terms of cumulative reward,DPPO also shows a trend of continuous optimization and robust convergence.The results indicate that the frame-work can effectively adapt to penetration testing tasks in large-scale static simulated networks,while its applicability to dynamic defense sce-narios and real network environments still requires further validation.

关键词

自动化渗透测试/双智能体协作/深度强化学习

Key words

automated penetration testing/dual-agent collaboration/deep reinforcement learning

分类

信息技术与安全科学

引用本文复制引用

李科,霍朝宾,贺敏超,杨继..基于DPPO的双智能体协同渗透测试方法研究[J].网络安全与数据治理,2026,45(6):1-8,8.

网络安全与数据治理

2097-1788

访问量0
|
下载量0
段落导航相关论文