| 注册
首页|期刊导航|航空兵器|面向大机动目标的高效强化学习拦截制导律研究

面向大机动目标的高效强化学习拦截制导律研究

张烨 涂远刚 郭正玉 王靖宇

航空兵器2026,Vol.33Issue(2):44-53,10.
航空兵器2026,Vol.33Issue(2):44-53,10.DOI:10.12132/ISSN.1673-5048.2025.0145

面向大机动目标的高效强化学习拦截制导律研究

Research on Efficient Reinforcement Learning Interception Guidance Law for Large Maneuvering Targets

张烨 1涂远刚 2郭正玉 3王靖宇1

作者信息

  • 1. 西北工业大学 航天学院,西安 710072||空基信息感知与融合全国重点实验室,河南 洛阳 471009
  • 2. 西北工业大学 航天学院,西安 710072
  • 3. 中国空空导弹研究院,河南 洛阳 471009||空基信息感知与融合全国重点实验室,河南 洛阳 471009
  • 折叠

摘要

Abstract

In response to the challenges posed by intercepting highly maneuverable targets,including difficulties in obtaining target acceleration information,insufficient adaptability of traditional guidance laws to high-dynamic environments,low exploration efficiency of reinforcement learning algo-rithms in wide spatial-temporal contexts,and poor training stability,this paper proposes a terminal guidance law based on a double random distillation network and real proximal policy optimization algo-rithm.Firstly,a Markov decision process is designed for the three-dimensional terminal guidance sce-nario.Subsequently,deep reinforcement learning algorithms are employed to train high-speed vehicles.To enhance the exploration efficiency of these vehicles within wide spatial-temporal con-texts,a double random distillation network is introduced that incentivizes the exploration of unknown states through internal rewards.To accelerate training speed and ensure training stability,the trust region rollback objective function of real proximal policy optimization algorithm is used to replace the clipping objective function for effectively improving the convergence rate while enhancing training sta-bility.Simulation results demonstrate that the proposed terminal guidance law exhibits generalization and robustness when faced with highly maneuverable targets,and can achieve efficient interception without requiring target acceleration information.

关键词

末制导律/强化学习/训练效率/近端策略优化/随机网络蒸馏/飞行器

Key words

terminal guidance law/reinforcement learning/training efficiency/proximal policy optimization/random distillation network/vehicle

分类

军事科技

引用本文复制引用

张烨,涂远刚,郭正玉,王靖宇..面向大机动目标的高效强化学习拦截制导律研究[J].航空兵器,2026,33(2):44-53,10.

基金项目

国家自然科学基金项目(02564897) (02564897)

航空兵器

1673-5048

访问量1
|
下载量0
段落导航相关论文