| 注册
首页|期刊导航|北京航空航天大学学报|基于模仿学习的机动目标拦截强化学习制导律

基于模仿学习的机动目标拦截强化学习制导律

任乐亮 鲜勇 刘振宇 张大巧 李冰 李少朋

北京航空航天大学学报2026,Vol.52Issue(6):2156-2171,16.
北京航空航天大学学报2026,Vol.52Issue(6):2156-2171,16.DOI:10.13700/j.bh.1001-5965.2024.0284

基于模仿学习的机动目标拦截强化学习制导律

Reinforcement learning guidance law for maneuvering target interception based on imitation learning

任乐亮 1鲜勇 1刘振宇 2张大巧 3李冰 3李少朋4

作者信息

  • 1. 火箭军工程大学 导弹工程学院,西安 710025||跨域飞行交叉技术实验室,绵阳 621000
  • 2. 火箭军工程大学 导弹工程学院,西安 710025
  • 3. 火箭军工程大学 作战保障学院,西安 710025
  • 4. 火箭军工程大学 导弹工程学院,西安 710025||清华大学 自动化系,北京 100084
  • 折叠

摘要

Abstract

The advancement of maneuver penetration technology necessitates an improvement in the design of interception guidance laws to a higher level.An echelon intelligent guidance framework of"terminal guidance law model based on imitation learning(IL)→ terminal guidance law evolution model based on reinforcement learning"is proposed to increase the interception probability,decrease the energy consumption,and improve the robustness of the guidance law for intercepting maneuvering targets.Firstly,a three-dimensional uncertain confrontation model between the maneuvering target and interceptor is established based on the interception collision triangle.Secondly,the IL method is used to mine the proportional navigation guidance(PNG)law,which provides a good initial policy for the subsequent reinforcement learning guidance law.Finally,a Markov decision model is established,and a process reward of energy consumption and a"soft"terminal reward model,including a"transition section,"are proposed.The proximal policy optimization(PPO)algorithm is used to fully explore the high-performance interception strategy.The findings of the Monte Carlo simulation show that the new guidance law is very stable and resilient,outperforming the conventional guidance algorithm in terms of interception probability and energy usage.Additionally,the single decision time is only 0.32 ms,making it of certain engineering value.

关键词

模仿学习/深度强化学习/机动目标拦截/近端策略优化/碰撞三角

Key words

imitation learning/deep reinforcement learning/maneuvering target interception/proximal policy optimization/collision triangle

分类

航空航天

引用本文复制引用

任乐亮,鲜勇,刘振宇,张大巧,李冰,李少朋..基于模仿学习的机动目标拦截强化学习制导律[J].北京航空航天大学学报,2026,52(6):2156-2171,16.

基金项目

国家自然科学基金(62103432) (62103432)

陕西省高校科协青年人才托举计划(20210108) (20210108)

跨域飞行交叉技术实验室开放基金(2024-KYKF-4004) National Natural Science Foundation of China(62103432) (2024-KYKF-4004)

Young Talent Fund of University Association for Science and Technology in Shaanxi,China(20210108) (20210108)

Open Fund of Key Laboratory of Cross-Domain Flight Interdisciplinary Technology(2024-KYKF-4004) (2024-KYKF-4004)

北京航空航天大学学报

1001-5965

访问量0
|
下载量0
段落导航相关论文