控制理论与应用2026,Vol.43Issue(8):1649-1657,9.DOI:10.7641/CTA.2025.40465
空中机器人悬停控制的强化学习算法
Reinforcement learning for aerial robot hovering control
摘要
Abstract
As flying carriers for aerial operations,aerial robots have high requirements for level flight performance.The equipment they carry and the impact of reaction forces during operation also require aerial robots to have strong stable pose capabilities.This article focuses on hovering control for aerial robots,with an emphasis on improving reinforcement learning efficiency for this task.Firstly,the actor-critic algorithm based on behavioural value Q(QAC)and the deep deter-ministic policy gradient algorithm(DDPG)are used for robot hovering control learning.Subsequently,a policy gradient algorithm with added supervisors was proposed,using a PID controller as the dynamic monitor.Finally,simulation was conducted and the algorithm was deployed to the actual environment.Simulation and experimental results show that the designed reinforcement learning algorithm can achieve self-learning of hovering control for aerial robots,and compared with QAC and DDPG,it has higher learning efficiency,faster convergence speed,and smoother hovering effect.关键词
空中机器人/策略梯度算法/动态监督器/悬停控制/强化学习Key words
aerial robot/policy gradient algorithm/dynamic monitor/hovering control/reinforcement learning引用本文复制引用
卓浩泽,杨忠,吴吉莹,何加辉..空中机器人悬停控制的强化学习算法[J].控制理论与应用,2026,43(8):1649-1657,9.基金项目
广西电网有限责任公司科技项目(GXKJXM20230169),国家自然科学基金面上项目(61473144)资助.Supported by the Science and Technology Project of Guangxi Power Grid Co.,Ltd.(GXKJXM20230169)and the General Program of National Natural Science Foundation of China(61473144). (GXKJXM20230169)