计算机应用研究2026,Vol.43Issue(8):2301-2307,7.DOI:10.19734/j.issn.1001-3695.2025.12.0501
融合课程学习与自适应奖励塑形的移动机器人导航方法
Mobile robot navigation method based on integrated course learning and adaptive reward shaping
摘要
Abstract
Sparse rewards in end-to-end mobile robot navigation often lead to low learning efficiency and unstable convergence in conventional reinforcement learning methods.This study developed a deep reinforcement learning approach that integrated curriculum learning and adaptive collision entropy reward shaping.It introduced collision entropy derived from 2D LiDAR as a dynamic shaping term to quantify environmental uncertainty and guide obstacle-aware exploration.An adaptive scheduling mechanism adjusted the shaping coefficient β according to the slope of success rate variation,which balanced exploration and convergence during training.A staged curriculum progressively increased environmental difficulty to stabilize learning.Simula-tion results show that the proposed method achieves an 84.5%success rate and reduces the collision rate to 14.5%in dense environments.The method effectively alleviates sparse reward issues and improves navigation performance.关键词
深度强化学习/碰撞熵/课程学习/自适应奖励塑形/移动机器人导航Key words
deep reinforcement learning/collision entropy/curriculum learning/adaptive reward shaping/mobile robot navigation分类
信息技术与安全科学引用本文复制引用
林玉杰,吴伟林,付占悦,蔡君颖,石少雄..融合课程学习与自适应奖励塑形的移动机器人导航方法[J].计算机应用研究,2026,43(8):2301-2307,7.基金项目
广西民族大学科研基金资助项目(21KJQD20) (21KJQD20)
广西重点研发计划资助项目(桂科AB25069215) (桂科AB25069215)
广西民族大学相思湖青年学者创新团队项目(2023GXUNXSHQN06) (2023GXUNXSHQN06)
国家自然科学基金资助项目(62241302) (62241302)