| 注册
首页|期刊导航|电子学报|面向自动驾驶的混合架构哈密顿-雅可比-贝尔曼近端策略优化方法研究

面向自动驾驶的混合架构哈密顿-雅可比-贝尔曼近端策略优化方法研究

王金强 宋利蓉 蒋远博 雍宾宾 李妍 周庆国

电子学报2026,Vol.54Issue(3):1024-1035,12.
电子学报2026,Vol.54Issue(3):1024-1035,12.DOI:10.12263/DZXB.20250977

面向自动驾驶的混合架构哈密顿-雅可比-贝尔曼近端策略优化方法研究

Research on Mixed Architecture Hamilton-Jacobi-Bellman Proximal Policy Optimization Method for Autonomous Driving

王金强 1宋利蓉 1蒋远博 1雍宾宾 1李妍 1周庆国1

作者信息

  • 1. 兰州大学信息科学与工程学院,甘肃 兰州 730000
  • 折叠

摘要

Abstract

Deep reinforcement learning(DRL)provides a powerful end-to-end learning framework for addressing complex sequential decision-making problems in autonomous driving,but the safety of vehicle control policies remains a core challenge.physics-informed reinforcement learning(PIRL)methods based on the hamilton-jacobi-bellman(HJB)equa-tion have demonstrated significant potential.However,such methods are severely limited in practice by the performance of the selected neural networks.Conventional multilayer perceptrons(MLPs)struggle to provide high-fidelity gradient signals for HJB physical constraints,thereby leading to training instability and model inefficiency issues.To address this challenge,we proposes a mixed architecture Hamilton-Jacobi-Bellman proximal policy optimization(MAHPO)algorithm tailored for autonomous driving tasks.This method innovatively constructs a heterogeneous Actor-Critic framework.Its policy network(Actor)uses an MLP to ensure efficient decision-making,while the value function network(Critic)is approximated by a kolmogorov-arnold network(KAN).Furthermore,the KAN-based value function representation network employs internal learnable smooth B-spline functions that can adaptively learn nonlinear transformations from trajectory data.This capability enables efficient modeling of complex value functions and their smooth gradient fields,thereby ensuring stable policy net-work updates.Experimental results in the MetaDrive simulation environment validate the efficacy of the MAHPO algo-rithm,which yields significant improvements over baselines across key performance metrics such as success rate,collision rate,and off-road rate.It has an average success rate improvement of 5.88%compared with the optimal benchmark soft ac-tor-critic(SAC),and the off-road rate has decreased by about 78.22%compared with the original HJBPPO algorithm.

关键词

深度强化学习(DRL)/自动驾驶/哈密顿-雅可比-贝尔曼(HJB)方程/混合架构/近端策略优化(PPO)

Key words

deep reinforcement learning(DRL)/autonomous driving/Hamilton-Jacobi-Bellman(HJB)equation/mixed architecture/proximal policy optimization(PPO)

分类

信息技术与安全科学

引用本文复制引用

王金强,宋利蓉,蒋远博,雍宾宾,李妍,周庆国..面向自动驾驶的混合架构哈密顿-雅可比-贝尔曼近端策略优化方法研究[J].电子学报,2026,54(3):1024-1035,12.

基金项目

兰州大学中央高校基本科研业务费专项资金(No.lzujbky-2024-eyt01) (No.lzujbky-2024-eyt01)

国家自然科学基金(No.61402210) (No.61402210)

甘肃省拔尖领军人才项目 Fundamental Research Funds for the Central Universities(No.lzujbky-2024-eyt01) (No.lzujbky-2024-eyt01)

National Natural Science Foundation of China(No.61402210) (No.61402210)

Gansu Province's Top Leading Talents ()

电子学报

0372-2112

访问量0
|
下载量0
段落导航相关论文