| 注册
首页|期刊导航|北京交通大学学报|网联环境下基于PPO的车队轨迹与信号配时协同优化方法

网联环境下基于PPO的车队轨迹与信号配时协同优化方法

邢凯然

北京交通大学学报2026,Vol.50Issue(3):152-163,12.
北京交通大学学报2026,Vol.50Issue(3):152-163,12.DOI:10.11860/j.issn.1673-0291.20250168

网联环境下基于PPO的车队轨迹与信号配时协同优化方法

Collaborative optimization method for platoon trajectory and signal timing based on proximal policy optimization in connected vehicle environments

邢凯然1

作者信息

  • 1. 北京交通大学 交通运输学院,北京 100044
  • 折叠

摘要

Abstract

Existing vehicle trajectory optimization methods based on deep reinforcement learning struggle to achieve the dynamic merging and splitting of platoons in connected,mixed-traffic environ-ments.Simultaneously,conventional signal timing optimization,which operates at the cycle or phase level,is hindered by response lag.The decoupled control between these two elements restricts further improvements in intersection efficiency.To address the necessity for the collaborative optimization of flexible vehicle trajectory control and real-time signal timing,this paper proposes a joint optimization method for platoon trajectories and signal timing based on Proximal Policy Optimization(PPO),termed FT-SIWO(Fleet Trajectory and Signal timing based on Waiting Offsetting).First,the vehicle trajectory optimization problem is formulated as a markov decision process.A state space integrating local vehicle dynamics and global traffic state variables,an action space encompassing car-following and multiple acceleration modes,and a multi-dimensional reward function considering factors such as driving comfort and safety are designed.This formulation enables the agent to comprehensively per-ceive the environment and learn efficient policies.Subsequently,the PPO-clip algorithm is employed for policy updates to enhance both training stability and sample efficiency.Finally,a real-time signal timing optimization mechanism based on Waiting Offsetting(WO)is introduced.By dynamically calcu-lating and balancing the"time loss"of vehicles in the green phase and the"time surplus"of vehicles in the red phase,this mechanism enables the dynamic adjustment of the green light duration at a second-level granularity,thereby achieving deep coordination with the vehicle trajectory optimization.Simula-tion results on a joint SUMO and Matlab platform demonstrate that,compared to a non-optimized baseline scheme,FT-SIWO reduces the average vehicle delay at intersections by 28%-32%and the average wait-ing time by 30%-35%.Additionally,it decreases the total emissions of various pollutants(CO2,CO,HC,NOx,and PMx)by 34%-38%and reduces fuel consumption by 19%.Furthermore,when com-pared to state-of-the-art distributed control or vehicle-infrastructure cooperation schemes,FT-SIWO achieves an additional 6%-10%improvement in control effectiveness.By enabling intelligent platoon formation and a second-level linkage with signal timing,this method effectively enhances green time utilization and the overall operational efficiency of traffic flow.It provides a viable collaborative optimi-zation solution for smart urban traffic management systems to handle mixed traffic flows,alleviate con-gestion,and reduce emissions.

关键词

智能交通/协同优化/车队轨迹/信号配时/近端策略优化

Key words

intelligent transportation/collaborative optimization/platoon trajectory/signal timing/proximal policy optimization

分类

交通工程

引用本文复制引用

邢凯然..网联环境下基于PPO的车队轨迹与信号配时协同优化方法[J].北京交通大学学报,2026,50(3):152-163,12.

基金项目

国家重点研发计划(2022YFB4301305) National Key R&D Plan(2022YFB4301305) (2022YFB4301305)

北京交通大学学报

1673-0291

访问量0
|
下载量0
段落导航相关论文