电工技术学报2026,Vol.41Issue(13):4386-4402,17.DOI:10.19595/j.cnki.1000-6753.tces.250696
基于时间与图嵌入强化学习的含5G基站储能配电网安全调度策略
A Safe Scheduling Strategy for Distribution Network Considering 5G Base Station Energy Storage Based on Time and Graph Embedding Reinforcement Learning
摘要
Abstract
With the large-scale deployment of 5G base stations,a number of base station energy storage systems have been integrated into the power grid.These storage devices possess a certain degree of dispatch flexibility and have the potential to facilitate the consumption of renewable energy.How to make full use of these energy storage resources in distribution network scheduling has become an important issue.Existing scheduling algorithms,however,suffer from long calculation times,resulting in poor real-time performance.In addition,traditional algorithms rely on precise mathematical models and are easily affected by environmental uncertainties,which leads to inaccurate solutions that deviate from the optimal solution.To address these challenges,this paper proposes a real-time scheduling algorithm based on reinforcement learning(RL),tailored to the characteristics of 5G base station energy storage,leveraging the fast decision capability of RL to handle the scheduling challenges. Firstly,based on the economic operation requirements of the distribution network,multiple cost components are considered,including the cost of purchasing electricity from the main grid,the cost of abandoning renewable energy,and the operating cost of the 5G base station energy storages.Taking into account the safety capacity constraints of 5G base station energy storages and physical constraints of distribution network,a mathematical optimization model for real-time scheduling is established.This model is then transformed into a Markov decision process(MDP),providing the foundation for an RL-based solution framework. Secondly,considering the characteristics of state information during the scheduling process,different types of information are processed separately.For temporal information,a novel Time2Vec(T2V)architecture is introduced to extract time-related features.For the graph-structured information,a graph convolutional network(GCN)is employed for feature extraction.By integrating these modules,the traditional deep deterministic policy gradient(DDPG)algorithm is improved,resulting in the proposed T2V-GCN-DDPG algorithm,which enhances the ability of agent to extract hidden features from the state,thereby improving decision-making performance.In addition,to ensure the safety of actions from agent,a safety constraint layer based on quadratic programming(QP)is designed to fine-tune the charging and discharging power of the base station energy storage.The QP model aims to minimize the magnitude of action adjustment while keeping the actions within the safe range. Finally,to validate the effectiveness of the proposed algorithm,experiments were conducted in modified IEEE 33-bus and 141-bus distribution networks incorporating 5G base station energy storage systems.By comparing the results with those obtained from intraday rolling optimization,proximal policy optimization(PPO),and behavior cloning,it was demonstrated that the proposed algorithm achieves superior decision-making performance.It effectively reduces the deviation in actions caused by environmental uncertainties and generates actions closer to the optimal solution.Moreover,comparisons of the algorithm before and after the proposed improvements further confirm the effectiveness of the enhancement method.Simulation results also indicate that the proposed algorithm,while ensuring the safety of decision actions,can reduce power grid operating costs and improve overall economic performance.Therefore,the proposed algorithm provides a reliable and efficient solution for the safe and economical operation of the power grid.关键词
强化学习/5G基站调度/图卷积网络/时间编码Key words
Reinforcement learning/5G base station energy storage scheduling/graph convolutional neural network/time embedding分类
信息技术与安全科学引用本文复制引用
裴青琦,陈逸诗,何智勇,刘清华,张森林..基于时间与图嵌入强化学习的含5G基站储能配电网安全调度策略[J].电工技术学报,2026,41(13):4386-4402,17.基金项目
中国长江三峡集团有限公司浙江分公司科研项目资助(合同编号ZJSL324009,项目编号 NBWL20240097). (合同编号ZJSL324009,项目编号 NBWL20240097)