南京理工大学学报(自然科学版)2026,Vol.50Issue(2):121-127,7.DOI:10.14177/j.cnki.32-1397n.2026.50.02.001
基于组合Q网络的确定性策略梯度
Deterministic policy gradient based on combined Q-network
摘要
Abstract
In terms of the issues that value function estimation bias and unstable policy learning process often occur in the training phase of off-policy Actor-Critic deep reinforcement learning algorithm,a deterministic policy gradient based on combined Q-network is proposed.On the one hand,referring to the problem of value function estimation bias,a combined Q-network mechanism is built to perform the adaptive weighted combination of outputs of multiple Q-networks,thus obtaining a more accurate temporal difference(TD)target.In the process of estimating the value function,the weights can be adaptively adjusted based on the deviation between the Q-value estimation and the discount return.On the other hand,the stability of policy learning process is enhanced by constraining the L2 norm between the old and new policies.The effectiveness and superiority of the proposed algorithm are verified by the simulation results of continuous control tasks on MuJoCo platform.关键词
组合Q网络/值函数估计偏差/自适应权值/确定性策略梯度Key words
combined Q-network/value function estimation bias/adaptive weight/deterministic policy gradient分类
信息技术与安全科学引用本文复制引用
孔毅,吴阳,程玉虎,王雪松..基于组合Q网络的确定性策略梯度[J].南京理工大学学报(自然科学版),2026,50(2):121-127,7.基金项目
国家自然科学基金(62006232) (62006232)
江苏省自然科学基金(BK2020063) (BK2020063)