南京大学学报(自然科学版)2026,Vol.62Issue(4):551-561,11.DOI:10.13232/j.cnki.jnju.2026.04.004
残缺条件下基于生成式轨迹建模的鲁棒模仿学习方法研究
Robust imitation learning via generative trajectory modeling from incomplete demonstrations
摘要
Abstract
Efficiently leveraging imperfect expert demonstration data is one of the key challenges in the field of imitation learning.Imperfect demonstrations,which entail the loss of fine-grained local details,can lead to error accumulation.To address scenarios with incomplete demonstrations,we propose a robust imitation learning method based on generative trajectory modeling.Our approach utilizes a Decision Transformer to generate continuous state-action sequences conditioned on incomplete demonstrations,thereby completing the trajectories.The quality of these generated trajectories is controlled by setting a target return.The completed trajectories are then fed into a diffusion model for further trajectory generation.Additionally,we introduce a novel scoring mechanism to evaluate whether the trajectories generated by the diffusion model conform to the environmental dynamics constraints present in the expert demonstrations.This model can generate high-reward trajectories that adhere to the dynamics constraints while effectively completing imperfect trajectories,thereby significantly reducing error accumulation and improving planning accuracy.Experiments show that our method outperforms existing offline reinforcement learning approaches.It successfully addresses the planning-to-reality mismatch problem in high-dimensional continuous spaces and demonstrates superior robustness and generalization capabilities.关键词
离线强化学习/模仿学习/扩散模型/轨迹补全Key words
offline reinforcement learning/imitation learning/diffusion models/trajectory completion分类
信息技术与安全科学引用本文复制引用
戴领,陈文韬,谭晓阳..残缺条件下基于生成式轨迹建模的鲁棒模仿学习方法研究[J].南京大学学报(自然科学版),2026,62(4):551-561,11.基金项目
国家自然科学基金(62476128) (62476128)