南京大学学报(自然科学版)2026,Vol.62Issue(4):577-591,15.DOI:10.13232/j.cnki.jnju.2026.04.006
面向端到端语音翻译的分阶段训练与策略优化方法
A staged training and policy optimization method for end-to-end speech translation
摘要
Abstract
In recent years,large speech models have demonstrated strong capabilities in cross-modal representation and cross-lingual generation,offering a new paradigm for end-to-end speech translation(E2E ST).Existing E2E ST methods have gradually leveraged these advantages to alleviate alignment and optimization challenges inherent in traditional training.However,directly generating translations from speech remains difficult,as the model is required to simultaneously perform acoustic modeling,semantic understanding,and cross-lingual generation within a single objective.This challenge becomes more severe when parallel data are limited,often leading to unstable training and degraded translation quality.To address these issues,this paper proposes a staged training and policy optimization method for end-to-end speech translation.Under a unified autoregressive generation framework,the proposed method organizes keywords,ASR transcriptions,and target translations into a structured output sequence and progressively enhances the model's speech understanding and translation generation capabilities through three stages.First,keyword prediction and speech transcription are jointly modeled to construct a dual-granularity source representation comprising a keyword-level semantic skeleton and a complete source transcription,thereby establishing stable acoustic-semantic correspondences.Second,the full translation objective is introduced,and mixed auxiliary labels are constructed by combining human annotations with predictions from the previous stage,improving the model's adaptability to intermediate-representation noise during inference.Finally,reinforcement learning based on Group Relative Policy Optimization(GRPO)is adopted,with reward modeling applied exclusively to the final translation subsequence,to further refine the translation generation strategy.Experiments on three MuST-C language pairs using Qwen2-Audio,Qwen2.5-Omni,and Qwen3-Omni demonstrate that the proposed method achieves significant improvements in both BLEU and COMET scores,confirming that the integration of staged training,mixed auxiliary labels,and translation-level policy optimization effectively enhances both the stability and cross-lingual generation quality of end-to-end speech translation.关键词
端到端语音翻译/分阶段训练/策略优化/强化学习Key words
end-to-end speech translation/staged training/policy optimization/reinforcement learning分类
信息技术与安全科学引用本文复制引用
朱烨,李军辉,周国栋..面向端到端语音翻译的分阶段训练与策略优化方法[J].南京大学学报(自然科学版),2026,62(4):577-591,15.基金项目
国家自然科学基金(62376178) (62376178)