| 注册
首页|期刊导航|南京大学学报(自然科学版)|面向端到端语音翻译的分阶段训练与策略优化方法

面向端到端语音翻译的分阶段训练与策略优化方法

朱烨 李军辉 周国栋

南京大学学报(自然科学版)2026,Vol.62Issue(4):577-591,15.
南京大学学报(自然科学版)2026,Vol.62Issue(4):577-591,15.DOI:10.13232/j.cnki.jnju.2026.04.006

面向端到端语音翻译的分阶段训练与策略优化方法

A staged training and policy optimization method for end-to-end speech translation

朱烨 1李军辉 1周国栋1

作者信息

  • 1. 苏州大学计算机科学与技术学院,苏州,215006
  • 折叠

摘要

Abstract

In recent years,large speech models have demonstrated strong capabilities in cross-modal representation and cross-lingual generation,offering a new paradigm for end-to-end speech translation(E2E ST).Existing E2E ST methods have gradually leveraged these advantages to alleviate alignment and optimization challenges inherent in traditional training.However,directly generating translations from speech remains difficult,as the model is required to simultaneously perform acoustic modeling,semantic understanding,and cross-lingual generation within a single objective.This challenge becomes more severe when parallel data are limited,often leading to unstable training and degraded translation quality.To address these issues,this paper proposes a staged training and policy optimization method for end-to-end speech translation.Under a unified autoregressive generation framework,the proposed method organizes keywords,ASR transcriptions,and target translations into a structured output sequence and progressively enhances the model's speech understanding and translation generation capabilities through three stages.First,keyword prediction and speech transcription are jointly modeled to construct a dual-granularity source representation comprising a keyword-level semantic skeleton and a complete source transcription,thereby establishing stable acoustic-semantic correspondences.Second,the full translation objective is introduced,and mixed auxiliary labels are constructed by combining human annotations with predictions from the previous stage,improving the model's adaptability to intermediate-representation noise during inference.Finally,reinforcement learning based on Group Relative Policy Optimization(GRPO)is adopted,with reward modeling applied exclusively to the final translation subsequence,to further refine the translation generation strategy.Experiments on three MuST-C language pairs using Qwen2-Audio,Qwen2.5-Omni,and Qwen3-Omni demonstrate that the proposed method achieves significant improvements in both BLEU and COMET scores,confirming that the integration of staged training,mixed auxiliary labels,and translation-level policy optimization effectively enhances both the stability and cross-lingual generation quality of end-to-end speech translation.

关键词

端到端语音翻译/分阶段训练/策略优化/强化学习

Key words

end-to-end speech translation/staged training/policy optimization/reinforcement learning

分类

信息技术与安全科学

引用本文复制引用

朱烨,李军辉,周国栋..面向端到端语音翻译的分阶段训练与策略优化方法[J].南京大学学报(自然科学版),2026,62(4):577-591,15.

基金项目

国家自然科学基金(62376178) (62376178)

南京大学学报(自然科学版)

0469-5097

访问量0
|
下载量0
段落导航相关论文