摘要
Abstract
To address the challenges faced by photorealistic digital human in cinematic applications,including low precision of cross-lingual lip-syncing and limited dynamic range of generated visuals,this paper proposes a complete end-to-end pipeline.In the speech synthesis and voice cloning module,the MiniMax-Speech model is integrated with retrieval-based voice conversion(RVC)to achieve high-fidelity voice cloning for low-resource languages.In the lip-sync module,a multilin-gual adaptive strategy extends the SyncTalk 2D model's compatibility with various speech recognition models,enhancing naturalness in cross-lingual scenarios.In the visual enhancement module,an inverse tone-mapping algorithm is incorpo-rated to convert standard dynamic range(SDR)video into high dynamic range(HDR)video compliant with the ITU-R BT.2100 standard.Experimental results show that on a single NVIDIA A10 GPU,the inference time is only 50%of the video duration,with objective and subjective quality surpassing baseline.The effectiveness of the system has been vali-dated in news-broadcasting scenarios at Xinhua News Agency,and it can serve as a technical reference for film produc-tion,virtual studio production,and related fields.关键词
数字分身/多模态算法/唇音同步/HDR影像/生成式人工智能Key words
Digital Human/Multimodal Algorithm/Lip Synchronization/HDR Cinematic/Generative Artificial Intelligence分类
信息技术与安全科学