| 注册
首页|期刊导航|计算机科学与探索|基于多任务预训练的情感音乐生成方法

基于多任务预训练的情感音乐生成方法

范舒柯 王海龙 柳林

计算机科学与探索2026,Vol.20Issue(7):2027-2036,10.
计算机科学与探索2026,Vol.20Issue(7):2027-2036,10.DOI:10.3778/j.issn.1673-9418.2507043

基于多任务预训练的情感音乐生成方法

Multi-task Pre-training Framework for Emotional Music Generation

范舒柯 1王海龙 2柳林1

作者信息

  • 1. 内蒙古师范大学 计算机科学技术学院,呼和浩特 010022
  • 2. 内蒙古师范大学 计算机科学技术学院,呼和浩特 010022||内蒙古师范大学 科学技术史研究院,呼和浩特 010022
  • 折叠

摘要

Abstract

AI-based affective music generation(AI-AMG)enables flexible music creation tailored to environmental contexts and listeners' physiological or emotional states,aiming to stimulate and regulate affective responses.It demonstrates strong potential in domains such as mental health therapy,immersive multimedia experiences,and intelligent music composition.Challenges remain,including the scarcity of affective music datasets,the difficulty of emotion annotation,the limited perceptual capability of existing models,and low accuracy in affective expression due to ineffective learning paradigms.A unified framework for affective music understanding and generation is proposed by integrating multi-task pre-training with fine-tuning.During the pre-training phase,proxy tasks,including bar-level attribute masking,mode prediction,and melody completion are designed to associate emotional characteristics with structural musical features using large-scale unlabeled data.Task-specific attention mechanisms enable joint learning of musical and affective representations.In the fine-tuning phase,an octuple representation incorporating valence-arousal dimensions is introduced,allowing emotion-controllable music generation on annotated datasets via a sequence-to-sequence attention mechanism.Experimental results show that the method improves four-class emotion classification accuracy by 4.2 percentage points and achieves gains of 17.48 and 3.40 percentage points in two-class classification tasks compared with baseline model.Subjective evaluation confirms enhanced perceived audio quality,and ablation experiments further validate the effectiveness of each pre-training sub-task for sentiment modeling.

关键词

音乐信息检索(MIR)/情感音乐生成/多任务预训练/音乐表示

Key words

music information retrieval(MIR)/emotional music generation/multi-task pre-training/music representation

分类

信息技术与安全科学

引用本文复制引用

范舒柯,王海龙,柳林..基于多任务预训练的情感音乐生成方法[J].计算机科学与探索,2026,20(7):2027-2036,10.

基金项目

国家自然科学基金(62566047) (62566047)

国家重点研发计划(2020YFC1523305) (2020YFC1523305)

内蒙古自治区自然科学基金(2024LHMS06015) (2024LHMS06015)

2022年度国家社科基金冷门绝学研究专项学术团队项目(22VJXT008). This work was supported by the National Natural Science Foundation of China(62566047),the National Key Research and Development Program of China(2020YFC1523305),the Natural Science Foundation of Inner Mongolia Autonomous Region(2024LHMS06015),and the 2022 Academic Team Project of the National Social Science Fund of China Special Program for Endangered and Obscure Disciplines(22VJXT008). (22VJXT008)

计算机科学与探索

1673-9418

访问量0
|
下载量0
段落导航相关论文