| 注册
首页|期刊导航|信息通信技术与政策|面向大模型训练的通信行业高质量数据集构建方法与实践

面向大模型训练的通信行业高质量数据集构建方法与实践

肖文彬 李雨霏 黄倚霄 马闻达

信息通信技术与政策2026,Vol.52Issue(5):41-49,9.
信息通信技术与政策2026,Vol.52Issue(5):41-49,9.DOI:10.12267/j.issn.2096-5931.2026.05.006

面向大模型训练的通信行业高质量数据集构建方法与实践

Construction methods and practice of high-quality datasets for telecommunications large model training

肖文彬 1李雨霏 2黄倚霄 1马闻达2

作者信息

  • 1. 中国移动通信集团广东有限公司,广州 510150
  • 2. 中国信息通信研究院人工智能研究所,北京 100191
  • 折叠

摘要

Abstract

With the rapid evolution of generative artificial intelligence technology,data quality has become the core bottleneck restricting the performance of industry-scale large language models.Telecom operators possess ZB-scale cross-domain data,providing inherent resource advantages for training vertical large language models.However,raw communications data generally faces issues such as multi-source heterogeneity,high redundancy,and scarcity of long-tail samples,which limits its effectiveness when directly applied to model training.To address this,the system proposes a high-quality dataset construction method encompassing"collection-governance-annotation-evaluation,"featuring key technologies such as deep semantic compression,long-tail data synthesis based on Wasserstein Generative Adversarial Network(WGAN)and Long Short-Term Memory(LSTM)network,domain ontology construction,and human-machine collaborative annotation.Meanwhile,a general and specialized knowledge data coordination mechanism is designed to effectively mitigate catastrophic forgetting during industry fine-tuning.Practice has proved that this method is effective and can provide reference for the construction of high-quality datasets in the telecommunications industry.

关键词

高质量数据集/通信行业/数据合成/领域本体/数据增强

Key words

high-quality dataset/telecommunications industry/data synthesis/domain ontology/data augmentation

分类

管理科学

引用本文复制引用

肖文彬,李雨霏,黄倚霄,马闻达..面向大模型训练的通信行业高质量数据集构建方法与实践[J].信息通信技术与政策,2026,52(5):41-49,9.

信息通信技术与政策

2096-5931

访问量0
|
下载量0
段落导航相关论文