| 注册
首页|期刊导航|南京理工大学学报(自然科学版)|基于图嵌入的关键词抽取方法

基于图嵌入的关键词抽取方法

管维亚 王球 李中烜 朱颀林 徐建

南京理工大学学报(自然科学版)2026,Vol.50Issue(3):304-311,8.
南京理工大学学报(自然科学版)2026,Vol.50Issue(3):304-311,8.DOI:10.14177/j.cnki.32-1397n.2026.50.03.007

基于图嵌入的关键词抽取方法

Graph embedding-based keyword extraction approach

管维亚 1王球 1李中烜 1朱颀林 2徐建2

作者信息

  • 1. 国网江苏省电力有限公司经济技术研究院,江苏 南京 210008
  • 2. 南京理工大学 计算机科学与工程学院,江苏 南京 210094
  • 折叠

摘要

Abstract

Keyword extraction is one of hot research topics in the field of text mining,and its accuracy has significant impact on downstream tasks.To enhance the accuracy of keyword extraction for test,this paper proposes a graph embedding-based keyword extraction(GeKWE)approach.A candidate word graph is built for a preprocessed document,where nodes denote candidate keywords while edges denote the strength of co-occurrence relationship between candidate words.On this basis,drawing on the embedding method of the Word2Vec model,this paper uses a random walk strategy to obtain sampling nodes sequences from the graph and then trains a skip-gram model to learn representations of candidate words graph,obtaining the semantic vectors of candidate words.Further,candidate words are divided into several clusters using the embedding representation.For each cluster,candidate words with a high score are chosen as the final keywords.Experiments on two benchmark datasets,Inspec and SemEval2010,are conducted and comparisons with six state-of-the-art approaches are made.Experimental results show that the proposed approach achieves the highest F1 score on both datasets,which are 61.26%and 42.73%respectively,better than the selected comparison approaches,and is able to effectively alleviate information overload.

关键词

关键词抽取/图嵌入/预训练语言模型/表示学习

Key words

keyword extraction/graph embedding/pre-trained language model/representation learning

分类

信息技术与安全科学

引用本文复制引用

管维亚,王球,李中烜,朱颀林,徐建..基于图嵌入的关键词抽取方法[J].南京理工大学学报(自然科学版),2026,50(3):304-311,8.

基金项目

国家自然科学基金(61872186) (61872186)

南京理工大学学报(自然科学版)

1005-9830

访问量0
|
下载量0
段落导航相关论文