情报杂志2026,Vol.45Issue(9):154-163,10.DOI:10.3969/j.issn.1002-1965.2026.09.019
基于文体计量和机器学习的英俄语仇恨言论自动识别研究
Automatic Identification of English and Russian Hate Speech Based on Stylometry and Machine Learning
摘要
Abstract
[Purpose]This study aims to investigate the common linguistic characteristics and strategic differences of cross-lingual hate speech,providing empirical insights for the development of an interpretable multilingual threat intelligence monitoring system.[Method]Based on a corpus of approximately 2.2 million English and Russian online hate speech tokens,this study combines stylometric analysis to examine common features across five dimensions:words,sentences,readability,complexity,and sentiment.This study extracted latent stylistic structures through exploratory factor analysis,validated the distribution of features in the feature space using UMAP,and quanti-fied the contribution of each micro-indicator to classification decisions by combining machine learning with the SHAP attribution model,thereby establishing a core indicator system for identifying cross-lingual hate speech.[Result/Conclusion]A total of 52 features were i-dentified as common linguistic characteristics for distinguishing hate speech,including 15 lexical features,19 sentence features,8 readabil-ity features,7 complexity features,and 3 sentiment features.The experiments revealed significant differences in stylistic strategies between English and Russian:English tends toward structural complexity to increase the difficulty of censorship,while Russian tends toward sim-plicity and accessibility to accelerate online dissemination.Attribution analysis confirmed that the proportion of negative sentiment and proper nouns serve as the core decision-making basis for cross-lingual identification.This reveals the fundamental commonality of cross-lingual cyber threats in targeting specific audiences and inciting negative emotions.It provides interpretable core feature support for the au-tomatic identification of online public opinion risks and early warning of perceived threats,offering empirical data to enhance national cy-bersecurity governance capabilities.关键词
仇恨言论/文体计量/机器学习/特征归因分析/跨语言识别/网络舆情监测Key words
hate speech/stylistic measurement/machine learning/characteristic attribution analysis/cross-language recognition/network public opinion monitoring分类
社会科学引用本文复制引用
张瀚文,原伟..基于文体计量和机器学习的英俄语仇恨言论自动识别研究[J].情报杂志,2026,45(9):154-163,10.基金项目
国家社会科学基金项目"生成式人工智能赋能语言学多模态学科知识图谱建设与应用研究"(编号:24BYY079) (编号:24BYY079)
国家社会科学基金项目"基于联邦法案的军语释解与本体知识库构建研究"(编号:2024-SKJJ-A-009) (编号:2024-SKJJ-A-009)
国家语委十四五科研规划部级重点委托项目"面向军地协同语言保障任务的多语种语言资源建设规划研究"(编号:WT145-68) (编号:WT145-68)
国防科技大学外国语学院2026年研究生优秀学位论文培育项目"中国军事典籍的大模型俄译质量评估与知识增强研究"研究成果. ()