| 注册
首页|期刊导航|数字中医药(英文)|CMM-EmbedCluster:一种融合大语言模型与中药药性理论的中药聚类框架

CMM-EmbedCluster:一种融合大语言模型与中药药性理论的中药聚类框架

何佳怡 谢佳东 胡孔法 李海燕

数字中医药(英文)2026,Vol.9Issue(2):278-289,12.
数字中医药(英文)2026,Vol.9Issue(2):278-289,12.DOI:10.1016/j.dcmed.2026.05.010

CMM-EmbedCluster:一种融合大语言模型与中药药性理论的中药聚类框架

CMM-EmbedCluster:a clustering framework for Chinese materia medica based on large language model and Chinese materia medica property theory

何佳怡 1谢佳东 2胡孔法 3李海燕4

作者信息

  • 1. 中国中医科学院中医药信息研究所,北京 100700,中国||南京中医药大学人工智能与信息技术学院,江苏 南京 210023,中国
  • 2. 南京中医药大学人工智能与信息技术学院,江苏 南京 210023,中国
  • 3. 南京中医药大学人工智能与信息技术学院,江苏 南京 210023,中国||南京中医药大学江苏省中医药防治肿瘤协同创新中心,江苏 南京 210023,中国||南京中医药大学江苏省智慧中医药健康服务工程研究中心,江苏 南京 210023,中国
  • 4. 中国中医科学院中医药信息研究所,北京 100700,中国
  • 折叠

摘要

Abstract

Objective This study proposes a clustering framework for Chinese materia medica(CMM)based on a large language model(LLM),aiming to explore potential compatibility patterns among CMMs from the semantic perspective of CMM property theory. Methods First,a CMM property knowledge base was constructed based on Chinese Materia Medica,including 567 commonly used CMMs characterized by four properties,five flavors,and meridian tropism.Then,49 CMMs derived from 10 prescriptions for Zangdu(脏毒,pathogenic toxins)recorded in Waike Zhengzong(《外科正宗》,Orthodox Manual of Exter-nal Medicine)and Yangke Xinde Ji(《疡科心得集》,Collected Insights on Ulcer Medicine)were selected as the experimental dataset.Five semantic representation methods—One-Hot,Word2Vec,Bidirectional Encoder Representations from Transformers(BERT),Beijing Acade-my of Artificial Intelligence General Embedding(BGE),and Qwen—were applied to encode CMM property information into vector representations.Subsequently,t-distributed Stochas-tic Neighbor Embedding(t-SNE)was used for nonlinear dimensionality reduction on high-di-mensional semantic vectors,followed by k-means clustering(k=7).Clustering performance was evaluated using the Silhouette Score(SS),Davies-Bouldin Index(DBI),and Calinski-Harabasz Index(CHI). Results The Qwen-based clustering method,CMM-EmbedCluster,achieved the highest SS(0.607 4)and CHI(158.057 2),as well as the lowest DBI(0.499 5),indicating improved cluster separation and compactness compared with other methods.Visualization of CMM clustering results showed that the clusters were well separated in the low-dimensional space,with strong inter-cluster discrimination and high intra-cluster functional consistency.Further in-terpretability analysis of CMM clustering results revealed stable structural differences among clusters in terms of four properties,five flavors,and meridian tropism,forming functional partitions consistent with CMM property theory. Conclusion CMM-EmbedCluster utilizes an LLM to achieve semantic-level representation and clustering of CMMs within the framework of CMM property theory,providing support for exploring potential compatibility patterns among CMMs from the perspective of CMM prop-erty semantics.

关键词

大语言模型/中药药性理论/中药聚类/语义表征/k-means聚类/CMM-EmbedCluster

Key words

Large language model/Chinese materia medica property theory/Chinese materia medica clustering/Semantic representation/k-Means clustering/CMM-EmbedCluster

引用本文复制引用

何佳怡,谢佳东,胡孔法,李海燕..CMM-EmbedCluster:一种融合大语言模型与中药药性理论的中药聚类框架[J].数字中医药(英文),2026,9(2):278-289,12.

基金项目

Frontier Technologies Research and Development Pro-gram of Jiangsu(BF2025076),Scientific and Technologi-cal Innovation Project of China Academy of Chinese Medical Sciences(CI2021B002),and National Natural Science Foundation of China(82575255). (BF2025076)

数字中医药(英文)

2096-479X

访问量0
|
下载量0
段落导航相关论文