现代情报2026,Vol.46Issue(6):76-88,13.DOI:10.3969/j.issn.1008-0821.2026.06.007
医学长文本因果证据链识别的分阶段微调方法研究
Two-Stage Fine-Tuning Method for Causal Evidence Chain Identification in Long-Form Medical Texts
摘要
Abstract
[Purpose/Significance]Under the background of the Healthy China strategy,the automatic identification of fine-grained knowledge elements and causal evidence chains in long-form medical texts is a key issue in medical infor-mation mining.[Method/Process]To address problems such as the complex structure of medical long texts and the diffi-culty in distinguishing the strength of causal relationships,this study proposed a recognition method based on two-stage fine-tuning.Medical literature from the PubMed Central(PMC)database was used as the data source.After screening,101 high-complexity medical long texts were selected to construct an annotated corpus containing 8005 labeled instances.A three-level annotation framework was established,including L1 entities,L2 attribute modifiers,and L3 causal evidence chains.Based on the Qwen3-8B architecture,the model was trained using parameter-efficient fine-tuning with LoRA and memory optimization techniques.A two-stage progressive fine-tuning strategy was adopted,in which entities and attri-butes were jointly extracted in the first stage,followed by causal strength classification and evidence span localization in the second stage.[Result/Conclusion]Experimental results show that the model achieves an F1 score of 83.89%in L1 drug entity recognition,an F1 score of 59.45%in L1 adverse drug reaction entity recognition,an F1 score of 51.12%in L2 attribute recognition,and a Macro-F1 score of 75.05%in L3 causal classification.The exact match rate for evidence span identification reaches 58.75%,which is superior to the baseline model.The proposed method enables coordinated identification of fine-grained knowledge elements and causal evidence chains,providing a technical approach for auto-mated medical information processing.关键词
医学长文本/细粒度知识抽取/因果证据链识别/大语言模型微调/循证医学Key words
long-form medical texts/fine-grained knowledge extraction/causal evidence chain recognition/large language model fine-tuning/evidence-based medicine分类
医药卫生引用本文复制引用
袁晓园..医学长文本因果证据链识别的分阶段微调方法研究[J].现代情报,2026,46(6):76-88,13.基金项目
江苏省社会科学面上项目"面向预防为主的医疗健康数据知识组织与知识服务平台构建研究"(项目编号22TQB003) (项目编号22TQB003)
江苏省教育厅哲学社会科学一般项目"基于DIKW转化体系的双一流学科知识抽取服务创新研究"(项目编号:2022JYB0004). (项目编号:2022JYB0004)