四川大学学报(自然科学版)2026,Vol.63Issue(4):862-876,15.DOI:10.19907/j.0490-6756.250369
知识增强与协同推理的视觉常识推理方法
Knowledge-augmented collaborative reasoning for visual commonsense reasoning
摘要
Abstract
Visual Commonsense Reasoning(VCR)aims to enable models to perform human-like cognitive reasoning based on images.Although existing methods primarily rely on large-scale cross-modal pre-trained models for implicit cross-modal modeling,they still fall short in terms of external knowledge incorporation,contextual modeling,and the consistency between answers and rationales.To address these limitations,the paper propose a knowledge-augmented collaborative reasoning framework for VCR.The proposed frame-work introduces a dual-layer knowledge injection mechanism comprising scene commonsense graphs and re-gional factual descriptions.By explicitly integrating scene-level commonsense with entity-level visual facts,this mechanism mitigates knowledge deficiency and semantic ambiguity.Meanwhile,a dual-task collabora-tive learning framework based on large language models is designed to jointly optimize question answering and rationale inference.A cognitive alignment loss is further employed to enhance the logical consistency be-tween the two tasks in the latent space.Experimental results on the VCR benchmark demonstrate that the proposed framework effectively improves model reasoning performance.Notably,it outperforms state-of-the-art baselines by 4.2 percentage points on the joint Q→AR task,thereby validating the effectiveness of the knowledge-augmented collaborative reasoning strategy.关键词
大语言模型/视觉常识推理/知识增强/协同学习Key words
Large language models/visual commonsense reasoning/knowledge enhancement/collaborative learning分类
信息技术与安全科学引用本文复制引用
李星悦,韩春燕,鲜雨成,琚生根,李勤..知识增强与协同推理的视觉常识推理方法[J].四川大学学报(自然科学版),2026,63(4):862-876,15.基金项目
国家自然科学基金重点项目(62137001) (62137001)