计算机工程与应用2026,Vol.62Issue(15):145-158,14.DOI:10.3778/j.issn.1002-8331.2505-0372
基于图像特征优化与自监督学习的视觉问答模型
Visual Question Answering Model Based on Image Feature Optimization and Self-Supervised Learning
摘要
Abstract
To tackle the challenges of insufficient image feature understanding and data bias in current visual question ans-wering(VQA)models,this paper proposes a novel VQA model featuring a dual visual feature encoder,text encoder,multi-modal feature fusion,and multimodal feature decoder,where the dual visual encoders extract multi-granularity features during question reasoning to address inadequate visual feature comprehension.In training,masked image reconstruction is employed to strengthen the correlation between local features and global structures,while generating highly irrelevant negative samples forces the model to learn true semantic associations between images and questions.The two-stage self-supervised training guides the model to deepen feature understanding from"pixel-level reconstruction"to"semantic-level alignment",optimizing image feature extraction and bias suppression.Experiments demonstrate that the model achieves 65.74%,67.50%,and 61.86%accuracies on VQA-CPv2,VQAv2,and OK-VQA datasets,surpassing state-of-the-art methods,with ablation studies and visual analyses validating the effectiveness of its modules.This work provides an effective approach for enhancing VQA models'generalization through rational architectural design and self-supervised training.关键词
视觉问答/图像特征优化/自监督训练/偏差抑制Key words
visual question answering/image feature optimization/self-supervised training/bias suppression分类
信息技术与安全科学引用本文复制引用
蔡谋熙,孙海春,张自勖..基于图像特征优化与自监督学习的视觉问答模型[J].计算机工程与应用,2026,62(15):145-158,14.基金项目
中国人民公安大学基本科研业务费(2024JKF02) (2024JKF02)
公安部技术研究计划基金项目(2024JSZ01). (2024JSZ01)