郑州大学学报(工学版)2026,Vol.47Issue(4):9-16,8.DOI:10.13705/j.issn.1671-6833.2026.04.002
跨模态时空注意力与上下文门控的情感分析
Sentiment Analysis with Cross-modal Spatio-Temporal Attention and Contextual Gating
摘要
Abstract
In multimodal sentiment analysis,it is difficult to capture the temporal dynamics of multimodal data by interaction inconsistencies due to modality heterogeneity,the complexity of linguistic scenarios,and the inability of static cross-modal attention,which limits deep modality correlation mining and sentiment classification perform-ance.To address these challenges,a multimodal sentiment analysis framework was proposed incorporating cross-modal spatio-temporal attention(CM-STA)to capture spatio-temporal dependencies among text,image,and audi-o,enhancing cross-modal interactions.Contextual gating(CG)was used to dynamically filter features strongly cor-related with emotional expressions,emphasizing key sentiment information.A Transformer cross-modal fusion inter-action(TCMFI)was used to leverage multi-head self-attention and bilinear pooling for efficient deep cross-modal fusion.Experiments on the TESS(audio)and MVSA-Multiple(text,image)datasets yielded an accuracy of 81.45%,an F1 score of 80.84%,and an AUROC of 96.40%,outperforming the best baseline model MISA by 0.95,0.24,and 7.91 percentage points,respectively.Computational complexity analysis revealed that the pro-posed model occupied 7.8 GB of GPU memory with a 98%GPU utilization rate,achieving efficient fusion with low spatial complexity and high GPU utilization,surpassing baseline models in performance.These results demonstrated the superior performance and robust effectiveness of the proposed model in complex multimodal sentiment analysis scenarios.关键词
多模态情感分析/时空注意力/上下文门控/Transformer跨模态融合/跨模态交互Key words
multimodal sentiment analysis/spatio-temporal attention/contextual gating/Transformer cross-modal fusion/cross-modal interaction分类
信息技术与安全科学引用本文复制引用
李丽红,李志勋,刘威伟,秦肖阳..跨模态时空注意力与上下文门控的情感分析[J].郑州大学学报(工学版),2026,47(4):9-16,8.基金项目
河北省数据科学与应用重点实验室项目(10120201) (10120201)