电子学报2026,Vol.54Issue(2):544-561,18.DOI:10.12263/DZXB.20250586
基于VLM凸优化的网络直播视频场景图生成
Scene Graph Generation of Livestreaming Video via VLM Convex Optimization
摘要
Abstract
Livestreaming video platforms have become an important medium for digital content dissemination,social interaction,and commercial activities.This is largely due to their large number of streamers,massive content supply,and ex⁃tremely high daily active user base.However,the real-time and unpredictable nature of livestreaming content poses serious challenges for online content supervision and regulation.Video scene graphs provide a structured representation for video understanding.They describe objects,attributes,and behavioral relationships within videos.By constructing a semantic net⁃work of"object-relation-action"in the spatiotemporal domain,video scene graphs enable structured modeling of video con⁃tent.In recent years,vision-language models(VLMs)have shown strong capabilities in cross-modal semantic understanding and complex scene reasoning.These advantages provide new technical support for livestreaming video scene graph genera⁃tion.Although VLMs can significantly improve semantic parsing accuracy in complex livestreaming scenarios,they still face an important challenge.Specifically,it is difficult to effectively capture the feature distribution patterns of livestream⁃ing videos.Convex optimization plays an important role in training VLMs.It helps guide the model to converge toward a global optimal solution.Based on this observation,this paper proposes a VLM-based convex optimization for scene graph generation(VCO-SGG).The method constructs a VLM-based approximately convex optimization framework that con⁃strains the geometric structure of the feature space for object semantics and their relationships,reducing feature distribution discrepancies and mitigating convergence oscillations during VLM training.A dynamic prototypical memory module is in⁃troduced,employing a parametric memory mechanism to strengthen the memory of key semantic elements'continuity and correlations across video frames.Furthermore,a feature association and relation filtering strategy is proposed to identify and filter redundant object indices online,which are generated in the scene graph due to dynamic changes,thereby enabling dy⁃namic generation and updating the scene graph.Experimental results demonstrate that our method achieves improvements of R@10 and mR@10 reaching 55.41%and 34.82%,on the self-built livestreaming video dataset BJUT-LGSD,respective⁃ly.In the publicly available datasets Mini Charades and Mini Action Genome datasets,R@10 and mR@10 are further im⁃proved to 48.19%/28.02%and 43.42%/26.02%,respectively,and the inference speed is 22.36 FPS.Overall,the results dem⁃onstrate greater competitiveness than other methods,indicating its capability to handle the task of generating scene graphs for livestreaming videos.关键词
网络直播视频/场景图生成/视觉语言模型/凸优化/动态原型记忆/特征联合与关系筛选Key words
livestreaming video/scene graph generation/vision-language models/convex optimization/dynamic pro⁃totype memory/feature association and relation filtering strategy分类
信息技术与安全科学引用本文复制引用
李文生,张菁,王艺晓,卓力..基于VLM凸优化的网络直播视频场景图生成[J].电子学报,2026,54(2):544-561,18.基金项目
国家自然科学基金(No.61971016,No.62471013) (No.61971016,No.62471013)
北京市自然科学基金(No.KZ201910005007) National Natural Science Foundation of China(No.61971016,No.62471013) (No.KZ201910005007)
Beijing Natural Science Foundation(No.KZ201910005007) (No.KZ201910005007)