计算机工程与应用2026,Vol.62Issue(11):17-40,24.DOI:10.3778/j.issn.1002-8331.2511-0018
视觉-语言-动作模型研究综述:迈向通用机器人
Survey of Vision-Language-Action Models:Towards General-Purpose Robots
摘要
Abstract
The vision-language-action(VLA)model,as an important research direction in embodied intelligence,aims to overcome the semantic fragmentation and generalization bottleneck of traditional robot systems in open environments by deeply integrating perception,understanding,and action.Based on the relevant progress in the VLA field,this paper first reviews the development history and current research status of VLA models.Then,from the perspective of system archi-tecture,existing models are classified into three types:monolithic,cascaded,and hierarchical,and the design ideas and representative progress of each type of architecture are analyzed in depth.At the same time,the relevant datasets and eval-uation benchmarks that support VLA research are systematically summarized.Finally,based on the representative achieve-ments in recent years,the core challenges faced by current VLA models are analyzed,and the future development direc-tions are prospected.关键词
具身智能/视觉-语言-动作模型/机器人/多模态融合Key words
embodied intelligence/vision-language-action model/robot/multimodal fusion分类
信息技术与安全科学引用本文复制引用
陈文祺,陈佳锋,支鹏翔,施露露,闻路红..视觉-语言-动作模型研究综述:迈向通用机器人[J].计算机工程与应用,2026,62(11):17-40,24.基金项目
宁波市自然科学基金(2024J217). (2024J217)