| 注册
首页|期刊导航|计算机工程与应用|视觉-语言-动作模型研究综述:迈向通用机器人

视觉-语言-动作模型研究综述:迈向通用机器人

陈文祺 陈佳锋 支鹏翔 施露露 闻路红

计算机工程与应用2026,Vol.62Issue(11):17-40,24.
计算机工程与应用2026,Vol.62Issue(11):17-40,24.DOI:10.3778/j.issn.1002-8331.2511-0018

视觉-语言-动作模型研究综述:迈向通用机器人

Survey of Vision-Language-Action Models:Towards General-Purpose Robots

陈文祺 1陈佳锋 2支鹏翔 2施露露 3闻路红3

作者信息

  • 1. 宁波大学 高等技术研究院,浙江 宁波 315211
  • 2. 宁波华仪宁创智能科技有限公司,浙江 宁波 315105
  • 3. 宁波大学 高等技术研究院,浙江 宁波 315211||宁波华仪宁创智能科技有限公司,浙江 宁波 315105
  • 折叠

摘要

Abstract

The vision-language-action(VLA)model,as an important research direction in embodied intelligence,aims to overcome the semantic fragmentation and generalization bottleneck of traditional robot systems in open environments by deeply integrating perception,understanding,and action.Based on the relevant progress in the VLA field,this paper first reviews the development history and current research status of VLA models.Then,from the perspective of system archi-tecture,existing models are classified into three types:monolithic,cascaded,and hierarchical,and the design ideas and representative progress of each type of architecture are analyzed in depth.At the same time,the relevant datasets and eval-uation benchmarks that support VLA research are systematically summarized.Finally,based on the representative achieve-ments in recent years,the core challenges faced by current VLA models are analyzed,and the future development direc-tions are prospected.

关键词

具身智能/视觉-语言-动作模型/机器人/多模态融合

Key words

embodied intelligence/vision-language-action model/robot/multimodal fusion

分类

信息技术与安全科学

引用本文复制引用

陈文祺,陈佳锋,支鹏翔,施露露,闻路红..视觉-语言-动作模型研究综述:迈向通用机器人[J].计算机工程与应用,2026,62(11):17-40,24.

基金项目

宁波市自然科学基金(2024J217). (2024J217)

计算机工程与应用

1002-8331

访问量0
|
下载量0
段落导航相关论文