| 注册
首页|期刊导航|计算机科学与探索|基于多模态大语言模型的Web自动化智能体设计

基于多模态大语言模型的Web自动化智能体设计

段文瑞 王福喜 黄坚 高涛 沈博 刘鹄云天

计算机科学与探索2026,Vol.20Issue(6):1716-1732,17.
计算机科学与探索2026,Vol.20Issue(6):1716-1732,17.DOI:10.3778/j.issn.1673-9418.2511003

基于多模态大语言模型的Web自动化智能体设计

Design of Web Automation Agent Based on Multimodal Large Language Models

段文瑞 1王福喜 1黄坚 2高涛 3沈博 3刘鹄云天1

作者信息

  • 1. 中国电子科技集团公司 第十五研究所,北京 100083||北京航空航天大学 软件学院,北京 100191
  • 2. 北京航空航天大学 软件学院,北京 100191
  • 3. 中国电子科技集团公司 第十五研究所,北京 100083
  • 折叠

摘要

Abstract

Traditional Web automation methods rely on scripted operation strategies,which suffer from poor robustness and high maintenance costs,making them difficult to adapt to the dynamic and complexity of modern Web systems.Multi-modal large language models offer advantages such as visual understanding,intent recognition and complex reasoning,providing a workable technical path for building Web agent with autonomous perception and execution capabilities.Therefore,this paper proposes a Web automation method based on multimodal large language models and designs a Web agent to automate the processes of interface understanding,instruction reasoning,and autonomous interaction.The main contributions are threefold.By introducing multimodal semantic fusion and instruction-wise context awareness mechanisms,the Web agent's interface understanding and operation decision accuracy are enhanced.Based on this,this paper designs a prototype system of Web-oriented multimodal agent to achieve an end-to-end automated interaction process from natural language to operation execution.Finally,this paper constructs an experimental dataset to verify the feasibility and effectiveness of the method,and empirically demonstrates its adaptability and application potential under constrained model conditions.Experimental results show that the method outperforms baseline methods in terms of task success rate.In addition,multimodal semantic fusion and the instruction-wise context awareness mechanisms effectively improve the task success rate and the element localization accuracy in complex interfaces.The findings provide practical methods and experimental basis for the integrated application of Web automation technology and multimodal large language models.

关键词

多模态大语言模型/Web智能体/Web自动化/语义融合/上下文工程/网络数据抽取

Key words

multimodal large language model/Web agent/Web automation/semantic fusion/context engineering/Web data extraction

分类

信息技术与安全科学

引用本文复制引用

段文瑞,王福喜,黄坚,高涛,沈博,刘鹄云天..基于多模态大语言模型的Web自动化智能体设计[J].计算机科学与探索,2026,20(6):1716-1732,17.

基金项目

杭州市北京航空航天大学国际创新研究院科研启动基金(2024KQ026). This work was supported by the Research Start-up Fund of Hangzhou International Innovation Institute of Beihang University(2024KQ026). (2024KQ026)

计算机科学与探索

1673-9418

访问量0
|
下载量0
段落导航相关论文