高电压技术2026,Vol.52Issue(7):2998-3009,12.DOI:10.13336/j.1003-6520.hve.20260524
MM-Power:面向电力巡检场景的视觉大模型多模态评测基准
MM-Power:A Multi-modal Benchmark for Evaluating Vision-language Models in Power Grid Inspection
摘要
Abstract
General-purpose vision-language models have shown strong performance on widely used multimodal bench-marks,yet a systematic understanding of their true capability boundaries in industrial fine-grained perception,visual reasoning,and hallucination resistance remains lacking.In power grid inspection,generic multimodal benchmarks often fail to cover the large variety of domain-specific targets,prominent small-target defects,complex imaging perspectives,and the high cost of false alarms.To address the mismatch between existing generic benchmarks and the demands of re-al-world power inspection,this paper presents MM-Power,a multimodal benchmark tailored to power grid inspection scenarios.MM-Power contains 2 708 multiple-choice questions organized into three tracks and 12 tasks,covering fi-ne-grained perception,visual-semantic consistency and hallucination assessment,and visual reasoning.The benchmark spans visible-light and infrared modalities,transmission/substation/distribution scenarios,five acquisition perspectives in-cluding pan-tilt cameras,unmanned aerial vehicles,robots,handheld cameras,and handheld infrared thermometers,and 130 categories of equipment states,defects,and anomalies.Based on a unified evaluation of 16 representative proprietary and open-weight vision-language models with both English and Chinese prompts,the best model achieves only 77.34%overall accuracy in English and 75.54%in Chinese.Significant weaknesses are observed in existence judgments,nega-tive-sample identification,and spatial relation reasoning,while prompt language also has a noticeable impact on the industrial-scenario performance of some models.These findings indicate that current vision-language models are still far from reliable deployment in practical power inspection,and MM-Power provides a unified benchmark for future evalua-tion,optimization,and deployment of multimodal models in the power domain.关键词
视觉大模型/电力巡检/多模态评测/幻觉评估/视觉推理/工业智能Key words
vision-language model/power grid inspection/multimodal benchmark/hallucination evaluation/visual rea-soning/industrial intelligence引用本文复制引用
闫云凤,齐冬莲,邓以恒,蔡芳仪,陈毅,龚世超,熊周智,曾子墨,蔡舒瑶,林嘉扬..MM-Power:面向电力巡检场景的视觉大模型多模态评测基准[J].高电压技术,2026,52(7):2998-3009,12.基金项目
国家自然科学基金(62476242) (62476242)
浙江省"尖兵""领雁"研发攻关计划项目(2025C01058).Project supported by National Natural Science Foundation of China(62476242),Pioneer R&D Program of Zhejiang Province(2025C01058). (2025C01058)