| 注册
首页|期刊导航|电子学报|CIMOT3D:基于中文引导的单目视角下三维多目标跟踪研究

CIMOT3D:基于中文引导的单目视角下三维多目标跟踪研究

王荣 胡海祥 魏弘凯 梁浩翔 钱晓伟 李凯飞 郭柯宇 宋翔宇 孙士杰

电子学报2026,Vol.54Issue(1):102-114,13.
电子学报2026,Vol.54Issue(1):102-114,13.DOI:10.12263/DZXB.20250826

CIMOT3D:基于中文引导的单目视角下三维多目标跟踪研究

CIMOT3D:Chinese-Instruction-Based Monocular 3D Multi-Object Tracking

王荣 1胡海祥 1魏弘凯 1梁浩翔 2钱晓伟 1李凯飞 1郭柯宇 1宋翔宇 3孙士杰3

作者信息

  • 1. 长安大学信息工程学院,陕西 西安 710064
  • 2. 长安大学电子与控制工程学院,陕西 西安 710064
  • 3. 长安大学数据科学与人工智能研究院,陕西 西安 710064
  • 折叠

摘要

Abstract

Natural language-driven object tracking parses human-like language descriptions and fuses them with visu-al information to achieve accurate recognition and continuous tracking of specific targets in complex environments.Howev-er,existing methods focus on 2D tracking or 3D single-target tracking,and they have not been effectively extended to 3D multi-target tracking.They lack the capability to align text with multiple candidate targets in 3D visual space and to estab-lish associations.In addition,existing natural language-driven 3D object tracking tasks suffer from redundancy in language descriptions,which makes it hard to track multiple specific targets using flexible and concise instructions as humans do.To address these challenges,this paper introduces a new task,chinese-instruction-based monocular 3D multi-object tracking(CIMOT3D).The paper also constructs a new dataset,CIMOT3D-5k,which contains 5 562 video sequences with human-like Chinese descriptions.Furthermore,this paper designs a neural network model chinese-instruction-based monocular 3D multi-object tracking synchronization tracker(CIMOT3D-SyncTracker)for this task,which consists of a multimodal feature extractor,a vision-language encoder-decoder,and a detection-tracking module.Compared with baseline methods,the pro-posed approach achieves an improvement of 4.1%in tracking accuracy and 5.0%in identity consistency metric on the CIMOT3D-5k dataset,verifying its performance advantage.This paper advances research on vision-language fusion in 3D multi-object tracking and offers new ideas for further exploration in related fields.

关键词

场景理解/三维目标跟踪/多目标跟踪/视觉语言模型/多模态学习/机器视觉

Key words

scene understanding/3D object tracking/multi-object tracking/vision-language model/multimodal learn-ing/machine vision

分类

信息技术与安全科学

引用本文复制引用

王荣,胡海祥,魏弘凯,梁浩翔,钱晓伟,李凯飞,郭柯宇,宋翔宇,孙士杰..CIMOT3D:基于中文引导的单目视角下三维多目标跟踪研究[J].电子学报,2026,54(1):102-114,13.

基金项目

国家重点研发计划(No.2023YFB4301800) (No.2023YFB4301800)

国家自然科学基金(No.62576050) (No.62576050)

国家资助博士后研究人员计划(No.GZC20241447) (No.GZC20241447)

长安大学中央高校基本科研业务费专项资金(No.300102325101) National Key Research and Development Program of China(No.2023YFB4301800) (No.300102325101)

National Natural Science Foundation of China(No.62576050) (No.62576050)

National Postdoctoral Researcher Program(No.GZC20241447) (No.GZC20241447)

Fun-damental Research Funds for the Central Universities,CHD(No.300102325101) (No.300102325101)

电子学报

0372-2112

访问量0
|
下载量0
段落导航相关论文