| 注册
首页|期刊导航|电子学报|MSGPose:基于多语义图卷积与图引导状态空间模型的单目单人三维人体姿态估计

MSGPose:基于多语义图卷积与图引导状态空间模型的单目单人三维人体姿态估计

李俊 李昱 陈黎

电子学报2026,Vol.54Issue(3):1118-1131,14.
电子学报2026,Vol.54Issue(3):1118-1131,14.DOI:10.12263/DZXB.20250860

MSGPose:基于多语义图卷积与图引导状态空间模型的单目单人三维人体姿态估计

MSGPose:Monocular Single-Person 3D Human Pose Estimation Via Multi-Semantic Graph Convolution and Graph-Guided State Space Models

李俊 1李昱 2陈黎1

作者信息

  • 1. 武汉科技大学计算机科学与技术学院,湖北 武汉 430065||智能信息处理与实时工业系统湖北省重点实验室,湖北 武汉 430065
  • 2. 武汉科技大学计算机科学与技术学院,湖北 武汉 430065
  • 折叠

摘要

Abstract

Monocular single-person 3D human pose estimation holds immense value in action recognition and hu-man-computer interaction.However,the lack of depth cues in monocular setups introduces inherent depth ambiguity,while self-occlusion and imaging noise further complicate the task.Robustly recovering 3D poses from erroneous 2D observations remains a major challenge.Existing graph convolutional network(GCN)methods often abstract the human skeleton into a predefined physical graph and mainly rely on this fixed topology.Consequently,they fail to express non-physical semantic relations,such as bilateral symmetry.Conversely,self-attention-based methods excel at long-range modeling.Yet,they suf-fer from quadratic complexity in long sequences.This complexity leads to increased computational cost and parameter over-head.To address these limitations,we propose MSGPose.This is a dual-stream parallel framework for parameter-efficient monocular 3D human pose estimation.It integrates a multi-semantic dynamic separable graph convolution(MSDG)module and a semantic graph-guided mamba(SGM)module.The framework jointly models and extracts spatial and temporal fea-tures from 2D pose sequences.The MSDG module tackles spatial and temporal relations.It constructs dynamic multi-level semantic graphs using three priors:self-connections,physical connections,and anatomical symmetry.MSDG assigns inde-pendent learnable weights to each semantic branch,which helps alleviate semantic feature coupling.Additionally,a weight-ed modification matrix dynamically mitigates the inductive bias of fixed topologies.For temporal dynamics,MSDG em-ploys a sparse dynamic temporal graph convolution.It builds this graph using a K-Nearest Neighbors(K-NN)strategy based on feature similarity.This enables the modeling of cross-joint and inter-frame dependencies during complex move-ments.The SGM module addresses the spatial topology limitations of standard Mamba architectures.Flattening spatial to-kens into 1D sequences may disrupt the natural topology of the human skeleton.To fix this,SGM introduces a multi-semantic graph convolution guidance mechanism.This mechanism operates before the causal 1D convolution and bidirectional state space scanning.This step explicitly injects anatomical structure priors into the sequence representation.It provides the subse-quent state space model with a geometry-aware feature space.This enables efficient modeling of long-range spatio-temporal dependencies with linear complexity.During the feature fusion stage,MSGPose employs an adaptive mechanism.Learnable weights complementarily integrate the outputs from the two streams.The framework utilizes joint training optimized by 3D position and velocity losses.The velocity loss limits differences between adjacent frames.This strategy improves the tempo-ral consistency and stability of the predicted poses.Extensive experiments demonstrate the effectiveness of MSGPose.On the Human3.6M dataset,it achieves a mean per joint position error(MPJPE)of 38.9 mm,representing a 0.3 mm improve-ment over MotionBERT while using only 13.3M parameters(approximately 31%of MotionBERT).On the challenging MPI-INF-3DHP dataset,MSGPose demonstrates strong generalization ability.It achieves an MPJPE of 14.5 mm,representing a 1.7 mm improvement over MotionAGFormer.Using noise-free ground-truth 2D annotations,the MPJPE on Human3.6M drops significantly to 12.7 mm.These results demonstrate the effectiveness of combining multi-semantic priors with a dual-stream architecture.This combination improves the performance of 2D-to-3D pose regression.

关键词

三维人体姿态估计/多语义图卷积/状态空间模型/时空建模/语义先验/深度学习

Key words

3D human pose estimation/multi-semantic graph convolution/state space model/spatio-temporal model-ing/semantic prior/deep learning

分类

信息技术与安全科学

引用本文复制引用

李俊,李昱,陈黎..MSGPose:基于多语义图卷积与图引导状态空间模型的单目单人三维人体姿态估计[J].电子学报,2026,54(3):1118-1131,14.

基金项目

国家自然科学基金(No.62271359) National Natural Science Foundation of China(No.62271359) (No.62271359)

电子学报

0372-2112

访问量0
|
下载量0
段落导航相关论文