Monocular 3D Object Detection Using Multi-Scale Context Fusion Attention

Authors

  • Sixu Yu School of Media and Communication, Shanghai Jiao Tong University, Shanghai, 200240, China

DOI:

https://doi.org/10.5755/j01.itc.55.2.43442

Keywords:

Monocular 3D Object Detection, Multi-Scale Feature Fusion, Attention Mechanism, Context Modelling, Autonomous Driving

Abstract

Monocular 3D object detection aims to infer three-dimensional spatial information from a single RGB image, but existing approaches are often challenged by depth ambiguity and insufficient interaction between global and local features. To mitigate these issues, a monocular 3D object detection framework based on multiscale context fusion attention is proposed in this paper. Multi-scale features are first extracted to capture objects at different spatial resolutions. Subsequently, global self-attention is employed to model long-range contextual dependencies, while a local context fusion module aggregates neighborhood features based on similarity-aware interactions, enabling more discriminative local representations. Through the coordinated global–local attention mechanism, depth reasoning capability is enhanced and background interference is effectively suppressed. Experiments conducted on the KITTI benchmark demonstrate that the proposed method achieves 27.12 / 17.68 / 14.21 AP for 3D detection and 34.72 / 23.58 / 19.67 AP for BEV detection, outperforming representative monocular baselines such as MonoDETR. These results indicate that the proposed 
approach improves detection accuracy and robustness, particularly for distant and partially occluded objects.

Downloads

Published

2026-07-23

Issue

Section

Articles