Monocular 3D Object Detection Using Multi-Scale Context Fusion Attention
DOI:
https://doi.org/10.5755/j01.itc.55.2.43442Keywords:
Monocular 3D Object Detection, Multi-Scale Feature Fusion, Attention Mechanism, Context Modelling, Autonomous DrivingAbstract
Monocular 3D object detection aims to infer three-dimensional spatial information from a single RGB image, but existing approaches are often challenged by depth ambiguity and insufficient interaction between global and local features. To mitigate these issues, a monocular 3D object detection framework based on multiscale context fusion attention is proposed in this paper. Multi-scale features are first extracted to capture objects at different spatial resolutions. Subsequently, global self-attention is employed to model long-range contextual dependencies, while a local context fusion module aggregates neighborhood features based on similarity-aware interactions, enabling more discriminative local representations. Through the coordinated global–local attention mechanism, depth reasoning capability is enhanced and background interference is effectively suppressed. Experiments conducted on the KITTI benchmark demonstrate that the proposed method achieves 27.12 / 17.68 / 14.21 AP for 3D detection and 34.72 / 23.58 / 19.67 AP for BEV detection, outperforming representative monocular baselines such as MonoDETR. These results indicate that the proposed
approach improves detection accuracy and robustness, particularly for distant and partially occluded objects.
Downloads
Published
Issue
Section
License
Copyright terms are indicated in the Republic of Lithuania Law on Copyright and Related Rights, Articles 4-37.


