Skip to main navigation Skip to search Skip to main content

MMA-Net: Multi-Modal Attention Network for 2-D Object Detection in Autonomous Driving

  • Abhilash Gaur
  • , Shubh Goel
  • , Kanishk Goel
  • , Seshan Srirangarajan
  • , Po Hsuan Tseng
  • , Kai Ten Feng

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

Abstract

Autonomous driving technology relies heavily on sensor data for environment perception. Heterogeneous sensors such as lidar, radar, and camera have their own strengths and limitations. Therefore, relying on any single sensor would restrict the effectiveness of autonomous driving technology. However, integrating data from such heterogeneous sensors poses challenges due to differences in their representations. This article outlines a deep learning network aimed at designing modality-agnostic multi-modal fusion architecture. We study sensor data from different modalities and learn fine-grained representations using modality-specific feature encoders independently. Then, a multi-modal attention-based network (MMA-Net) is proposed to fuse the data from heterogeneous modalities. The proposed MMA-Net fuses multi-modal sensor data by jointly exploiting the inter-modality and intra-modality relationships among camera, lidar, and radar sensors. The effectiveness of the proposed multi-modal fusion architecture is demonstrated using 2-D object detection metrics through extensive experiments on a dataset generated using the CARLA simulator.

Original languageEnglish
Title of host publication2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025 - Proceedings
EditorsBhaskar D Rao, Isabel Trancoso, Gaurav Sharma, Neelesh B. Mehta
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798350368741
DOIs
StatePublished - 2025
Event2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025 - Hyderabad, India
Duration: 6 Apr 202511 Apr 2025

Publication series

NameICASSP, IEEE International Conference on Acoustics, Speech and Signal Processing - Proceedings
ISSN (Print)1520-6149

Conference

Conference2025 IEEE International Conference on Acoustics, Speech, and Signal Processing, ICASSP 2025
Country/TerritoryIndia
CityHyderabad
Period6/04/2511/04/25

Keywords

  • autonomous driving
  • CARLA simulator
  • multi-modal learning
  • object detection
  • Sensor fusion

Fingerprint

Dive into the research topics of 'MMA-Net: Multi-Modal Attention Network for 2-D Object Detection in Autonomous Driving'. Together they form a unique fingerprint.

Cite this