Journal of Scientific Reports

Deep Learning Multi-Target Underwater Acoustic Intelligence System

Author (s)

Saio Alusine Marrah, Abu Bakarr Koroma, Sayo Nakeleh Turay, Gibrilla Deen Kamara, Mabinty Marrah, Mohamed Thoronka, & Paul Conteh

Abstract

Increased complexity of underwater acoustic space poses great difficulty in the accurate and real-time identification of the target, especially in low signal-to-noise ratio (SNR) environments and in mixed-source sound waves. This research combines a Deep Learning Multi-Target Underwater Acoustic Intelligence System that incorporates a Convolutional Neural Network (CNN) with a Transformer-based attention encoder to obtain robust and explainable multi-classification of underwater acoustic signals. The model uses a trained resnet-50 backbone that computes spectral features (localized) in the form of Short-Time Fourier transform (STFT) spectrograms and then uses a multi-head self-attention mechanism to note long-term temporal features. A parallel attention fusion layer is used to facilitate both spatial and temporal representations, and the Focal Loss reconstruction makes weak or minority classes, including low-energy biological calls, more sensitive. The data set, containing real and simulated underwater records in five categories, namely ships, submarines, marine mammals, ambient noise and environmental interference was supplemented to display a variety of SNR conditions between 0 and 25 dB. Experimentally, the hybrid model has been found to be more precise and energy-efficient than CNN-only, LSTM, and Transformer-only baselines because it has 98.1 percent classification using a fixed modulus and above 90 percent classification in 0 dB SNR. Moreover, the model exhibits an average throughput rate of 370 frames per second with a real-time inference rate of 3.4 ms per frame, and hence the model is suitable in autonomous underwater vehicles (AUVs) and marine surveillance systems. Grad-CAM images verify that the attention module is concentrated on acoustical significance spectral areas, which proves the interpretability and cognitive openness of the model. On the whole, the hybrid framework represents a considerable breakthrough in the underwater acoustic intelligence sphere as it combines strengths, precision, and comprehensibility, and preconditions the emergence of intelligent sensing and real-time maritime observation platforms of the next generation.

Key words: Underwater acoustic intelligence, Deep learning, CNN–Transformer hybrid, Multi-target classification, Signal-to-noise robustness.

Download PDF

Title: Deep Learning Multi-Target Underwater Acoustic Intelligence System
Author: Saio Alusine Marrah, Abu Bakarr Koroma, Sayo Nakeleh Turay, Gibrilla Deen Kamara, Mabinty Marrah, Mohamed Thoronka, & Paul Conteh
Journal Name: Journal of Scientific Reports
Website: http://ijsab.com/jsr
ISSN: ISSN: 2708-7085 (Online), ISSN: 3079-9317 (Print)
Publisher IJSAB International
DOI: https://doi.org/10.58970/JSR.1152
Media: Online
Volume: 12
Issue: 1
Acceptance Date: 07/11/2025
Date of Publication: 11/11/2025
PDF URL: http://ijsab.com/wp-content/uploads/1152.pdf
Free download: Available
Page: 20-39
First Page: 20
Last Page: 39
Paper Type: Research Paper
Current Status: Published

Cite This Article:

Marrah, S. A., Koroma, A. B., Turay, S. N., Kamara, G. D., Marrah, M., Thoronka, M., & Conteh, P. (2026). Deep Learning Multi-Target Underwater Acoustic Intelligence System, Journal of Scientific Reports, 12(1), 20-39.  DOI: https://doi.org/10.58970/JSR.1152

 

About Author (s)

Saio Alusine Marrah, School of Software Engineering, University of Electronic Science and Technology of China (UESTC), Chengdu, China. ORCID: https://orcid.org/0009-0008-4727-2480

Abu Bakarr Koroma (Corresponding Author), Center for West African Studies of University of Electronic Science and Technology of China (CWAS of UESTC), School of Management and Economics (SME), (UESTC), Chengdu, China. &  Ernest Bai Koroma University of Science and Technology (EBKUST), Sierra Leone. ORCID: https://orcid.org/0009-0005-3613-104X

Sayo Nakeleh Turay, Institute of Public Administration and Management, University of Sierra Leone, Freetown, Sierra Leone. ORCID: https://orcid.org/0009-0005-9199-0154

Gibrilla Deen Kamara, School of Information and Communication Engineering, University of Electronic Science and Technology of China (UESTC), Chengdu, China. ORCID: https://orcid.org/0009-0005-1732-6407

Mabinty Marrah, Limkokwing University of Creative Technology, Freetown, Sierra Leone. ORCID: https://orcid.org/0009-0003-1413-4843

Mohamed Thoronka, College of Software Engineering, Nankai University, Tianjin, China. ORCID: https://orcid.org/0009-0000-7570-3964

Paul Conteh, College of Software Engineering, Nankai University, Tianjin, China.

 

Download PDF

DOI: https://doi.org/10.58970/JSR.1152

This Post Has Been Viewed 558 Times

Summarize this article with:
ChatGPT
ChatGPT
Perplexity
Perplexity
Mistral
Mistral
HuggingChat
HuggingChat
You.com
You.com
Grok
Grok