Deep Learning Multi-Target Underwater Acoustic Intelligence System
Author (s)
Saio Alusine Marrah, Abu Bakarr Koroma, Sayo Nakeleh Turay, Gibrilla Deen Kamara, Mabinty Marrah, Mohamed Thoronka, & Paul Conteh
Abstract
Increased complexity of underwater acoustic space poses great difficulty in the accurate and real-time identification of the target, especially in low signal-to-noise ratio (SNR) environments and in mixed-source sound waves. This research combines a Deep Learning Multi-Target Underwater Acoustic Intelligence System that incorporates a Convolutional Neural Network (CNN) with a Transformer-based attention encoder to obtain robust and explainable multi-classification of underwater acoustic signals. The model uses a trained resnet-50 backbone that computes spectral features (localized) in the form of Short-Time Fourier transform (STFT) spectrograms and then uses a multi-head self-attention mechanism to note long-term temporal features. A parallel attention fusion layer is used to facilitate both spatial and temporal representations, and the Focal Loss reconstruction makes weak or minority classes, including low-energy biological calls, more sensitive. The data set, containing real and simulated underwater records in five categories, namely ships, submarines, marine mammals, ambient noise and environmental interference was supplemented to display a variety of SNR conditions between 0 and 25 dB. Experimentally, the hybrid model has been found to be more precise and energy-efficient than CNN-only, LSTM, and Transformer-only baselines because it has 98.1 percent classification using a fixed modulus and above 90 percent classification in 0 dB SNR. Moreover, the model exhibits an average throughput rate of 370 frames per second with a real-time inference rate of 3.4 ms per frame, and hence the model is suitable in autonomous underwater vehicles (AUVs) and marine surveillance systems. Grad-CAM images verify that the attention module is concentrated on acoustical significance spectral areas, which proves the interpretability and cognitive openness of the model. On the whole, the hybrid framework represents a considerable breakthrough in the underwater acoustic intelligence sphere as it combines strengths, precision, and comprehensibility, and preconditions the emergence of intelligent sensing and real-time maritime observation platforms of the next generation.
Key words: Underwater acoustic intelligence, Deep learning, CNN–Transformer hybrid, Multi-target classification, Signal-to-noise robustness.
| Title: | Deep Learning Multi-Target Underwater Acoustic Intelligence System |
|---|---|
| Author: | Saio Alusine Marrah, Abu Bakarr Koroma, Sayo Nakeleh Turay, Gibrilla Deen Kamara, Mabinty Marrah, Mohamed Thoronka, & Paul Conteh |
| Journal Name: | Journal of Scientific Reports |
| Website: | http://ijsab.com/jsr |
| ISSN: | ISSN: 2708-7085 (Online), ISSN: 3079-9317 (Print) |
| Publisher | IJSAB International |
| DOI: | https://doi.org/10.58970/JSR.1152 |
| Media: | Online |
| Volume: | 12 |
| Issue: | 1 |
| Acceptance Date: | 07/11/2025 |
| Date of Publication: | 11/11/2025 |
| PDF URL: | http://ijsab.com/wp-content/uploads/1152.pdf |
| Free download: | Available |
| Page: | 20-39 |
| First Page: | 20 |
| Last Page: | 39 |
| Paper Type: | Research Paper |
| Current Status: | Published |
Cite This Article:
Marrah, S. A., Koroma, A. B., Turay, S. N., Kamara, G. D., Marrah, M., Thoronka, M., & Conteh, P. (2026). Deep Learning Multi-Target Underwater Acoustic Intelligence System, Journal of Scientific Reports, 12(1), 20-39. DOI: https://doi.org/10.58970/JSR.1152
About Author (s)
Saio Alusine Marrah, School of Software Engineering, University of Electronic Science and Technology of China (UESTC), Chengdu, China. ORCID: https://orcid.org/0009-0008-4727-2480
Abu Bakarr Koroma (Corresponding Author), Center for West African Studies of University of Electronic Science and Technology of China (CWAS of UESTC), School of Management and Economics (SME), (UESTC), Chengdu, China. & Ernest Bai Koroma University of Science and Technology (EBKUST), Sierra Leone. ORCID: https://orcid.org/0009-0005-3613-104X
Sayo Nakeleh Turay, Institute of Public Administration and Management, University of Sierra Leone, Freetown, Sierra Leone. ORCID: https://orcid.org/0009-0005-9199-0154
Gibrilla Deen Kamara, School of Information and Communication Engineering, University of Electronic Science and Technology of China (UESTC), Chengdu, China. ORCID: https://orcid.org/0009-0005-1732-6407
Mabinty Marrah, Limkokwing University of Creative Technology, Freetown, Sierra Leone. ORCID: https://orcid.org/0009-0003-1413-4843
Mohamed Thoronka, College of Software Engineering, Nankai University, Tianjin, China. ORCID: https://orcid.org/0009-0000-7570-3964
Paul Conteh, College of Software Engineering, Nankai University, Tianjin, China.
DOI: https://doi.org/10.58970/JSR.1152
