REAL-TIME SOUND EVENT LOCALIZATION AND DETECTION USING IOT RASPBERRY PI DEVICES BASED ON SINGLE-STAGE CRNN

Authors

  • Syed Owais Shah
  • Muneeba Darwaish
  • Mohsin Shah*
  • Muhammad Asad Khan
  • Muhammad Shujaat

Abstract

Sound event localization and detection (SELD) plays a vital role in understanding the environment. Recently, the SELD problem has received increasing interest from the research community. The state-of-the-art models for the DCASE 2020 for the SELD task have achieved good performance in terms of accuracy, but these models are generally based on multistage deep neural networks, which require substantial computational power and memory, making them unsuitable for deployment on low-cost Internet of Things (IoT) edge devices. In this paper, we propose a single-stage convolutional recurrent neural network (SS-CRNN) designed for real-time implementation of SELD on resource-constrained devices like Raspberry Pi. We also conducted a comprehensive analysis of different feature representations in terms of both accuracy and computational efficiency. Our results demonstrate that the SS-CRNN outperforms other models on the DCASE 2020 SELD dataset in terms of real-time factor (RTF), with only a slight trade-off: a 1% reduction in frame recall performance and a 6-degree decrease in localization accuracy compared to state-of-the-art methods. Additionally, we employ SpecMix augmentation to further enhance the model’s performance, which helps to boost our model’s performance during training.

https://doi.org/10.5281/zenodo.17818760

Downloads

Published

2025-04-30

How to Cite

Syed Owais Shah, Muneeba Darwaish, Mohsin Shah*, Muhammad Asad Khan, & Muhammad Shujaat. (2025). REAL-TIME SOUND EVENT LOCALIZATION AND DETECTION USING IOT RASPBERRY PI DEVICES BASED ON SINGLE-STAGE CRNN. Spectrum of Engineering Sciences, 3(4), 1020–1033. Retrieved from https://thesesjournal.com/index.php/1/article/view/1603