ENHANCING SPEECH EMOTION RECOGNITION WITH DEEP LEARNING THROUGH DATA FUSION, SPECTROGRAM AUGMENTATION, AND HYBRID FEATURE INTEGRATION

Authors

  • Muhammad Talha Jahangir
  • Mujahid Hussain
  • Nashitah Alwaz
  • Muhammad Musawir Saeed
  • Waheed Ahmad
  • Uzair Ahmad
  • Hammad Toheed Khan

Keywords:

Deep Learning, Data Fusion, Convolutional Neural Network, Bidirectional Long Short-Term Memory, Human-Computer Interaction (HCI), Spectrogram Augmentation, Speech Emotion Recognition, MFCC, Mel Spectrogram, Root mean square

Abstract

Speech Emotion Recognition (SER), which lets computers decode human feelings using vocal clues, is among the most vital elements of affective computing. The range of speech patterns, lack of data, and difficulty of emotional expression make it still difficult to get excellent SER accuracy. Data fusion from four baseline datasets RAVDESS, TESS, CREMA-D, and SAVEE is used by our proposed deep learning-based SER architecture. The suggested model design combines Convolutional Neural Networks (CNNs) with Bidirectional Long Short-Term Memory (BiLSTM) to efficiently capture spatial and temporal characteristics. With a remarkable classification accuracy of 98%, the proposed framework improving SER performance and giving computers the ability to immediately detect and respond to human feelings that helps our system foster a more sympathetic and flexible human-computer connection.

Downloads

Published

2025-09-24

How to Cite

Muhammad Talha Jahangir, Mujahid Hussain, Nashitah Alwaz, Muhammad Musawir Saeed, Waheed Ahmad, Uzair Ahmad, & Hammad Toheed Khan. (2025). ENHANCING SPEECH EMOTION RECOGNITION WITH DEEP LEARNING THROUGH DATA FUSION, SPECTROGRAM AUGMENTATION, AND HYBRID FEATURE INTEGRATION. Spectrum of Engineering Sciences, 3(9), 1048–1067. Retrieved from https://thesesjournal.com/index.php/1/article/view/1100