DeepAgeEmotionNet: A Multi-Task Deep Learning Framework for Speech-Based Age and Emotional State Assessment
Main Article Content
Abstract
The task of automatic speaker profiling based on speech signals becomes increasingly crucial in human-computer interaction, clinical voice assessment, and affective computing. Still, the tasks of speaker age group classification and emotion recognition are usually studied separately despite similar acoustic characteristics of both tasks. In this paper, we propose DeepAgeEmotionNet, a multi-task deep neural network architecture which learns to classify age groups and speech emotions in an integrated manner by utilizing a shared feature representation. The network merges log-Mel spectrogram and MFCC features via a spectro-cepstral projection block, learns local time-frequency patterns via convolutional layers, and captures temporal dependencies in speech signal via a two-layer BiLSTM network. Task-specific feature representations for age and emotion classification are formed via attention and gated fusion mechanisms. The network was evaluated with speaker-independent five-fold cross-validation experiments conducted on RAVDESS, CREMA-D, TESS, SAVEE, and EAS-Corp datasets. DeepAgeEmotionNet achieves 91.3% age-group accuracy and 88.7% emotion weighted F1-score, outperforming seven baselines while having less number of parameters than the most successful baseline based on Transformers.
