An Explainable ConvNeXt-V2 Architecture for Facial Emotion Recognition from Children with Autism under Real-Time Visual Environments
Main Article Content
Abstract
Autism Spectrum Disorder (ASD) is a neurodevelopmental disorder caused by differences in brain development that affect a child’s ability to communicate, interact socially, and interpret emotional cues. Individuals with ASD often experience difficulties in recognizing facial expressions, making emotion recognition a significant component for early screening and diagnosis. Moreover, conventional machine learning and deep learning methodologies face limitations such as poor generalization, high computational complexity, and difficulty in capturing subtle facial features. In addition, explainable artificial intelligence (XAI) approaches are incorporated to improve the transparency of model decisions. Consequently, there is a need for more robust and effective computational models that improve emotion recognition accuracy and support early ASD detection, ultimately helping to reduce social and emotional difficulties. This manuscript presents an Interpretable Explainable Deep Convolutional Framework-based Autism Facial Emotion Recognition (IDCF-AFER) framework. The proposed model aims to design an effective approach for the accurate identification and classification of facial emotions such as neutral, anger, fear, joy, sadness, and surprise. The proposed model employs a ConvNeXt-V2 framework to effectively extract deep spatial features from facial images. Moreover, the stacked sparse autoencoder with a softmax layer is used for classification, which enables the model to effectively learn discriminative feature representations. Furthermore, Harris Hawks Optimization (HHO) is applied to fine-tune the model parameters and improve classification performance. Eventually, Score-CAM is utilized to highlight significant regions in facial images. Extensive experimental evaluation of the IDCF-AFER model was conducted on the benchmark Autism Facial Emotion Recognition dataset, and the results demonstrate its superior performance with an accuracy of 95.39% compared to existing methods.
