Emotionnet: A Deep Learning Framework for Real-Time Video-Based Facial Emotion Recognition

Main Article Content

Amit Rehapade, Prafulla Ajmire, Kiran Varma

Abstract

EmotionNet is a neural network that is used to recognize facial emotions in real-time and under dynamic and unconstrained conditions using videos. The suggested system combines both spatial features extraction and time modeling to ensure the capture of facial appearance and motion-based emotion features. A convoluted neural network (CNN) backbone that is lightweight is used to extract discriminative spatial characteristics of successive video frames and a bidirectional long short-term memory (Bi-LSTM) network is used to learn the time related characteristics of frame sequences. In order to increase resistance to changes in illumination, face alignment, and pose, the framework will include data augmentation, face alignment, and attention-based feature refinement. EmotionNet is designed to be used in the real-time surveillance systems and with edge devices, which is optimized to react to the low-latency inference. The model is trained and tested with benchmark facial emotion datasets through a multi-class classification model of the emotions of happiness, sadness, anger, fear, surprise, disgust, and neutrality. It is proven by experimental success that it performs better than standard CNN and single recurrent models with better accuracy, macro F1-score and temporal stability with less computational cost. Moreover, the role of temporal modeling and attention mechanisms in the improvement of recognition is justified by ablation experiments.

Article Details

Section
Articles