$$\rightleftharpoonup{xx}$$
$$\longleftharp{xx}$$,
$$\longrightharp{xx}$$,
The study protocol was approved by the Sushruta Hospitals Ethics Committee, Hubballi, Karnataka, India (EC registration number: ECR/372/INST/KA/2013/RR-19), and written informed consent was obtained from all participants prior to enrollment.
The FEA algorithm is a proprietary ML-based method for real-time analysis of facial microexpressions and upper-body movements to infer emotional states using the Navarasa framework. The Navarasa framework, derived from classical Indian aesthetic theory, delineates nine primary emotions: shringara (love), haasya (joy), karuna (sorrow/pensiveness), adbhuta (surprise), shanta (peace), raudra (anger), veera (courage), bhayanaka (fear), and vibhitsa (disgust). The FEA algorithm uses a machine-learning model trained on facial microexpressions and upper-body movements extracted from images of actors portraying the nine Navarasa emotions. The algorithm infers an individual’s emotional state by analyzing subtle facial dynamics and micromovements from visual input captured by a single RGB camera and generates quantitative scores for each of the nine primary emotions. This non-invasive configuration enables unobtrusive, naturalistic emotional assessment, making it suitable for a wide range of research and applied settings.
Model development
The FEA algorithm is based on a neural network (NN) architecture. The training dataset was collected in-house and consists of several hours of recorded video data of Indian actors portraying the nine Navarasa emotions. The dataset was derived from controlled video recordings of several actors/participants, with multiple recordings/tests conducted for each actor/participant across different elicited emotional states. Following pre-processing, including facial landmark detection, temporal consistency checks, quality filtering, and removal of unusable or low-confidence frames, approximately 25,000 structured training samples were retained. During dataset curation, recordings were stratified by more than ten demographic and recording-related parameters, including age, sex, and other relevant attributes, and intentionally balanced to achieve equitable representation across these parameters and minimize demographic bias. Sex distribution was maintained at approximately 52% male and 48% female. Participants broadly covered adolescent and adult age groups, with representation across younger, middle, and older adult cohorts. Actors/participants were based in India and represented multiple Indian regional and cultural backgrounds, as well as multiple international ethnicities present in India. Owing to proprietary restrictions, the exact number of actors/participants, actor/participant-level identity, exact video/recording counts per actor, precise frame counts, and the internal annotation methodology cannot be disclosed, as these form part of Nihilent Ltd.’s proprietary dataset and model-development intellectual property; this limitation is explicitly acknowledged in the Discussion section.
Facial images extracted from these recordings were processed using image processing techniques to derive quantitative features, which were subsequently used as inputs to a neural network for prediction. This model was then evaluated on a target population, and the insights gained were used to iteratively refine subsequent versions. This cycle of development, testing, and improvement was repeated until model performance stabilized. The final predictive model is implemented as a multilayer perceptron (MLP) comprising an input layer with approximately 300 selected features, hidden layers, and an output layer producing a nine-dimensional score vector corresponding to the nine Navarasa emotion categories. More than 700 features were initially derived from the captured data and subsequently reduced through feature selection and model optimization procedures. The MLP output score is then passed to proprietary post-processing algorithms to derive the final aggregated FEA-algorithm scores. Further architectural details—including the precise number of trainable parameters and specific layer dimensions—are proprietary and cannot be disclosed; this is acknowledged as a limitation of the present report. A detailed description of the model development is provided in the supplementary material (Supplementary File 1 and Supplementary Figure 1).
During iterative development, model performance was assessed using cross-validation on the in-house dataset, achieving 96% cross-validated classification accuracy across the nine Navarasa categories. Subsequent independent validation was conducted on more than 15,000 min of video data. A normalized 9 × 9 confusion matrix summarising classification performance across all nine Navarasa categories is provided in the supplementary material (Supplementary File 1).
Following stabilization, the model underwent controlled internal validation under expert supervision. This approach emphasizes empirical validation through direct human subject testing, enabling assessment of real-world applicability and robustness across diverse populations. The model development is summarized in Figure 1.
The patient Health Questionnaire-9 (PHQ-9) and the Generalized Anxiety Disorder-7 (GAD-7) scale: Participants completed self-report questionnaires under the supervision of a trained psychologist to assess their current emotional states. The instruments administered included the Patient Health Questionnaire-9 (PHQ-9), and the Generalized Anxiety Disorder-7 (GAD-7) scale. A score above 9 on either scale was considered clinically significant.
TRENDS
The Tool for Recognition of Emotions in Neuropsychiatric Disorders (TRENDS) is a culturally validated, ecologically sensitive instrument designed to assess facial emotion recognition abilities11. It has been validated for use in Indian populations and consists of 40 photographs of actors depicting the universal basic emotions of neutrality, happiness, fear, anger, and sadness. Images from TRENDS were used as reference stimuli to validate the FEA algorithm’s emotion classification outputs.
Experimental design
The experiment was designed to elicit an algorithmic response to emotional stimuli. Two emotion-elicitation paradigms were developed and used for this study.
TRENDS images paradigm
In the first paradigm, a video sequence initially presented a blank screen for 60 s, followed by the display of emotional stimuli. The emotional stimuli were selected from the TRENDS tool. A total of 20 images were presented, representing neutrality, happiness, sadness, anger, and fear × two genders × two age groups. Each stimulus was displayed for 5 s. The total stimulus presentation time was therefore 20 × 5 s, with an additional 10 s blank screen interval between consecutive images, resulting in a total duration of 300 s. During each 10 s blank interval, participants were required to identify the emotion depicted in the preceding image by selecting one of the five emotion labels displayed on the screen.
Navarasa paradigm
A second, analogous video paradigm was constructed based on the Navarasa framework of nine rasas. A total of 18 images were presented (nine emotions × two genders). Each stimulus was displayed for 5 s, with a 10 s blank screen between images, yielding a total duration of 270 s. The nine emotions were categorized as neutral (shanta), pleasant (haasya, veera, shringara, adbhuta), or unpleasant (Krodha, vibhitsa, bhayanaka, karuna). During each 10 s blank interval, participants were required to classify the emotion depicted in the preceding image by selecting one of three options presented on the screen: neutral, pleasant, or unpleasant. The nine Navaras were grouped into three categories to facilitate the recording of participants’ responses to emotion identification.
Participants
Healthy adult volunteers were recruited from workplace settings and educational institutions. Inclusion criteria comprised individuals aged 18–50 years, of any sex/gender. Exclusion criteria included a history of diagnosed psychiatric disorders, neurological disorders, significant visual impairment, or physical disabilities that restricted head and neck movements.
Procedure
Participants were seated such that the camera was at eye level and positioned 60–70 cm away. The FEA algorithm was used in the study. The application was opened on the laptop, and the researcher entered the subject ID. Throughout the experiment, the FEA algorithm continuously recorded video data to enable subsequent analysis of facial expressions and micro-movements. Initially, a 60 s baseline recording was obtained, during which participants were instructed to relax and refrain from overt expressive behaviors. Following baseline acquisition, participants were exposed to the two video paradigms (TRENDS images followed by Navarasa). All participants viewed the paradigms in the same fixed order to minimize potential confounding effects related to presentation sequence; however, this fixed order could have introduced order effects. During the recording sessions, the researcher remained outside the room to reduce distraction. The researcher monitored the experiment remotely via a mobile application while outside the room. Participants were offered a brief rest interval between the two paradigms (Figure 2).
Analysis
The collected data were mapped to the subject ID entered in the application to ensure data fidelity. The FEA algorithm utilizes a structured, end-to-end computational pipeline for automatic emotion inference, comprising five principal stages: video acquisition, feature detection and extraction, data preprocessing, emotion classification, and output interpretation. Extracted features underwent a series of pre-processing operations, including noise reduction, normalization across subjects and illumination conditions, and interpolation of missing frames. Temporal smoothing was applied to enhance signal stability. The resulting feature vectors were temporally aligned for downstream ingestion by the classification model. These time-aligned, denoised feature matrices served as inputs to the emotion classification component.
Feature extraction
The extraction framework employs a multi-feature extraction-stage algorithmic pipeline that translates standard 2D monocular video image feeds into a reconstructed 3D coordinate space. Rather than relying on raw 2D pixel tracking, the system maps localized spatial coordinates across the face and upper body. The features derived fall into three broad categories: (i) facial landmark regions, encompassing geometric descriptors of facial structure and dynamics (including eye-region cues and facial micro-expression correlates drawn from the Navarasa framework, which is rooted in the Natyashastra—the earliest formalized framework of human emotional expression, recognized by UNESCO); (ii) head-pose information, capturing orientation and movement of the head; and (iii) upper-body movement features, capturing resultant movement of the upper body. Specific feature definitions, including precise landmark indices and keypoint identifiers, are proprietary and cannot be fully disclosed; this is acknowledged as a limitation. The system does not use Facial Action Coding System (FACS) action unit labels directly; instead, it derives geometric and temporal features from facial and body movement. These 3D spatial transformations are aggregated over time to generate a comprehensive behavioral summary vector, which serves as input to the MLP classifier. In total, more than 700 candidate features were initially derived from the captured data and subsequently reduced to approximately 300 selected features through feature selection and model optimization. The structural pipeline proceeds from video acquisition through 2D-to-3D spatial reconstruction, temporal aggregation, and feature normalization prior to classification.
Emotion classification model
The pre-processed feature vectors were provided to a supervised multilayer perceptron (MLP) neural network architecture trained on the in-house annotated dataset corresponding to the nine Navarasa emotional categories: Śṛṅgāra (love), Hāsya (joy), Karuṇā (sorrow), Raudra (anger), Vīra (courage), Bhayānaka (fear), Bībhatsa (disgust), Adbhuta (wonder), and Śānta (peace). The model was explicitly designed to capture spatiotemporal dependencies in expressive behavioral data. At each time step, the architecture produced a vector of predicted emotion intensities for the nine categories.
Post-processing of model outputs
The raw model outputs were subsequently post-processed to yield stable and interpretable emotion scores. Three sequential operations were applied: (i) temporal smoothing to reduce short-term fluctuations, (ii) normalization to ensure comparability across time and across emotion dimensions, and (iii) aggregation to derive summary measures over specified temporal windows or entire sequences. The post-processing procedures ensure cross-subject comparability and control for inter-individual baseline variation.
Output interpretation and visualization
The FEA algorithm generates quantitative and visual representations of emotional dynamics inferred from video data. The two primary forms of output are temporal emotion trajectories. Emotion intensity values are visualized as time-series plots showing the evolution of each emotion throughout the video. This representation provides a continuous view of affective variation, facilitating the examination of transient changes, peak activations, and recovery or de-escalation phases in emotional arousal. Comparative emotional scores: Aggregated emotion values for each of the nine Navarasa categories are displayed as comparative score plots, illustrating the relative magnitude of each emotion across the analyzed interval. These comparative visualizations enable direct comparisons across emotions and provide a concise summary of the overall affective profile. The FEA-algorithm scores generated at baseline were inspected for any flatline or absence of meaningful change, which would indicate inadequate recording.
Statistical analysis
Scores for the nine Navarasa emotion domains, generated by the FEA algorithm during the 10 s interval following each image presentation in the paradigm, were treated as outcome variables. FEA-algorithm scores were generated for each s in the 10 s window (T0, T1, T2… T9). The FEA-algorithm score at T0 was considered the baseline score. The highest FEA algorithm score among the recordings from T1–T9 was defined as Tmax. The change in FEA-algorithm score from baseline (Tmax − T0) was computed and used as the outcome variable in the analyses. The change in FEA-algorithm scores over the time series for each emotion is depicted in Figure 3A–D.
For the TRENDS images, the corresponding target emotion categories were mapped as follows: Shanta > Neutral; Haasya > Happy; Krodha > Anger; and Bhayanaka > Fear. For the Navarasa images, FEA-algorithm scores were grouped and averaged into three composite categories: neutral (Shanta), pleasant (Haasya, Veera, Shringara, Adbhuta), and unpleasant (Krodha, Vibhitsa, Bhayanaka, Karuna).
Both group-level and subject-level analyses were conducted.
Group-level analysis: For each emotion category described above, the change in FEA-algorithm score from baseline to Tmax was compared using paired t-tests. A significant change from baseline in the FEA-algorithm score would indicate good concordance with emotion recognition. However, a significant difference would not necessarily indicate strong classification accuracy.
Subject-level analysis: Any increase in an FEA-algorithm emotion-domain score above the baseline value was classified as a positive emotional response. Accuracy indices were then derived—sensitivity and specificity. Sensitivity and specificity quantify, respectively, the probability that the test is positive when the target condition is present (+) and negative when it is absent (−). When both estimates approach 1, the test performs well in terms of diagnostic accuracy. Estimated sensitivity and specificity were defined as the proportions of participants with and without the reference condition (Navarasa expressions), respectively, who were correctly classified.
A subject was considered to have the target condition present (+) if they correctly identified at least 75% of the depicted emotions. This classification was based on validation studies of TRENDS, in which an accuracy of 60–80% was considered adequate discriminant validity. Receiver operating characteristic (ROC) analysis for a continuous predictor was performed to estimate the area under the curve (AUC) and the corresponding sensitivity and specificity values. A sensitivity above 80% was considered good discriminative ability, and below 80% was considered modest.