Experimental results for age and gender classification were evaluated using MFCC-based and hybrid feature representations with ML and DL models. The dataset was partitioned into training and testing sets at an 80:20 ratio, with 80% used for training and 20% for testing. Speaker-level partitioning was applied to prevent recordings from the same speaker from appearing in both sets.
The hyperparameters and implementation details of the proposed framework are summarized in Table 3. The experiments were conducted on a system equipped with an Intel i7-8265U CPU, 16 GB of RAM, and an NVIDIA Tesla K80 GPU running Windows 11. Python and the Scikit-learn library were used to develop the ML models, whereas the proposed DL framework was implemented using PyTorch. These models were evaluated for gender and age classification from voice data.
The Adam optimizer was used with an initial learning rate of 1 × 10⁻4, β₂ = 0.999, and ε = 1 × 10⁻8. A batch size of 64 and a maximum of 50 epochs were used, with early stopping set to a patience of four epochs to reduce overfitting. A dropout rate of 0.3 was applied to the fully connected layers for regularization. Categorical cross-entropy was used as the loss function for the classification tasks. The ReduceLROnPlateau learning-rate scheduler was used to monitor the validation loss.
To prevent data leakage and ensure a fair evaluation, speaker-level splitting was enforced so that no speaker appeared in both the training and testing sets. In addition, five-fold cross-validation was performed to assess the generalizability and robustness of the proposed model.
Results of models using the original dataset
Table 4 presents the performance of several ML and DL models for gender classification on a dataset comprising 28 languages, without MFCC features. Precision, recall, F1-score, and accuracy were used as the evaluation metrics.
Among the evaluated models, EfficientNet-Lite achieved the highest accuracy of 83.48%, with a precision 82.87%, a recall 84.28%, and an F1-score 83.54%. MobileNet and InceptionV3 also performed well, with MobileNet achieving 81.33% accuracy and 83.11% F1-score, and InceptionV3 achieving 80.77% accuracy and 82.52% F1-score. The CNN model achieved an accuracy of 75.71% and an F1-score of 77.43%, while ResNet achieved an accuracy of 74.37% and an F1-score of 75.54%. SVM, RF, LR, and ETC achieved accuracies ranging from 71.42% to 73.59% and F1-scores ranging from 72.48% to 74.34%. VGG19 achieved a comparatively lower accuracy of 70.56%, with a recall of 75.07% and an F1-score of 74.17%. Overall, EfficientNet-Lite and MobileNet achieved the strongest performance among the evaluated models when using the original features.
Results of models using MFCC features
MFCC features were subsequently evaluated for gender classification. Table 5 compares the performance of the ML and DL models using MFCC features on the dataset comprising 28 languages.
EfficientNet-Lite achieved the highest performance, with an accuracy of 92.64%, precision of 93.79%, recall of 94.92%, and F1-score of 94.84%. MobileNet achieved an accuracy of 91.39% and an F1-score of 91.07%, whereas InceptionV3 achieved an accuracy of 90.24% and an F1-score of 89.83%. ResNet achieved an accuracy of 89.43% and an F1-score of 83.67%, while CNN achieved an accuracy of 88.60% and an F1-score of 87.35%. VGG19 achieved an accuracy of 87.48% and an F1-score of 85.58%.
Among the traditional ML models, SVM achieved an accuracy of 85.47% and an F1-score of 84.29%; RF achieved an accuracy of 84.57% and an F1-score of 87.42%; LR achieved an accuracy of 83.25% and an F1-score of 84.98%; and ETC achieved an accuracy of 83.59% and an F1-score of 85.43%. Overall, the inclusion of MFCC features improved the performance of most models compared with the results obtained using the original features. EfficientNet-Lite achieved the highest overall performance in this experiment.
Results of models using hybrid MFCC–wav2vec 2.0 features with the English-language dataset
The next set of experiments evaluated gender classification using the hybrid MFCC–wav2vec 2.0 feature representation on the English-language subset. The performance of the evaluated models is presented in Table 6. EfficientNet-Lite achieved the highest performance, with an accuracy of 97.87%, precision of 98.61%, recall of 98.70%, and F1-score of 98.65%.
Among the traditional ML models, SVM achieved an accuracy of 92.58% and an F1-score of 91.52%; ETC achieved an accuracy of 91.49% and an F1-score of 92.79%; LR achieved an accuracy of 90.64% and an F1-score of 91.34%; and RF achieved an accuracy of 88.36% and an F1-score of 90.86%. Overall, the models demonstrated strong performance on the English-language subset, with EfficientNet-Lite achieving the highest values across the reported evaluation metrics.
Results of five-fold cross-validation
Five-fold cross-validation was performed to evaluate the stability and generalizability of the proposed framework. The cross-validation results obtained using the English-language voice dataset are presented in Table 7. EfficientNet-Lite achieved a standard deviation of ±0.0033 across the five folds. The low variability across folds indicated stable performance under the evaluated cross-validation setting.
Age classification using voice data
In addition to gender classification, age classification was evaluated using multiple models and MFCC features. Table 8 presents the performance of the evaluated models for age classification.
The results showed that the transfer-learning models achieved accuracies above 85%, with MobileNet achieving 85.23% and InceptionV3 achieving 86.67%. EfficientNet-Lite achieved the highest reported performance, with an accuracy of 88.42%, precision of 86.56%, and recall of 87.21%.
Validation of the proposed framework
To evaluate the proposed framework on an external dataset, additional experiments were conducted using the Mozilla Common Voice dataset51. The dataset contained approximately 500 hours of crowdsourced speech recordings with associated demographic metadata, including age, gender, and accent information. The corpus was organized into predefined training, development, and test subsets. Audio recordings were processed using the same preprocessing pipeline as that used in the primary experiments, including resampling to 16 kHz, extraction of 40-dimensional MFCC features, and generation of 768-dimensional wav2vec 2.0 embeddings. The two representations were concatenated to produce the 808-dimensional hybrid feature vector used by VoiceNet-RAI. The validation results are presented in Table 9.
EfficientNet-Lite with the hybrid low- and high-level feature representation achieved a gender classification accuracy of 98.27%, higher than that of the other evaluated models. Among the ML models, RF achieved 90.47% accuracy, SVM 90.69%, and LR 89.75%. Among the other DL models, ResNet achieved 92.19% accuracy, whereas VGG19 achieved 92.12%. The corresponding precision, recall, and F1-score values are reported in Table 9.
Ablation study on feature representation and fusion
An ablation study was conducted to evaluate the contribution of the feature representations used in the proposed framework. Four configurations were considered: raw/original features, MFCC-only features, wav2vec 2.0-only embeddings, and the proposed hybrid MFCC–wav2vec 2.0 representation. EfficientNet-Lite was used as the representative model because it achieved the highest performance across the evaluated experimental settings. The ablation results are presented in Table 10.
The baseline model using raw features achieved 83.48% accuracy, whereas the MFCC-based representation achieved 92.64% accuracy and 94.84% F1-score. The hybrid MFCC–wav2vec 2.0 representation achieved an accuracy of 97.87% and an F1-score of 98.65%. Under identical English-language data partitions and training conditions, the hybrid MFCC–wav2vec 2.0 representation achieved 97.87% accuracy, compared with 92.64% for the MFCC-only representation.
Controlled experiments were conducted using identical data partitions, training settings, and the EfficientNet-Lite architecture to evaluate the contribution of each feature representation. Raw/original features, MFCC-only features, wav2vec 2.0-only embeddings, and the hybrid MFCC–wav2vec 2.0 representation were compared under these conditions. This experimental design was used to assess the individual and combined contributions of the feature representations.
Statistical significance analysis
All feature configurations in this comparison were evaluated using the same English-language subset, identical speaker-level data partitions, and the same five cross-validation folds. The proposed model achieved significantly higher accuracy than the raw-feature (97.86 ± 0.23% vs. 83.48 ± 0.24%; t (4) = 384.32, p < 0.001), MFCC-only (97.86 ± 0.23% vs. 92.60 ± 0.32%; t (4) = 131.50, p < 0.001), and wav2vec 2.0-only configurations (97.86 ± 0.23% vs. 95.34 ± 0.24%; t (4) = 126.00, p < 0.001), confirming that the improvements from MFCC–wav2vec 2.0 fusion were statistically significant.
Responsible AI considerations and deployment guidelines
Responsible AI in voice-based demographic classification required consideration of fairness, transparency, privacy, potential misuse, and deployment-related risks in addition to predictive performance. Although subgroup analysis provided an empirical assessment of performance differences across the evaluated gender and language groups, fairness could not be established solely from similar classification accuracy. Therefore, the VoiceNet-RAI framework was evaluated from multiple Responsible AI perspectives.
Fairness and bias analysis
A fairness analysis was conducted across gender and language subgroups to evaluate the Responsible AI characteristics of the proposed VoiceNet-RAI framework. The results are presented in Table 11. For gender-wise evaluation, model performance was analyzed separately for male and female speakers. The results showed less than 1.5% variation in accuracy and F1-score between the evaluated gender groups.
For language-wise evaluation, performance was assessed across the evaluated language groups within the dataset of 28 languages. The results showed an average variation of less than 2% in accuracy across the evaluated groups, with greater variation observed for some lower-resource languages with fewer available samples.
Fairness was further assessed using the demographic parity difference, the equal opportunity difference, the False Positive Rate (FPR), and the False Negative Rate (FNR). Demographic parity difference quantified differences in positive prediction rates between groups, whereas equal opportunity difference quantified differences in True Positive Rates (TPRs). FPR and FNR were used to characterize subgroup-specific error patterns. These metrics complemented subgroup-level accuracy and F1-score, providing additional information on differences in model performance across the evaluated demographic and linguistic groups. Under the evaluated dataset and subgroup definitions, the results indicated relatively small performance differences across the assessed gender and language groups. However, these findings were limited to the demographic and linguistic groups represented in the evaluated datasets and did not establish fairness across unrepresented populations or deployment settings.
Explainability and transparency
VoiceNet-RAI combined interpretable MFCC-based acoustic features with contextual wav2vec 2.0 embeddings. The ablation study evaluated the contributions of the individual and combined feature representations. However, wav2vec 2.0 embeddings remained comparatively difficult to interpret. Post-hoc explanation methods could therefore be investigated in future studies to identify the features that contributed most strongly to individual predictions.
Privacy preservation
Because voice recordings contained sensitive biometric information, privacy-preserving deployment required data-minimization practices, including processing only the information required for the classification task and avoiding unnecessary storage of raw audio. Secure storage, controlled access, and local inference could further reduce privacy risks. Formal privacy-preserving approaches, including differential privacy, were not evaluated in the present study and remained an area for future investigation.
Responsible deployment
Deployment of VoiceNet-RAI would require human oversight, confidence-aware decision-making, and continuous performance monitoring. Low-confidence predictions should not be used independently for critical decisions, and model performance should be re-evaluated when deployment conditions differ substantially from those represented in the training and validation datasets.
Comparison with state-of-the-art techniques
A comparative evaluation of the proposed framework and previously published approaches is presented in Table 12. Studies published between 2019 and 2024 were considered, including approaches based on ML, DL, and transfer learning. The comparison included previously reported studies in the same application domain. As shown in Table 12, the proposed framework achieved higher reported accuracy for gender and age classification than the compared approaches. However, because the studies may have differed in datasets, preprocessing procedures, data partitions, and evaluation protocols, the comparison reflected reported performance rather than a controlled head-to-head experimental comparison.
DATA AVAILABILITY:
The raw speech data used in this study are publicly available from the BVC Challenging Voice Set hosted on GitHub and the Mozilla Common Voice dataset. The BVC dataset can be accessed at https://github.com/OgeNI/BVC_Challenging_Voice_Set, while the Common Voice dataset is available at https://www.kaggle.com/datasets/mozillaorg/common-voice/data.

Figure 1: Architecture Diagram. The complete workflow methodology of the proposed framework. Please click here to view a larger version of this figure.
| References | Models | Datasets | Results |
| 19 | GBC, DT, RF, SVM, NN | VoxForge dataset | GBC 90% |
| 20 | Robustscalar, PCA, and Logistic
regression, Sequential model | voice.mozilla.org | Sequential Model 91%. |
| 21 | CNN, CRNN, TCN | Mozilla Voice dataset | 80% CNN |
| 22, 23 | Backpropagation, DT, Bagging,
RF, KNN, DNN, DL | Kaggle, Mozilla Voice dataset,
Speech Accent Archive | Kaggle: Bagging, KNN and RF
98.10%. Voice common: Bagging 55.39%. Speech accent: Bagging 78.94%. |
| 24 | DNN, MLP, KNN, SVC, DT, RF,
SVC RBF kernel, ADA, QDA, GNB, ResNET34 and ResNET50. | Mozilla Common Voice | ResNET50 98.57%. |
| 25 | ANN, SVM, RF, DT, NB, KNN | OGI Kids Speech corpus, Private
corpus, Specific Dataset, | For Gender detection SVM 98%,
For age detection RF 95% |
| 26 | SVM, DA, NB, DT, KNN, Ensem-
ble, LightGBM | Mozilla Common Voice, VoxCeleb | LGBM 91% |
| 27 | KNN, SVM, LR, SGD, Stacked
model | Various gender sources | Stacked model 99.64% |
| 28 | DNN, SVM | Various gender sources | SVM 97% |
| 29 | SVM, RF, CART, KNN | Dataset contain 3169 instances | RF 93% |
| 30 | CNN | Nigerian subjects dataset | CNN 85.48%. |
| 31 | DT, RF, KNN, LR, SVM, ANN,
CNN2D | 3 dataset from kaggle | Gender Prediction: ANN 85.51%,
Age Prediction: ANN 89.20%. |
Table 1: Summary of existing studies on voice-based gender and age classification, including models, datasets, and reported performance.
| Dataset Characteristic | Description / Value |
| Total number of speakers | 526 |
| Male speakers | 336 |
| Female speakers | 190 |
| Total recordings | Approximately 3,964 |
| Number of languages | 28 |
| Language-wise distribution | English and 28 native languages (Afemai, Akoko-Edo, Annang, Efik, Ekoi, Fulani, Hausa, Ibibio, Idoma, Igala, Igbo, Igede, Ijaw, Ika, Ikom, Ikwerre, Ishan, Kaire-Kaire, Kanuri, Lokaa, Urhobo, Yoruba, Obudu, Ogoni, Okobo, Okirika, Tiv and Ukwani. The dominant tribe amongst these native languages in the dataset is Igbo.) |
| Age groups | three age groups: young people (teens and twenties), adults (thirties, forties and fifties), and seniors (sexagenarians, heptagenarians and octogenarians) |
| Age-group distribution | Teens13–19 yearsLow (< 10%)Twenties20–29 yearsVery High (~35–45%)Thirties30–39 yearsHigh (~20–30%)Forties40–49 yearsModerate (~10–15%)Fifties50–59 yearsLow (~5–10%)Sixties+60–80+ yearsSparse (< 5%) |
| Recording duration | 3-10 seconds |
| Recording conditions | 16 bit PCM |
| Original sampling rate | 44.1 kHz |
| Processing sampling rate | 16 kHz |
| MFCC frame length | 25 ms |
| MFCC frame overlap / step | 10 ms |
| MFCC representation | 39-dimensional temporally averaged vector |
| wav2vec 2.0 representation | 768-dimensional temporally mean-pooled vector |
| Final fused representation | 807 dimensions |
| Train–test split | 80:20 |
| Data partitioning | Speaker-independent splitting |
| Cross-validation | 5-fold cross-validation |
| External validation dataset | Mozilla Common Voice |
Table 2: Demographic characteristics, language coverage, data partitioning, and preprocessing settings of the speech dataset used in VoiceNet-RAI.
| Parameter | Value / Description |
| Framework | PyTorch |
| Hardware | NVIDIA GPU (CUDA-enabled) |
| Random seed | 42 |
| Optimizer | Adam |
| Initial Learning Rate | 1 × 10−4 |
| β₁ / β₂ | 0.9 / 0.999 |
| Adam epsilon | 1 × 10⁻⁸ |
| Batch Size | 64 |
| Epochs | 50 (Early stopping patience = 8) |
| Loss Function | Categorical Cross-Entropy |
| Normalization | Z-score normalization |
| Learning Rate Scheduler | ReduceLROnPlateau |
| Dropout Rate | 0.3 |
| Training time | 20min |
| Inference time/sample | 13 seconds |
| Feature Extraction | Feature Extraction |
| MFCC Coefficients | 40 |
| Window Size | 25 ms |
| Stride | 10 ms |
| wav2vec Variant | Base (pretrained) Embedding Dimension |
| Pooling Strategy | Mean pooling |
| Feature Fusion | Concatenation (MFCC + wav2vec) |
| MFCC Coefficients | 40 |
| Evaluation Protocol | Evaluation Protocol |
| Train-Test Split | Speaker-level separation |
| Cross-Validation | 5-Fold |
| Datasets | Mozilla Common Voice (28 languages + English subset) and BVC Gender & Age from Voice Challenging Dataset |
| Tasks | Gender and Age Classification |
Table 3: Training configuration, feature-extraction settings, hyperparameters, and evaluation protocol used for VoiceNet-RAI.
| Models | Accuracy | Precision | Recall | F1-score |
| RF | 71.42 | 72.79 | 74.43 | 73.52 |
| LR | 73.59 | 74.22 | 75.40 | 74.34 |
| ETC | 72.44 | 72.33 | 74.52 | 73.41 |
| SVM | 73.52 | 72.60 | 72.39 | 72.48 |
| CNN | 75.71 | 76.59 | 77.34 | 77.43 |
| VGG19 | 70.56 | 73.28 | 75.07 | 74.17 |
| ResNet | 74.37 | 73.49 | 76.68 | 75.54 |
| EfficientNet-Lite | 83.48 | 82.87 | 84.28 | 83.54 |
| MobileNet | 81.33 | 83.19 | 82.14 | 83.11 |
| InceptionV3 | 80.77 | 81.34 | 82.67 | 82.52 |
Table 4: Gender-classification performance of ML and DL models using the 28-language dataset without MFCC features.
| Models | Accuracy | Precision | Recall | F1-score |
| RF | 84.57 | 86.53 | 87.39 | 87.42 |
| LR | 83.25 | 84.42 | 85.64 | 84.98 |
| ETC | 83.59 | 84.25 | 86.36 | 85.43 |
| SVM | 85.47 | 83.37 | 85.26 | 84.29 |
| CNN | 88.60 | 86.75 | 88.46 | 87.35 |
| VGG19 | 87.48 | 85.49 | 85.80 | 85.58 |
| ResNet | 89.43 | 84.39 | 82.34 | 83.67 |
| EfficientNet-Lite | 92.64 | 93.79 | 94.92 | 94.84 |
| MobileNet | 91.39 | 90.64 | 91.15 | 91.07 |
| InceptionV3 | 90.24 | 88.81 | 90.70 | 89.83 |
Table 5: Gender-classification performance of ML and DL models using MFCC features on the 28-language dataset.
| Models | Accuracy | Precision | Recall | F1-score |
| RF | 88.36 | 90.91 | 89.94 | 90.86 |
| LR | 90.64 | 92.24 | 91.30 | 91.34 |
| ETC | 91.49 | 92.80 | 92.79 | 92.79 |
| SVM | 92.58 | 92.75 | 90.64 | 91.52 |
| CNN | 93.69 | 94.69 | 94.53 | 94.34 |
| VGG19 | 91.37 | 92.64 | 93.48 | 93.21 |
| ResNet | 95.28 | 96.39 | 95.52 | 95.43 |
| EfficientNet-Lite | 97.87 | 98.61 | 98.70 | 98.65 |
| MobileNet | 96.41 | 95.51 | 96.39 | 96.24 |
| InceptionV3 | 96.35 | 95.29 | 96.84 | 96.08 |
Table 6: Gender-classification performance of ML and DL models using the hybrid MFCC–wav2vec 2.0 feature representation on the English-language subset.
| Models | Accuracy | Precision | Recall | F1-score |
| First-fold | 97 | 93 | 96 | 94 |
| Second-fold | 96 | 94 | 96 | 95 |
| Third-fold | 97 | 94 | 95 | 95 |
| Fourth-fold | 96 | 93 | 95 | 94 |
| Fifth-fold | 97 | 94 | 96 | 95 |
| Average | 96.6 | 93.6 | 95.6 | 94.6 |
Table 7: Five-fold cross-validation performance of VoiceNet-RAI on the English-language subset.
| Models | Accuracy | Precision | Recall | F1-score |
| RF | 80.15 | 78.68 | 79.27 | 79.04 |
| LR | 81.28 | 80.69 | 80.43 | 80.54 |
| ETC | 80.58 | 81.21 | 80.82 | 81.04 |
| SVM | 82.54 | 81.31 | 82.14 | 81.78 |
| CNN | 85.27 | 84.39 | 84.47 | 84.43 |
| VGG19 | 83.73 | 82.97 | 83.53 | 83.48 |
| ResNet | 82.68 | 81.42 | 80.78 | 81.95 |
| EfficientNet-Lite | 88.42 | 86.56 | 87.21 | 86.97 |
| MobileNet | 85.23 | 85.64 | 85.95 | 85.81 |
| InceptionV3 | 86.67 | 85.86 | 86.78 | 86.35 |
Table 8: Age-classification performance of ML and DL models in terms of accuracy, precision, recall, and F1-score.
| Models | Accuracy | Precision | Recall | F1-score |
| RF | 90.47 | 92.82 | 90.85 | 91.32 |
| LR | 89.75 | 90.44 | 89.61 | 90.03 |
| ETC | 92.63 | 91.29 | 91.55 | 91.42 |
| SVM | 90.69 | 91.84 | 91.81 | 91.82 |
| CNN | 94.63 | 95.52 | 94.24 | 94.94 |
| VGG19 | 92.12 | 91.34 | 90.52 | 91.10 |
| ResNet | 92.19 | 91.98 | 92.27 | 92.25 |
| EfficientNet-Lite | 98.27 | 98.63 | 98.44 | 98.56 |
| MobileNet | 97.28 | 96.36 | 97.01 | 96.40 |
| InceptionV3 | 95.85 | 96.18 | 97.35 | 97.17 |
Table 9: External validation results for gender classification on the Mozilla Common Voice dataset using the hybrid MFCC–wav2vec 2.0 feature representation.
| Feature Configuration | Accuracy (%) | Precision (%) | F1-Score ( %) |
| Raw features (No MFCC) | 83.48 | 82.87 | 83.54 |
| MFCC only | 92.64 | 93.79 | 94.84 |
| wav2vec 2.0-only | 95.34 | 96.32 | 95.18 |
| MFCC + wav2vec (Proposed) | 97.87 | 98.61 | 98.65 |
Table 10: Ablation analysis of feature representations using EfficientNet-Lite.
| Evaluation Group | SubGroup | Accuracy (%) | F1-score (%) | FPR | FNR | Demographic Parity Difference | Equal Opportunity Difference |
| Gender | Male | 97.95 | 97.95 | 0.021 | 0.024 | _ | _ |
| Gender | Female | 97.62 | 98.50 | 0.024 | 0.027 | _ | _ |
| Gender disparity | Male vs. Female | _ | _ | _ | _ | 0.012 | 0.003 |
| High-resource languages | 98.10 | 98.80 | 0.018 | 0.021 | _ | _ |
| Low-resource languages | 96.85 | 97.90 | 0.031 | 0.034 | _ | _ |
| Language Disparity | High vs. Low | _ | _ | _ | _ | 0.028 | 0.014 |
Table 11: Fairness evaluation across gender and language-resource subgroups. Abbreviations: FPR = false positive rate; FNR = false negative rate. Demographic parity difference and equal opportunity difference represent between-group disparities, so they are reported for subgroup comparisons rather than individual groups. Lower disparity values indicate more consistent model behavior across groups.
| Reference | Approach | Accuracy |
| 16 | GBC | 90%. |
| 17 | Sequential Model | 91% |
| 18 | CNN | 80% |
| 19 | Bagging, KNN and RF | 98.10%. |
| 20 | ResNET50 | 98.57%. |
| 21 | SVM | 98% |
| 22 | LGBM | 91% |
| 23 | Stacked model | 99.64% |
| 24 | SVM | 97% |
| 25 | RF | 93% |
| 26 | CNN | 85.48% |
| 27 | ANN | 89.20% |
| This study | EfficientNet-Lite | 97.87% |
Table 12: Accuracy comparison of VoiceNet-RAI with previously reported state-of-the-art approaches for voice-based gender and age classification.