Voice-based prediction of prediabetes using classical machine learning models.
Level 4 - case-series / case-control
Cross-sectional diagnostic machine learning development and validation study
PubMed 41395633 · doi:10.3389/fcdhc.2025.1697769
What was done
Participants were recruited from clinical sites in India and a community college in Canada. Glycemic status was determined by HbA1c levels. Participants recorded a standardized spoken phrase multiple times daily through a mobile application. Audio was preprocessed to remove silence and uninformative sections, yielding 167 extracted acoustic features that were averaged per participant. Sex-specific models were evaluated under six configurations varying by dataset balance (age/BMI-matched vs. unbalanced) and BMI inclusion. Following LASSO feature selection and SMOTE oversampling, 12 machine learning classifiers were evaluated via leave-one-subject-out cross-validation on the Indian dataset, followed by testing on an internal holdout set and the independent Canadian dataset.
What was found
In cross-validation on the Indian dataset, the best female model (XGBoost, balanced, without BMI) achieved a balanced accuracy of 0.78, while the best male model (Random Forest, balanced, without BMI) achieved a balanced accuracy of 0.68. On the Indian holdout set, the male XGBoost model trained on an unbalanced dataset outperformed the cross-validated model. On the independent Canadian dataset, the models failed to generalize effectively, with several configurations unable to correctly identify prediabetic participants. The abstract does not report participant counts or sensitivity and specificity numbers.
Why it matters
This study shows that while voice features may detect prediabetes within a single geographic cohort, models trained in one demographic population fail to translate across different geographic or cultural settings.
Limits
The abstract does not report the total sample size (n), participant demographics, or detailed diagnostic performance metrics like sensitivity, specificity, and AUC. The external validation failed, indicating vulnerability to acoustic confounding from differences in language, accent, recording environment, or population demographics.