A Novel Machine Learning-Driven Voice and Clinical Biomarkers Framework for Robust Prediction of Type 2 Diabetes Mellitus.
Level 3 - non-randomized controlled study
Cross-sectional machine learning derivation and validation study comparing individuals with and without type 2 diabetes.
PubMed 41087282 · doi:10.1016/j.jvoice.2025.09.033
What was done
Researchers developed and evaluated machine learning models to screen for type 2 diabetes mellitus (T2DM) using data from 3,129 individuals (1,158 with T2DM and 1,971 without). Voice recordings were analyzed using the openSMILE toolkit to extract 88 acoustic features across prosodic, spectral, cepstral, and voice quality domains, alongside 30 clinical and biochemical variables. Multiple feature selection methods and machine learning algorithms (including XGBoost, Random Forest, TabNet, and TabTransformer) were evaluated using cross-validation and an independent test set, with model interpretability assessed via SHAP analysis.
What was found
Models trained on clinical features alone reached an AUC of approximately 69%. Acoustic-only models performed higher, with a LASSO plus XGBoost model achieving an AUC of 80.8%. The multimodal fused model (combining acoustic and clinical features) achieved the highest performance with 94.1% accuracy, 93.6% recall, and an AUC of 95.2%. SHAP analysis identified HbA1c, fasting blood glucose, HOMA-IR, and acoustic parameters including jitter and shimmer as the most influential predictive features.
Why it matters
This study shows that automated acoustic voice analysis carries a diagnostic signal for T2DM and can substantially boost classification performance when paired with clinical baseline data.
Limits
Key top-contributing features in the fused model included standard diagnostic blood tests (HbA1c, fasting glucose, HOMA-IR), which limits the noninvasive screening utility of the combined model over standard diagnostics. The abstract does not report confidence intervals, demographic characteristics of the cohort, recording environment generalizability, or prospective validation in an undiagnosed real-world screening setting.