Yang · European journal of cardiovascular nursing 2026 · multicenter observational cohort study · n=?

The voice of decompensation: an explainable machine learning approach to assessing heart failure severity.

Level 3 - non-randomized controlled study

Multicenter observational cohort study evaluating a predictive machine learning model.

PubMed 41495965 · doi:10.1093/eurjcn/zvag003 · record verified 2026-08-26

What was done

Patients hospitalized with acute heart failure (AHF) were recruited across four centers. Voice recordings were gathered using professional devices and consumer smartphones (Apple, Huawei, Vivo) across 11 speech tasks. Researchers extracted 11 acoustic feature categories (e.g., Mel frequency cepstrum coefficients, chroma, spectral, and glottal features) and trained XGBoost models with Shapley Additive Explanations (SHAP) to classify clinical status from admission to discharge and mid-hospitalization to discharge, selecting four tasks for the final model.

What was found

The admission-to-discharge model achieved an accuracy of 0.76, sensitivity of 0.86, specificity of 0.65, and an AUC of 0.77. The mid-hospitalization-to-discharge model showed an accuracy of 0.81, sensitivity of 0.89, specificity of 0.71, and an AUC of 0.80. SHAP analysis showed that clinical recovery tracked with shifts from unstable, noisy vocal patterns to more stable, harmonic-rich, and brighter expressions. Consistency across smartphone brands was high.

Why it matters

This demonstrates the feasibility of using standard smartphone voice recordings to track heart failure status non-invasively, providing a potential tool for nurse-led remote triage and post-discharge monitoring.

Limits

The total sample size and patient demographics are not reported in the abstract. Specificity was relatively low (0.65 to 0.71), and the model was evaluated only in hospitalized patients during clinical recovery rather than in a true ambulatory outpatient cohort.