Lyu · BMC medical informatics and decision making 2026 · Diagnostic prediction model development and validation study · n=553

Multimodal deep learning for cardiovascular disease detection using pulse wave and vocal signals: a prediction model development and validation study.

Level 3 - non-randomized controlled study

Diagnostic prediction model development with independent cohort validation

PubMed 42032577 · doi:10.1186/s12911-026-03518-w · record verified 2026-08-26

What was done

Radial artery pulse waves and vocal signals (sustained /a:/ phonation) were recorded under standardized conditions from 553 participants (healthy, coronary artery disease, and heart failure). Researchers evaluated single-modality and multimodal fusion models across four deep learning architectures (MLP, GAN-Discriminator, ResNet-MLP, Bi-LSTM) using ten-fold cross-validation, validated the fused model on an independent external cohort of 194 subjects, and analyzed feature importance via SHAP and LIME.

What was found

In the independent external validation cohort (n = 194), the fused multimodal model achieved an accuracy of 0.7165 and an AUC of 0.8454. Model predictions relied on both pulse dynamic variables (e.g., t3/tmax) and vocal acoustic biomarkers (e.g., MFCCs).

Why it matters

This study demonstrates the feasibility of combining digitized radial pulse waveforms and vocal acoustics into a non-invasive AI screening tool for coronary artery disease and heart failure in primary care settings.

Limits

The total sample size is modest (553 subjects). The abstract does not report sensitivity, specificity, positive predictive values, or the baseline demographic and clinical characteristics of the cohorts. An accuracy of 71.65% entails a substantial rate of misclassification, and real-world clinical implementation was not evaluated.