Cardiovascular risk stratification in the community based on a proteomics panel identified by machine learning survival analysis
Scientific Reports, 2026
Santana E., Wils E., Ntalianis E., Cauwenberghs N., Kuznetsova T.
| Disease area | Application area | Sample type | Products |
|---|---|---|---|
CVD | Patient Stratification | Serum | Olink Target 96 |
Abstract
Proteomics offers potential biomarkers that reflect underlying pathophysiological processes associated with cardiovascular (CV) outcomes. We aimed to identify proteins predictive of CV adverse events using machine learning survival analysis methods and to define clinically relevant protein-based clusters. We analyzed data from 768 community-based participants (mean age: 58.0 ± 10.8 years; 49.7% women) from FLEMENGHO who underwent echocardiography and proteomic profiling. CV events were tracked over a median of 10.9 years. Using Random Survival Forests and Gradient Boosting Survival Analysis, 15 proteins were among the most predictive for CV events. They were involved in key CV biological processes, including extracellular matrix organization (MMP-12), angiogenesis (ANGPT1, PlGF), blood pressure regulation and kidney function (KIM-1, REN, ACE2, BNP), apoptosis (TRAIL-R2), autophagy (CTSL1), proteolysis (PAPPA), inflammatory and immune response (CD40-L, Dkk-1, Gal-9, RAGE), and metabolism (AGRP). Using Gaussian Mixture Model based on those proteins levels, three phenogroups were identified. Cluster 3 exhibited the worst baseline CV risk profile. Although clusters 1 and 2 consisted of relatively younger individuals, cluster 2 was associated with significantly higher CV risk (HR 1.77, 95% CI 1.27–2.46, p < 0.001). Differences in cardiac structure and function were noted cross-sectional and longitudinally. Therefore, we identified proteins predictive of CV events that stratified CV risk in the community.