The numbers, honestly measured

Every algorithm below was scored the same unforgiving way: repeated 5-fold cross-validation (5 folds × 4 repeats) on unique profiles — exact duplicate questionnaire rows are dropped first, so no test row is a copy of a training row — and 95% confidence intervals from 2,000 bootstrap resamples. No accuracy alone, no single lucky split.

Cross-validated performance of 10 algorithms and a stacking ensemble on the diabetes dataset

↓ Brier is a loss: lower is better. Sensitivity = recall on positive cases — in screening this matters most, because a missed case is the expensive error.

How to read this

Loading benchmark data…

ROC-AUC measures ranking quality across all thresholds (1.0 = perfect). Brier rewards being well-calibrated — saying 70% only when it's truly 70%. The ★ row is the model this site ships: a stack of five experts whose stacker is fit on out-of-fold expert probabilities. These cross-validation numbers describe the recipe — they are not a test of the exported weights, which were refit on all profiles.