Can a GAN improve training? An honest experiment

Loading the results…

The idea. A GAN is two networks playing a game: a generator invents synthetic rows of patient data, a discriminator tries to spot the fakes. They improve until the fakes are statistically indistinguishable from real records — a way to multiply scarce medical data. The tabgan protocol adds a trick: filter the synthetic rows by how closely they resemble a held-back slice of real rows, without using its labels.

The setup, made trustworthy. Five folds on unique profiles (exact duplicate rows dropped first); CTGAN fitted only on part of each fold's training split (labels included, the standard practice); synthetic rows keep their generated labels; adversarial RF filtering against a separate slice of the training split — the held-out test fold is never touched before scoring. Then the deployed ensemble recipe trained on (real + synthetic) and scored on data it never saw. Sibling arms: plain baseline, GAN-generated-data-only (stress test), and classic minority-class oversampling (duplicating the smaller class until balanced).

Mean out-of-fold metrics per augmentation strategy

Mean across 5 folds ± std. ROC-AUC is the headline; Brier is a loss (lower is better). Dots under each row are the 5 per-fold values, so you see the variance, not just the mean.

The honest answer

Computing…

Reading the result

Computing…

Why this matters more than a win would

With so few unique profiles, the model already extracts nearly all the signal — its out-of-fold AUC leaves almost no gap for a generator to fill. Synthetic rows mostly re-encode existing patterns: a generator good enough to add genuinely new information is indistinguishable from having more real data, which is exactly what we don't have. And on medical tabular data the generator can amplify the minority class's noise. The tabgan paper itself found sampling strategies improve some datasets and hurt others (their Table 1.2); if ours lands on the "hurts" side, knowing that precisely beats pretending a fancy technique helped.

What WOULD help this dataset: more real records, blood-test features (HbA1c/glucose), and symptom scales instead of Yes/No. Generative tricks don't manufacture information that was never collected.

Experiment code: train/gan_experiment.py in the repo.