Given only a spreadsheet of real ovarian-cancer patients, and no guidance, DataStoryMD was asked two questions: does it recover what the field already knows, and can it reproduce a specific published paper on that paper’s own cohort? It does both, and applies methodological safeguards the original analyses did not.
Cohort: the sequenced subset of TCGA-OV (n = 316), the exact patients analysed by Birkbak et al. Reproduction re-run and verified 2026-08-03; every figure below regenerates from a committed script, not a saved copy.
The paper also reports two raw median mutation counts. We do not compare those, on purpose: a median is a property of the exact variable measured, and our mutation count is measured differently from the paper’s, so comparing the two values would compare two different quantities. The models and tests carry no such problem, which is why all ten reproduce.
Before a tool can be trusted on a novel question, it should recover the canonical ones. From the raw records alone, DataStoryMD identified the accepted predictors of ovarian-cancer survival, each concordant with the literature, each adjusted for the others.
A forest plot the system produced from the raw cohort. Each row is a candidate predictor, shown as an adjusted hazard ratio with its 95% confidence interval on a log scale; points left of 1 indicate longer survival. The established predictors separate in the expected directions, with HR-deficiency and BRCA reaching significance.
Actual output from the live demo, rendered by the tool on the synthetic cohort (known ground truth).
Pointed at Birkbak et al.’s own 316 patients, the system independently recomputed the paper’s analyses. Each comparable analysis appears below, beside the published value; survival models are compared on the test that governs the paper’s claims, and effect sizes agree once expressed on a common scale (see note).
| Reported quantity | DataStoryMD | Birkbak | Verdict |
|---|---|---|---|
| Whole cohort | |||
| Chemoresistance rate, low / high mutation burden | 40.0% / 25.3% | 40.2% / 23.9% | concordant |
| Mutation burden → PFS / OS (p) | .008 / .0008 | .013 / .0014 | concordant |
| BRCA-mutated subset | |||
| Burden → PFS / OS, univariate (p) | .001 / .010 | .002 / .011 | concordant |
| Burden → PFS / OS, adjusted (p) | .014 / .005 | .027 / .008 | concordant |
| BRCA wild-type subset: the paper’s null results | |||
| Burden → PFS / OS (significance) | ns / ns→sig* | ns / ns→sig* | concordant |
A median is a descriptive property of the exact variable measured. Our mutation count includes non-synonymous variants only; the paper’s also counts the synonymous fraction, so the two are different measurements and their medians describe different quantities. Comparing those raw values would not be a like-for-like reproduction, so we do not report it as one. The models are unaffected: a linear rescaling of a continuous covariate leaves the Cox partial-likelihood test invariant, which is exactly why every model and test above reproduces. Ingesting the full variant set (GDC MAF) would align the raw counts too, but changes none of the conclusions.
*On effect size: Birkbak quotes the adjusted BRCA-mutant hazard ratio per ~10 mutations (0.821), with a confidence interval stated per single mutation. Expressed per single mutation, our estimate is 0.971, inside the paper’s published interval (0.966, 0.995). Same effect, common unit.
Concordance is only half the claim; the other half is that the reproduction is more defensible than the original. On every finding, by default, the system applies:
On this cohort the system detects a superset of the standard prognostic picture: 27 findings across overall and progression-free survival and platinum sensitivity. Two of the paper’s analyses place the mutation count as an outcome (median-by-group); producing a continuous molecular variable on the left-hand side is a current engine boundary, disclosed rather than worked around. And this is a reproduction on public, de-identified data: a demonstration of capability, not a claim of clinical effect. Findings on any single retrospective cohort remain exploratory until externally validated, as the tool states on every report it writes.