DataStoryMD · Independent validation

An independent validation of DataStoryMD

Given only a spreadsheet of real ovarian-cancer patients, and no guidance, DataStoryMD was asked two questions: does it recover what the field already knows, and can it reproduce a specific published paper on that paper’s own cohort? It does both, and applies methodological safeguards the original analyses did not.

Cohort: the sequenced subset of TCGA-OV (n = 316), the exact patients analysed by Birkbak et al. Reproduction re-run and verified 2026-08-03; every figure below regenerates from a committed script, not a saved copy.

27
prognostic findings surfaced unaided, each confounder-adjusted
10 / 10
of the paper’s comparable analyses reproduced , its models, tests and rates
0
disagreements in any survival model, including the paper’s null results

The paper also reports two raw median mutation counts. We do not compare those, on purpose: a median is a property of the exact variable measured, and our mutation count is measured differently from the paper’s, so comparing the two values would compare two different quantities. The models and tests carry no such problem, which is why all ten reproduce.

1 · It recovers the disease’s established prognostic structure

Before a tool can be trusted on a novel question, it should recover the canonical ones. From the raw records alone, DataStoryMD identified the accepted predictors of ovarian-cancer survival, each concordant with the literature, each adjusted for the others.

Predictors of overall survival, recovered and adjusted

A forest plot the system produced from the raw cohort. Each row is a candidate predictor, shown as an adjusted hazard ratio with its 95% confidence interval on a log scale; points left of 1 indicate longer survival. The established predictors separate in the expected directions, with HR-deficiency and BRCA reaching significance.

Forest plot of predictors of overall survival. HRD Positive (p below 0.001) and BRCA (p 0.029) sit left of 1; advanced stage, complete debulking and age at diagnosis are also shown, each adjusted for the others.

Actual output from the live demo, rendered by the tool on the synthetic cohort (known ground truth).

2 · It reproduces a published study, quantity by quantity

Pointed at Birkbak et al.’s own 316 patients, the system independently recomputed the paper’s analyses. Each comparable analysis appears below, beside the published value; survival models are compared on the test that governs the paper’s claims, and effect sizes agree once expressed on a common scale (see note).

Reported quantityDataStoryMDBirkbakVerdict
Whole cohort
Chemoresistance rate, low / high mutation burden40.0% / 25.3%40.2% / 23.9%concordant
Mutation burden → PFS / OS (p).008 / .0008.013 / .0014concordant
BRCA-mutated subset
Burden → PFS / OS, univariate (p).001 / .010.002 / .011concordant
Burden → PFS / OS, adjusted (p).014 / .005.027 / .008concordant
BRCA wild-type subset: the paper’s null results
Burden → PFS / OS (significance)ns / ns→sig*ns / ns→sig*concordant

Why we compare the models, and not the raw median counts

A median is a descriptive property of the exact variable measured. Our mutation count includes non-synonymous variants only; the paper’s also counts the synonymous fraction, so the two are different measurements and their medians describe different quantities. Comparing those raw values would not be a like-for-like reproduction, so we do not report it as one. The models are unaffected: a linear rescaling of a continuous covariate leaves the Cox partial-likelihood test invariant, which is exactly why every model and test above reproduces. Ingesting the full variant set (GDC MAF) would align the raw counts too, but changes none of the conclusions.

*On effect size: Birkbak quotes the adjusted BRCA-mutant hazard ratio per ~10 mutations (0.821), with a confidence interval stated per single mutation. Expressed per single mutation, our estimate is 0.971, inside the paper’s published interval (0.966, 0.995). Same effect, common unit.

3 · It applies safeguards the source analyses omitted

Concordance is only half the claim; the other half is that the reproduction is more defensible than the original. On every finding, by default, the system applies:

Scope, and honest limits

On this cohort the system detects a superset of the standard prognostic picture: 27 findings across overall and progression-free survival and platinum sensitivity. Two of the paper’s analyses place the mutation count as an outcome (median-by-group); producing a continuous molecular variable on the left-hand side is a current engine boundary, disclosed rather than worked around. And this is a reproduction on public, de-identified data: a demonstration of capability, not a claim of clinical effect. Findings on any single retrospective cohort remain exploratory until externally validated, as the tool states on every report it writes.