Clinical data intelligence

The findings are
already in your data.

DataStoryMD reads a messy real-world clinical spreadsheet, recovers the associations that matter, adjusts them for confounders, and writes them up. It also flags the ones that are artifacts of study design rather than biology.

Real published data, no signup — or start with the synthetic cohort, where the ground truth is known.
89 published numbers checked. Reproduces a standard textbook to within its own rounding, and two published hazard ratios on the paper’s own cohort.
BRCA2 Carrier → Overall Survivalsurvival
aHR 0.48  (95% CI 0.33–0.70)  ·  q < 0.001  ·  n = 300
Cox model, adjusted for age, FIGO stage, histology, pretreatment CA-125
See how it works
From the spreadsheet you actually have

Real-world data is a mess. That is the starting point, not the excuse.

Coded columns, value legends jammed into headers, sentinels, mixed languages. DataStoryMD maps every column to clinical meaning before it models anything.

Data Table · 300 rows · 133 columns 126 mapped7 unmapped
Age Ethnicity Family History Diabetes Debulking
67Ashkenazi JewishNoNo DiabetesR0 (complete)
60Non-AshkenaziYesNo DiabetesR2
41Ashkenazi JewishNoDiabetesR0 (complete)
72UnknownNoNo DiabetesR1
55Ashkenazi JewishYesDiabetesNo surgery
BRCA1 carrier → Platinum sensitivityendpoint
OR 6.2 (95% CI 2.4–16.1) · q < 0.001 · n = 300

Named, correctly typed, modeled, and written up in a sentence a clinician can read aloud.

Beyond the findings

Ask it questions. Export the answers. Bring your own specialty.

The parts a findings list can't give you.

Ask in plain English

Type a clinical question and get a real answer with the statistics behind it. No query language, no pivot tables.

"Does complete debulking help platinum-resistant patients?"

Export for grants and talks

One click to a formatted Word or PowerPoint document, figures, methods, and citations included.

It already knows your field

Each specialty pack encodes the causal structure, temporal order, confounders, valid endpoints, and the traps. The statistical judgment is built in, so your clinical expertise is all it takes to run rigorous research.

Findings, organized

Not a flat list. Findings grouped into themes, ranked by impact.

Dozens of related results collapse into a handful of clinical themes, each scored and ready to open.

154comparisons tested 154eligible to model 65significant, FDR-corrected 88stories · 89 not significant

Counts from the bundled 300-patient synthetic cohort. TCGA-OV, with 28 mapped clinical fields, tests 19.

Predictors of Overall Survival

6 converging stories
6 exploratory
Overall Survival differs by BRCA2 Carrier
BRCA2 Carrier
IMPACT97
BRCA / HRD group: 36.4-month gap in OS
BRCA / HRD Group
IMPACT60
Overall Survival differs by HRD Score
HRD Score
IMPACT40

Predictors of Platinum Sensitivity

4 converging stories
4 exploratory
BRCA1 Carrier: higher Platinum Sensitivity
BRCA1 Carrier · OR 6.2
IMPACT96
LOF Variant: higher Platinum Sensitivity
Loss-of-Function · OR 4.1
IMPACT96
Founder Mutation: higher sensitivity
Founder Mutation · OR 3.9
IMPACT44

Then keep going: the pathway map, cohort comparison, and live subgroup drill-down are built to be explored, not screenshotted. See them in the live demo →

Rigor that earns trust

It separates real effects from artifacts of study design.

Any tool can surface a correlation. The harder question is which associations reflect biology and which are artifacts of how the data was collected.

On this cohort DataStoryMD set aside 77 outcome candidates as ineligible and excluded hundreds of biased comparisons, each with a stated reason.

Excludednumber of treatment lines
Immortal-time bias

More treatment lines means the patient lived long enough to receive them. Survival drives the count, not the reverse.

No black box

Every number is one you could defend.

Adjusted models, honest reporting, and your data staying with you are the defaults, not add-ons. (This demo runs on a synthetic cohort with planted effects, so you can watch the tool recover exactly what was built in.)

300
synthetic patients, 133 raw columns
65
findings surfaced in one run
77
ineligible outcomes correctly excluded
Checked against numbers we did not produce

A synthetic cohort proves it finds what was planted. Published research proves it is right.

So we ran it on cohorts whose results are already in the literature and compared every number we could check — against a standard textbook, a peer-reviewed paper, and the consortium that assembled the data.

89
published numbers checked
100%
textbook coefficients matched within 0.1%
2 of 3
published hazard ratios reproduced to the digit; the third differs by estimand, not direction
Ovarian cancer · TCGA, 316 patients Published DataStoryMD
Age at diagnosis → progression-free survival 0.995 (0.982–1.009) 0.995 (0.982–1.008) reproduced
Age at diagnosis → overall survival 1.019 (1.005–1.033) 1.017 (1.004–1.031) reproduced
Tumour stage → overall survival 1.325 (0.960–1.828) 2.041 (0.851–4.899) estimand differs
Hazard ratios from Birkbak et al., PLoS One 2013, reproduced on the same 316 patients that paper analysed. Re-run on the 316 patients Birkbak analysed, rebuilt by the same script that builds the bundled cohort. The two age terms reproduce to the digit — age → OS 1.017 (1.004–1.031) against their 1.019 (1.005–1.033), and age → DFS 0.995 against their 0.995. The stage row is a difference of ESTIMAND rather than of result: their 1.325 sits inside our interval and both analyses agree the effect is not significant, but our surface summarises eight FIGO levels with a k-sample log-rank where the paper fitted stage as an ordinal Cox term, so the point estimates are not the same quantity and we do not badge it as a match. The full comparison, including one we do not match and why, is in the validation report.

Sources: Hosmer & Lemeshow, Applied Logistic Regression 2nd ed. · Birkbak et al., PLoS One 2013;8(11):e80023 · TCGA Research Network, Nature 2011;474:609–615. Based upon data generated by the TCGA Research Network.

Get started

Three ways in, depending on where you are.

Run it on your own cohort

Private beta. Your data never leaves your machine — the tool runs locally and the analysis is offline.

Request beta access
or write to hello@datastorymd.com
Already have access

Install, point it at a spreadsheet, and read the findings. Ten minutes end to end.

Open the quickstart
Just looking

No signup. TCGA-OV is real published data you can check against the literature; the synthetic cohort has known ground truth. Either runs in about fifteen seconds.

Open the TCGA-OV cohort Open the synthetic cohort

Best on a laptop or desktop. The analysis workspace uses wide tables, survival curves and a pathway map, so on a phone this page reads fine, but the tool itself wants a bigger screen.

DataStoryMD · clinical narrative intelligence · demo on synthetic data, no PHI