Clinical data intelligence

The findings are
already in your data.

DataStoryMD reads a messy real-world clinical spreadsheet, recovers the associations that matter, adjusts them for confounders, and writes them up. It also flags the ones that are artifacts of study design rather than biology.

Live demo runs on a synthetic 300-patient cohort. No patient data.
89 published numbers checked. Reproduces a standard textbook to within its own rounding, and three published hazard ratios on the paper’s own cohort.
BRCA2 Carrier → Overall Survivalsurvival
aHR 0.48  (95% CI 0.33–0.70)  ·  q < 0.001  ·  n = 300
Cox model, adjusted for age, FIGO stage, histology, pretreatment CA-125
See how it works
From the spreadsheet you actually have

Real-world data is a mess. That is the starting point, not the excuse.

Coded columns, value legends jammed into headers, sentinels, mixed languages. DataStoryMD maps every column to clinical meaning before it models anything.

Data Table · 300 rows · 133 columns 126 mapped7 unmapped
Age Ethnicity Family History Diabetes Debulking
67Ashkenazi JewishNoNo DiabetesR0 (complete)
60Non-AshkenaziYesNo DiabetesR2
41Ashkenazi JewishNoDiabetesR0 (complete)
72UnknownNoNo DiabetesR1
55Ashkenazi JewishYesDiabetesNo surgery
BRCA1 carrier → Platinum sensitivityendpoint
OR 6.2 (95% CI 2.4–16.1) · q < 0.001 · n = 300

Named, correctly typed, modeled, and written up in a sentence a clinician can read aloud.

Beyond the findings

Ask it questions. Export the answers. Bring your own specialty.

The parts a findings list can't give you.

Ask in plain English

Type a clinical question and get a real answer with the statistics behind it. No query language, no pivot tables.

"Does complete debulking help platinum-resistant patients?"

Export for grants and talks

One click to a formatted Word or PowerPoint document, figures, methods, and citations included.

It already knows your field

Each specialty pack encodes the causal structure, temporal order, confounders, valid endpoints, and the traps. The statistical judgment is built in, so your clinical expertise is all it takes to run rigorous research.

Findings, organized

Not a flat list. Findings grouped into themes, ranked by impact.

Dozens of related results collapse into a handful of clinical themes, each scored and ready to open.

154comparisons tested 154eligible to model 65significant, FDR-corrected 88stories · 89 not significant

Counts from the bundled 300-patient synthetic cohort. TCGA-OV, with 28 mapped clinical fields, tests 19.

Predictors of Overall Survival

6 converging stories
6 exploratory
Overall Survival differs by BRCA2 Carrier
BRCA2 Carrier
IMPACT97
BRCA / HRD group: 36.4-month gap in OS
BRCA / HRD Group
IMPACT60
Overall Survival differs by HRD Score
HRD Score
IMPACT40

Predictors of Platinum Sensitivity

4 converging stories
4 exploratory
BRCA1 Carrier: higher Platinum Sensitivity
BRCA1 Carrier · OR 6.2
IMPACT96
LOF Variant: higher Platinum Sensitivity
Loss-of-Function · OR 4.1
IMPACT96
Founder Mutation: higher sensitivity
Founder Mutation · OR 3.9
IMPACT44

Then keep going: the pathway map, cohort comparison, and live subgroup drill-down are built to be explored, not screenshotted. See them in the live demo →

Rigor that earns trust

It separates real effects from artifacts of study design.

Any tool can surface a correlation. The harder question is which associations reflect biology and which are artifacts of how the data was collected.

On this cohort DataStoryMD set aside 77 outcome candidates as ineligible and excluded hundreds of biased comparisons, each with a stated reason.

Excludednumber of treatment lines
Immortal-time bias

More treatment lines means the patient lived long enough to receive them. Survival drives the count, not the reverse.

No black box

Every number is one you could defend.

Adjusted models, honest reporting, and your data staying with you are the defaults, not add-ons. (This demo runs on a synthetic cohort with planted effects, so you can watch the tool recover exactly what was built in.)

300
synthetic patients, 133 raw columns
65
findings surfaced in one run
77
ineligible outcomes correctly excluded
Checked against numbers we did not produce

A synthetic cohort proves it finds what was planted. Published research proves it is right.

So we ran it on cohorts whose results are already in the literature and compared every number we could check — against a standard textbook, a peer-reviewed paper, and the consortium that assembled the data.

89
published numbers checked
100%
textbook coefficients matched within 0.1%
3 of 3
published hazard ratios reproduced, each inside the other's interval
Ovarian cancer · TCGA, 316 patients Published DataStoryMD
Age at diagnosis → progression-free survival 0.995 (0.982–1.009) 0.995 (0.982–1.008) matches
Age at diagnosis → overall survival 1.019 (1.005–1.033) 1.017 (1.004–1.031) matches
Tumour stage → overall survival 1.325 (0.960–1.828) 1.267 (0.893–1.797) matches
Hazard ratios from Birkbak et al., PLoS One 2013, reproduced on the same 316 patients that paper analysed. Every published estimate falls inside the interval DataStoryMD produced — and the stage row reproduces its non-significance too, not just its point estimate. Three further terms match; the full comparison, including one we do not match and why, is in the validation report.

Sources: Hosmer & Lemeshow, Applied Logistic Regression 2nd ed. · Birkbak et al., PLoS One 2013;8(11):e80023 · TCGA Research Network, Nature 2011;474:609–615. Based upon data generated by the TCGA Research Network.

See it for yourself

Watch it read the spreadsheet.

Open the live analysis of the synthetic ovarian-cancer cohort: the findings, the survival curves, and every association it set aside.

Open the live demo Run it locally

Best on a laptop or desktop. The analysis workspace uses wide tables, survival curves and a pathway map, so on a phone this page reads fine, but the tool itself wants a bigger screen.

DataStoryMD · clinical narrative intelligence · demo on synthetic data, no PHI