DataStoryMD reads a messy real-world clinical spreadsheet, recovers the associations that matter, adjusts them for confounders, and writes them up. It also flags the ones that are artifacts of study design rather than biology.
Coded columns, value legends jammed into headers, sentinels, mixed languages. DataStoryMD maps every column to clinical meaning before it models anything.
Named, correctly typed, modeled, and written up in a sentence a clinician can read aloud.
The parts a findings list can't give you.
Type a clinical question and get a real answer with the statistics behind it. No query language, no pivot tables.
One click to a formatted Word or PowerPoint document, figures, methods, and citations included.
Each specialty pack encodes the causal structure, temporal order, confounders, valid endpoints, and the traps. The statistical judgment is built in, so your clinical expertise is all it takes to run rigorous research.
Dozens of related results collapse into a handful of clinical themes, each scored and ready to open.
Counts from the bundled 300-patient synthetic cohort. TCGA-OV, with 28 mapped clinical fields, tests 19.
Then keep going: the pathway map, cohort comparison, and live subgroup drill-down are built to be explored, not screenshotted. See them in the live demo →
Any tool can surface a correlation. The harder question is which associations reflect biology and which are artifacts of how the data was collected.
On this cohort DataStoryMD set aside 77 outcome candidates as ineligible and excluded hundreds of biased comparisons, each with a stated reason.
More treatment lines means the patient lived long enough to receive them. Survival drives the count, not the reverse.
Adjusted models, honest reporting, and your data staying with you are the defaults, not add-ons. (This demo runs on a synthetic cohort with planted effects, so you can watch the tool recover exactly what was built in.)
So we ran it on cohorts whose results are already in the literature and compared every number we could check — against a standard textbook, a peer-reviewed paper, and the consortium that assembled the data.
Sources: Hosmer & Lemeshow, Applied Logistic Regression 2nd ed. · Birkbak et al., PLoS One 2013;8(11):e80023 · TCGA Research Network, Nature 2011;474:609–615. Based upon data generated by the TCGA Research Network.
Open the live analysis of the synthetic ovarian-cancer cohort: the findings, the survival curves, and every association it set aside.
Best on a laptop or desktop. The analysis workspace uses wide tables, survival curves and a pathway map, so on a phone this page reads fine, but the tool itself wants a bigger screen.