Clinical research intelligence

Research at the
speed of thought.

Bring the cohort you already have. See what matters. Follow it immediately. DataStoryMD maps the data, brings your field’s knowledge to the first view, and surfaces patterns before you have to formulate the perfect question. As the investigation moves, the definitions, the methods and the safeguards move with it.

A clinician-native research workspace for real-world clinical data.

One click opens a full cohort: full-feature simulated, where the ground truth is known, or published TCGA-OV data. Your own cohort runs on your own machine.
Two published studies, re-run on the exact patients each analysed. Every quantity we could compare is reported beside the published one, including the one that differs, and why. →
Predictors of overall survivalforest plot
Multivariable Cox forest plot: hazard ratios for five predictors of overall survival
HRD-positive aHR 0.48  (95% CI 0.33–0.70)  ·  p < 0.001  ·  BRCA1/2 p = 0.029
Multivariable Cox proportional hazards · adjusted for stage, debulking, age, histology, CA-125 · n = 181 of 300
See how it works
From the spreadsheet you actually have

Real-world data is a mess. That is the starting point, not the excuse.

Coded columns, value legends jammed into headers, sentinels, mixed languages. DataStoryMD maps the columns to clinical meaning before it models anything, and names the ones it could not: on the demo cohort that is 132 of 136.

0=negative BRAC1=1 BRCA2=2 UK=3→BRCA status HRD: 0=Negative, 1=Positive, UK=3→HRD status platinum: Resist=0, Sens=1, UK=2→Platinum sensitivity ca125 after 6 cycles→CA-125 · cycle 6 ECOG PS (Performance Status)→ECOG status
Data Table · 300 rows · 136 of 149 columns 132 mapped4 unmapped
✓Age ✓Ethnicity ✓Family History ✓Diabetes ✓Debulking
67Ashkenazi JewishNoNo DiabetesR0 (complete)
60Non-AshkenaziYesNo DiabetesR2
41Ashkenazi JewishNoDiabetesR0 (complete)
72UnknownNoNo DiabetesR1
55Ashkenazi JewishYesDiabetesNo surgery
BRCA1 carrier → Platinum sensitivityendpoint
OR 6.2 (95% CI 2.4–16.1) · q = 0.001 · n = 300

Named, correctly typed, modeled, and written up in a sentence a clinician can read aloud.

The bottleneck is every turn

The hour should be spent thinking scientifically.

Having an hour free is not the same as having an hour to research. Most of it goes into operating the machinery: preparing the data, specifying variables, choosing a method, briefing a statistician, waiting, then reconstructing where you were. Every new question pays that cost again, and the questions that never feel worth the cost are the ones that quietly go unasked.

Natural-language tools make each request easier to execute, and that is a real gain. It is not the same as making the next question cheap to ask. You still have to formulate it, hold the context yourself, steer the system and check that it understood you. DataStoryMD is built so the next scientifically valid question costs almost nothing to follow, and is held to the same standard as the one you planned.

What you have to bring

Open the file. Start thinking.

Most research tools begin after the hard part: once someone has framed the question, chosen the outcome, cleaned the columns and picked the test. This one begins at the file, and lets the thinking go wherever it goes.

From the first minute

  • →Arrive with curiosity, not a hypothesis. It reads what is in the cohort and hands you the questions back, ranked by what matters clinically.
  • →Every variable is already a candidate. It works out what can serve as an outcome, what belongs on the other side, and which pairings are worth a model.
  • →The method arrives with the question. The test, the model and the adjustment set are chosen for what is being asked, and shown to you.

And underneath, all along

  • ✓Time order checked before a model is fit. A predictor measured after its outcome is refused outright. That is an eligibility rule, not a claim that the remaining effect is causal.
  • ✓Adjusted where the question and the data support it. Effects come from a multivariable model carrying the confounders your specialty pre-specifies; where a cohort cannot support one, the card says so instead of implying otherwise.
  • ✓Read against the literature. 164 curated associations declared in the ovarian pack (154 carrying a literature reference), and a check that our number and theirs are even the same quantity before either is called a match.
  • ✓Corrected for multiplicity. Across every pair tested, not only the ones that survived.
  • ✓Sorted for your eye. Grouped into clinical themes and ranked by impact, so pattern recognition, the thing you are already expert at, has something it can work on.

Starting this easily usually costs you depth. That is the trade every simple tool makes. Here you get both.

Findings, organized

What matters, in the shape a clinician reads.

Dozens of related results collapse into a handful of clinical themes, each scored and ready to open. Overview first, detail when you ask for it, so the pattern recognition you already have gets something to act on, instead of a table to grind through.

362comparisons tested → 362eligible to model → 66significant, FDR-corrected → 67stories · 295 not surfaced

Counts from the bundled 300-patient simulated cohort. TCGA-OV, with 28 mapped clinical fields, tests 19.

Predictors of Overall Survival

6 converging stories
6 exploratory
Overall Survival differs by BRCA2 Carrier
BRCA2 Carrier
IMPACT97
BRCA / HRD group: 36.4-month gap in OS
BRCA / HRD Group
IMPACT60
Overall Survival differs by HRD Score
HRD Score
IMPACT40

Predictors of Platinum Sensitivity

4 converging stories
4 exploratory
BRCA1 Carrier: higher Platinum Sensitivity
BRCA1 Carrier · OR 6.2
IMPACT96
LOF Variant: higher Platinum Sensitivity
Loss-of-Function · OR 4.1
IMPACT96
Founder Mutation: higher sensitivity
Founder Mutation · OR 3.9
IMPACT44

Then keep going: the pathway map, cohort comparison, and live subgroup drill-down are built to be explored, not screenshotted. See them in the live demo →

Why it is built this way

Friction is not a matter of taste. It is measurable, and it costs you findings.

The interaction model is not a design preference. It follows established results on how people actually investigate data, including two controlled experiments that bear directly on what this product is for.

Half a second

Adding 500 ms of latency to exploratory analysis measurably reduced how much of the data people covered, how many observations and generalisations they made, and how many hypotheses they formed. Delay does not just annoy. It suppresses discovery.

Liu & Heer, InfoVis 2014
Speed alone fails

A visual interactive tool for clinical research produced hypotheses faster, and rated lower on feasibility and quality. Removing friction without keeping the methodology honest buys volume at the cost of worth.

Jiang et al., 2023

Those two experiments are the argument for this product in one line: speed changes how much you find, and guardrails decide whether it was worth finding. What these results establish is the design; whether this product improves research outcomes is a question we intend to measure, not one we claim answered.

Who does what

The friction moves out of the way. The research stays yours.

Everything that has to happen for a cohort to become a finding, and which side of the screen it happens on. The machinery takes the preparation, the scanning and the bookkeeping. The judgement never leaves your hands.

Stage
DataStoryMD prepares and checks
You think and explore
01
Bring the cohort

Opens the file you already have. Profiles every field, maps local names and codes, resolves units, missingness and repeated measures into clinical concepts, and says plainly what it could not resolve.

Confirm anything consequential. You start from your dataset, not from an analysis-ready statistical specification you had to write about it first.

02
See the landscape

Builds the overview (variables, outcomes, timelines and cohort structure) before any single model takes over the screen.

Notice what is interesting. Patterns, gaps and distributions, in the context that makes them mean something.

03
Discover

Scans broadly, behind gates: associations, subgroups, survival, thresholds, interactions, trajectories, and only where the design permits the question.

Decide what deserves attention. Whether a result is clinically plausible, and whether it matters, stays a human judgement.

04
Explore

Recomputes without rebuilding. The filter, the focal variable, the outcome and the adjustment set all stay attached to the same analysis.

Test the idea directly. Split it, compare, drill down, change what it adjusts for, challenge the explanation you were given.

05
Follow the thread

Keeps the context. A finding stays linked to its methods, assumptions and evidence instead of becoming a screenshot in a folder.

Turn one observation into the next question, while everything you built to get there is still standing.

06
Draft

Assembles methods, results and figures from the same analytical record you just explored, so the write-up matches what was actually run.

Interpret, edit, conclude. It accelerates the writing. The science, and the authorship, remain yours.

The left column makes statistical decisions and states every one of them: which model, which adjustment set, what it refused and why. What it does not do is decide whether a result is clinically plausible, or what it means, and the second kind is the reason you were the one looking at this cohort in the first place.

Recognition, not continuous prompting

You are not waiting for an analysis. You are inside one.

Every finding opens. See a pattern, open it, split it, compare it, change the population, follow a pathway, challenge it, inspect the evidence. This is the part you do yourself, and you do it by acting on what you notice rather than by composing an instruction for each move. No request, no queue, no one to ask.

Every control above is in the live demo, and every number is one the engine produced on the bundled 300-patient cohort. The same adjustment sets, multiplicity correction and time-order rules apply whichever way you got here, so a question you followed on impulse is held to the same standard as the one you planned, and is labelled exploratory when it is.

Beyond the findings

Ask when it is faster. Export what you need. The field knowledge is already there.

The parts a findings list can't give you.

Ask in plain English when it is faster

Most of a session is recognising something and following it. Sometimes a question is quicker said than found, so type it and get a real answer with the statistics behind it. No query language, no pivot tables.

"Does complete debulking help platinum-resistant patients?"

Export for grants and talks

One click to a formatted Word or PowerPoint document, figures, methods, and citations included.

Domain knowledge before the first click

Governed Domain Packs encode how a clinical field is structured: what the variables mean, what matters, how time is ordered, which relationships deserve attention and which analytical safeguards apply. The knowledge changes what the product does, not merely the terminology it displays.

164 curated associations declared, 154 carrying a literature reference · BRCA → platinum sensitivity (Alsop 2012) · BRCA → survival (Bolton 2012)
Rigor that earns trust

It separates real effects from artifacts of study design.

Any tool can surface a correlation. The harder question is which associations reflect biology and which are artifacts of how the data was collected.

On this cohort it set aside 67 outcome candidates as ineligible and excluded hundreds of biased comparisons, each with a stated reason.

Deterministic by design. The same data and the same rules produce the same analysis, from a fixed statistical engine rather than a language model.

Every result is traceable: the estimate, its uncertainty, what it was adjusted for, and the decisions behind it.

Excludednumber of treatment lines
Immortal-time bias

More treatment lines means the patient lived long enough to receive them. Survival drives the count, not the reverse.

One engine, many specialties

Everything above is ovarian cancer. None of it is built into the engine.

The vocabulary, the chronology, the endpoints, the confounders and the literature all live in a governed Domain Pack the engine reads, so a specialty is something it can be taught. This is not generic AI with medical vocabulary: the pack changes the research behavior of the system. Depth differs by pack, and we say which is which.

Oncology

Ovarian cancervalidated

Survival, HRD and BRCA, platinum sensitivity, debulking, and its estimates are reported beside the published TCGA-OV literature.

Open the live demo →
Obstetrics · maternal-fetal

Pregnancyin validation

Trimester uterine- and umbilical-artery Doppler, sFlt-1/PlGF, pre-eclampsia and growth-restriction screening. The engine reads real-world windowed columns automatically. In active clinical validation before it goes public.

Ask for early access
Your domain next

Breast, cervical, endometrial…

Name your field’s concepts, endpoints and timelines once. The same engine then handles the messy data, the adjusted models and the caveats, on your specialty.

Ask about your domain
Checked against numbers we did not produce

A simulated cohort proves it finds what was planted. Published research proves it is right.

So we ran it on public cohorts whose results are already in the literature, and checked every number we could, against the peer-reviewed paper and the consortium that analysed those exact patients.

2
published studies compared, on the exact patients each one analysed
3
quantities in the table below, each shown beside its published value
Ovarian cancer · TCGA, 316 patients Published DataStoryMD
Age at diagnosis → disease-free survival 0.995 (0.982–1.009) 0.995 (0.982–1.008) endpoint unverified
Age at diagnosis → overall survival 1.019 (1.005–1.033) 1.017 (1.004–1.031) compared
Tumour stage → overall survival 1.325 (0.960–1.828) 2.041 (0.851–4.899) estimand differs
Hazard ratios from Birkbak et al., PLoS One 2013, compared against our own estimates on the same 316 patients that paper analysed. Re-run on the 316 patients Birkbak analysed, rebuilt by the same script that builds the bundled cohort. Our estimate on the first row is computed from the cohort file’s disease-free survival columns, which are the only time-to-event fields it carries besides overall survival. Disease-free and progression-free survival are different endpoints, so until we have verified which one the paper reported, that row is shown without a verdict rather than badged as a match. The stage row is a difference of ESTIMAND rather than of result: their 1.325 sits inside our interval and both analyses agree the effect is not significant, but our surface summarises eight FIGO levels with a k-sample log-rank where the paper fitted stage as an ordinal Cox term, so the point estimates are not the same quantity and we do not badge it as a match. The full comparison, including one we do not match and why, is in the validation report.

Read the full validation report →

What you open is what we checked. The demo carries this exact 316-patient cohort, the same patients Birkbak et al. analysed, and the same file these hazard ratios were computed on. Like for like, no larger-export caveat.

Sources: Hosmer & Lemeshow, Applied Logistic Regression 2nd ed. · Birkbak et al., PLoS One 2013;8(11):e80023 · TCGA Research Network, Nature 2011;474:609–615. Based upon data generated by the TCGA Research Network.

Three ways in

Start with the demo, or go straight to your own data.

The live demo

No signup, no upload. Three cohorts are waiting inside: a simulated ovarian cohort with every field populated and the ground truth known, the real TCGA-OV patients the hazard ratios above were computed on, and a simulated obstetric cohort. Pick one and it runs to a full report.

Open the live demo →
Run it on your own cohort

Private beta. Your data never leaves your machine: the tool runs locally and the analysis is offline.

Request beta access →
Questions, collaborations and press go through the same form.
I have beta access

Install, point it at a spreadsheet, and read the findings. Ten minutes end to end.

Open the quickstart →

Clinical intuition, amplified.

Best on a laptop or desktop. The analysis workspace uses wide tables, survival curves and a pathway map, so on a phone this page reads fine, but the tool itself wants a bigger screen.

DataStoryMD · clinical research intelligence · demo on simulated and public de-identified cohorts, no PHI