Automated data extraction for systematic reviews

Screening's done.
Now it's just you and forty PDFs.

Every mean, SD, and N, located and typed by hand. The SD that turns out to be an SEM. The subgroup that lives only in a figure. The one transposed digit that reaches print with your name on it. TrialExtract does that pass in minutes, and hands you a receipt for every number: the sentence it read, the formula it ran, the check that ran. You review instead of retype.

Founding cohort access. A free preview of any paper you bring, before you spend a credit.

Free preview on every paper before you spend a credit · exports to metafor & RevMan

The table you've been dreading. Already filled.

Every outcome from the PDFs, one row apiece: the native effect, the standardized SMD, and where each was read.

StudyOutcomeNativeSMD d [95% CI]Status
Stough 2008Mental control (Wechsler Memory Scale)Mean difference 0.60 score1.14 [0.43, 1.86] derived
Calabrese 2008Heart rate (bpm)Mean difference −2.60 beats/min−0.27 [−0.84, 0.29] derived
Calabrese 2008Total adverse events (all reported AEs)Mean difference −5.00 N reported
Stough 2008Percentage improvement category: > 21% (0–12 weeks, total score)Responder proportion 0.56 % reported
  • Stough 2008 Mental control (Wechsler Memory Scale)
    1.14 [0.43, 1.86] derived
  • Calabrese 2008 Heart rate (bpm)
    −0.27 [−0.84, 0.29] derived
  • Calabrese 2008 Total adverse events (all reported AEs)
    reported
  • Stough 2008 Percentage improvement category: > 21% (0–12 weeks, total score)
    reported
Bacopa memory review · four outcomes, two trials · every value has a receipt

The receipt behind every number

The numbers are the machine's job. The judgment's yours.

Open any value and it answers for itself. Every cell in that table traces straight back to the paper: the exact sentence it was read from, the formula that computed it, and whether a second reading agreed.

Click a tab for the three shapes a number arrives in: reported and verified, recovered from the arm means, or recovered from a confidence interval. This is what you'd be signing off on.

Receipt Calabrese 2008 · AVLT delayed recall · 12 wk

Only per-arm means reported. The effect is computed, every step shown.

Calabrese 2008 · AVLT delayed recall · 12 wk
MD 0.7 words
SMD g = 0.170 · 95% CI [−0.397, 0.737] lossy derived
Table 1: AVLT Delayed Recall, 12 wk
Source evidence
Table 1, AVLT Delayed Recall (# of words): Bacopa 7.6 (3.9) vs Placebo 6.9 (4.2) at +12 Weeks (F=5.4; 1,21; p=0.03*). Narrative: “Controlling for baseline cognitive deficit using the Blessed Orientation–Memory–Concentration test, Bacopa participants had enhanced AVLT delayed word recall memory scores relative to placebo.”
7.6 Bacopa mean 3.9 Bacopa SD 6.9 placebo mean 4.2 placebo SD

Not located in the quote (not restated here): 24 Bacopa n,24 placebo n

Table 1 (Means for Cognitive, Affective, and Physiologic Measures), AVLT Delayed Recall row, +12 Weeks column Table 1 (Means for Cognitive, Affective, and Physiologic Measures)Article narrative (Europe PMC full text) Highlights locate each value in the quoted text; locating a number is not the same as checking it.
How this value was derived
g = 0.170 95% CI [−0.397, 0.737] derived · from arm means
  1. 1 Reported arm summaries Bacopa: n=24, mean=7.6, SD=3.9 · placebo: n=24, mean=6.9, SD=4.2 N=48
  2. 2 Pooled SD (Cochrane 5.3.a) √[((24−1)·3.9² + (24−1)·4.2²) / (24+24−2)] s_pooled = 4.0528
  3. 3 Mean difference 7.6 − 6.9 Δ = 0.7
  4. 4 Cohen's d Δ / s_pooled = 0.7 / 4.0528 d = 0.1727
  5. 5 Hedges small-sample correction (Hedges & Olkin 1985) g = d · J, J = 1 − 3/(4·48−9) = 0.98361 g = 0.1699
  6. 6 Variance & 95% CI (Hedges & Olkin 1985) Var(g) = (N)/(n₁·n₂) + g²/(2N) = 0.0836; CI = g ± 1.96·√Var [−0.397, 0.737]

Independently re-derived: the point and 95% CI both match the reported values.

Cochrane summary-data SMD (Handbook v6.5 §6.5.1); pooled SD per Cochrane 5.3.a; Hedges & Olkin (1985) small-sample correction
Independent verification
not run

Independent verification has not run for this value. Provenance located the number in the source; that is not the same as a second reader agreeing with it.

A real extraction from a real published trial.

The extraction phase

Screening ends. The grind begins.

Extraction by hand

  • Two to three hours per trial: finding the right table, the right arm, the right timepoint
  • Usually alone. If you can staff a second extractor, a reconciliation meeting on top of it
  • SEM or SD? Median and IQR to convert? Endpoint or change score? Decided at 11pm, cell by cell
  • Six weeks later, a grid you still have to spot-check against forty PDFs

With TrialExtract

  • Upload your library. The first pass is done in minutes, not weeks
  • Every value arrives with the exact sentence it came from, highlighted where it sits in the paper
  • SEM to SD, median and IQR to mean and SD, effect sizes: converted with cited formulas, every step shown
  • Second-extractor rigor without a second extractor. You review what the machine did, and decide

How it works

You bring the library. It brings back the month.

  1. 1

    Import your studies

    RIS, BibTeX, .nbib, CSV, or drop the PDFs. Open-access full texts are fetched for you.

  2. 2

    Preview, free

    It reads each paper and shows what's extractable: the tables, figures, and outcomes it found. No credit spent.

  3. 3

    Run the extraction

    Priced per paper, quoted up front. The quote is exactly what you're charged.

  4. 4

    Review by outcome

    Every value lands in a grid, organized by outcome. Open any one for its receipt, then accept, edit, or reject it.

  5. 5

    Pool and export

    DerSimonian–Laird random-effects pooling, then metafor, RevMan, or a workbook, each with the full audit trail attached.

Free on every paper

Bring your ugliest PDF. It reads the whole thing, free.

The scanned table. The SEM where you wanted an SD. The subgroup buried in a figure. Give it the paper you're sure will break it. It reads the whole thing: a plain-language synopsis, and every table, figure, and supplementary block it found, tagged by how it was read. Reading is free. You spend a credit only after you've watched it work.

Preview Read · free Peth-Nui 2012 · Bacopa 300/600 mg vs placebo
Generated source description

A 12-week, double-blind, three-arm randomized controlled trial comparing standardized Bacopa monnieri extract (300 mg/day and 600 mg/day) with placebo in 60 healthy older adults. It reports memory, sustained-attention, and reaction-time outcomes with per-arm means and standard deviations, alongside adverse-event counts, primarily in Tables 2–4 and Figure 1.

14
pages
38,204
text chars
2
tables
1
figure
1
image-table

Readable structure we found

source_blocks · by shape
4 narrative1 table1 image-table1 figure1 supplementary 2 read by vision
narrative
Abstract
Abstract Text layer
1,240ch
narrative
Methods · design & randomization
Methods Text layer
4,980ch
narrative
Results · adverse events
Safety narrative Text layer
2,110ch
narrative
Discussion
Text layer
3,120ch
table
Table 2 · cognitive outcomes
Results table Text-layer table
2,980ch
image-table vision
Table 3 · reaction time
Results table Transcribed by vision
1,610ch
figure vision
Figure 1 · AVLT trajectory
Figure Transcribed by vision
720ch
supplementary
Protocol · randomization
Supplementary Supplementary document
3,400ch
The free read of a real published trial: the machine's synopsis, and every readable table, figure, and supplementary block it found in the source.

The deliverable, in full

The whole table, with receipts.

Six weeks of typing, done. One row per outcome, per arm, per timepoint: the effect as the paper reported it, the standardized SMD with its confidence interval, and where each value was read. Derived values are marked as derived. A value the paper doesn't report enough to compute is shown as reported, never invented.

StudyRunOutcomeContrastTimepointFamilyNativeSMD d [95% CI]StatusVerifyDecision
Stough 2008Table 1 — Mental control (Wechsler Memory Scale)SBME vs Placebo12 weeks (TREATMENT)ContinuousMean difference 0.60 score1.14 [0.43, 1.86] derived
Calabrese 2008Table 1 — Heart rate (bpm)Bacopa vs placebo6 weeks (TREATMENT)ContinuousMean difference −2.60 beats/min−0.27 [−0.84, 0.29] derived
Calabrese 2008Article narrative — Total adverse events (all reported AEs)Bacopa vs placebo12 weeks (TREATMENT)ContinuousMean difference −5.00 N reported
Stough 2008Table 1 — Percentage improvement category: > 21% (0–12 weeks, total score)SBME vs Placebo12 weeks (TREATMENT)ResponderResponder proportion 0.56 % reported
Table 1 — Mental control (Wechsler Memory Scale) SBME vs Placebo
1.14 [0.43, 1.86]
derived
Stough 2008 12 weeks (TREATMENT) Continuous
Table 1 — Heart rate (bpm) Bacopa vs placebo
−0.27 [−0.84, 0.29]
derived
Calabrese 2008 6 weeks (TREATMENT) Continuous
Article narrative — Total adverse events (all reported AEs) Bacopa vs placebo
Mean difference −5.00 N
reported
Calabrese 2008 12 weeks (TREATMENT) Continuous
Table 1 — Percentage improvement category: > 21% (0–12 weeks, total score) SBME vs Placebo
Responder proportion 0.56 %
reported
Stough 2008 12 weeks (TREATMENT) Responder
Four outcomes from two randomized bacopa trials, exactly as the machine read them, before you've reviewed a single row. Every value carries a receipt of its own.

What you walk away with

Pooled, plotted, and ready to defend.

The forest plot you'd drop straight into the manuscript: two randomized bacopa memory trials, pooled into one random-effects estimate. And when the studies genuinely disagree, it shows you, instead of hiding it behind false precision.

Memory (Bacopa monnieri vs placebo)

k=2 · N=83 · DL random-effects · SMD (Hedges g) · I²=77% considerable · provisional

Stough2008
N=35 1.14 [0.43, 1.85]
Stough 2008: 1.14 [0.43, 1.85], significant: CI excludes the null, weight from √n fallback, weight 47.4% of pool, effect derived.
Calabrese2008
N=48 0.17 [−0.40, 0.74]
Calabrese 2008: 0.17 [−0.40, 0.74], not significant: CI crosses the null, weight 52.6% of pool, effect derived.
Pooled (DL)
N=83 0.63 [−0.32, 1.58]
Pooled DL random-effects estimate 0.63 [−0.32, 1.58], inconclusive: CI crosses the null, I²=77% considerable heterogeneity, provisional until accepted. No prediction interval (k<3).
Pooled (DL)Pinned summary
0.63 [−0.32, 1.58] k=2 · N=83 · I²=77% considerable
Pooled DL random-effects estimate 0.63 [−0.32, 1.58], inconclusive: CI crosses the null, I²=77% considerable heterogeneity, provisional until accepted. No prediction interval (k<3).
Stough2008
1.14 [0.43, 1.85] N=35 · wt 47.4%
Stough 2008: 1.14 [0.43, 1.85], significant: CI excludes the null, weight from √n fallback, weight 47.4% of pool, effect derived.
Calabrese2008
0.17 [−0.40, 0.74] N=48 · wt 52.6%
Calabrese 2008: 0.17 [−0.40, 0.74], not significant: CI crosses the null, weight 52.6% of pool, effect derived.
Pooled (DL)
0.63 [−0.32, 1.58] k=2 · N=83 · I²=77% considerable
Pooled DL random-effects estimate 0.63 [−0.32, 1.58], inconclusive: CI crosses the null, I²=77% considerable heterogeneity, provisional until accepted. No prediction interval (k<3).
CI excludes null inconclusive weight from √n derived effect pooled · hue = heterogeneity
Calabrese 2008 and Stough 2008, pooled with DerSimonian–Laird.

Defensibility

Built for the day Reviewer 2 asks where a number came from.

Every value here answers for itself: the sentence it was read from, the formula that computed it, and its verification status. When a paper doesn't report a number cleanly, the derivation is shown, not hidden. Open any number and check it yourself, even two years on at revision, long after the details have left your head.

Suspect numbers stay out of your pool

Quality checks run on every extracted value. One that fails them is quarantined: surfaced with the reason, and excluded from every pooled estimate until you rule on it. Your forest plot never quietly includes a value you haven't seen flagged.

Your decisions, on the record

Accepts, edits, and rejections are stamped into the export. When a co-author asks whether you included the per-protocol arm, the answer is in the trail, not in anyone's memory.

Nothing enters silently

No value reaches your evidence table without a source. If a paper doesn't report enough to compute an effect, you see that stated. You'll never discover a guessed number at galley proofs.

Statistical rigor

It recovers the numbers a paper leaves out. No typing, no guessing.

When a paper reports an SEM instead of an SD, or a median where you needed a mean, TrialExtract recovers the real value with the Cochrane Handbook's own formula: every step shown on the receipt, nothing estimated, nothing assumed. Hedges' g applies its small-sample correction as a visible step. Pooling names its estimator on the plot. Math your methods section can cite, and a statistical reviewer can check line by line.

Derivation · recorded on the receipt

SE = (CIupper − CIlower) / (2 · z0.975)

SD = SE / √(1/n₁ + 1/n₂)

Cochrane Handbook §6.5.2.3 · obtaining SDs from CIs

g = J · d,  J = 1 − 3/(4·df − 1)

Hedges & Olkin · small-sample correction, shown as a step

Pricing

Credits, not subscriptions.

Reviews are episodic, so your tooling shouldn't bill monthly. At two to three hours a paper, forty trials is over a month of work by hand. A credit takes one of them from PDF to a cited row you can review in minutes, each value carrying its own receipt. Buy a pack when a review starts. Nothing recurs.

Starter

$49

3 papers · $16.33 per paper

Pilot it on your own trials

Outcome matrix save 9%

$149

10 papers

A typical meta-analysis

Review pack save 19%

$399

30 papers

The full systematic review

Previews are free on every paper, so you see exactly what's extractable before a credit ever leaves your balance.

Questions reviewers ask

Before you trust it with a review.

What study designs does it handle?
Randomized controlled trials: parallel and crossover, continuous, dichotomous, and time-to-event outcomes, including the multi-arm and multi-dose shapes common in supplement and nutrition trials. RCTs are the one study type TrialExtract is built and tuned for, and that focus is why the receipts hold up.
Will a journal accept tool-assisted extraction?
What editors want from any extraction, by hand or tool-assisted, is transparency about who checked what. That's exactly what the audit trail and the methods paragraph give them: every value reviewed by a named person, on the record.
What happens to my PDFs?
They're used to run your extraction, stored for your review, and deleted on your schedule. Your library isn't training data.

For your manuscript

Ready for your methods section.

Report tool-assisted extraction the honest way. Paste this and set your reviewer count:

Outcome data were extracted using TrialExtract (Reseda Labs), which records a source quotation, derivation chain, and verification status for each value. All extracted values were reviewed by [N] reviewer(s), and pooled estimates were computed using DerSimonian–Laird random-effects models.

Your evidence table, by this afternoon.

Every number leaves with its receipt, so it holds up whenever it's questioned. Founding seats are limited and go in the order they're claimed, so the earlier you're in, the sooner you're running.

Claim your founding seat before it's taken: