DrugEvidence

Methodology

How a page is built, which trials it counts, and what a human decides

What a page is

Every page on this site has an identity of its own — a stable identifier and URL that does not change when an upstream vocabulary changes. A page is not a MeSH term, a sponsor record or a registry entry; it is a subject we chose to cover, mapped to the upstream identifiers that describe it.

That mapping is a table, not code. When ClinicalTrials.gov renames a term, splits one into two, or labels a trial incorrectly, we edit one row and recompute. The page, its URL and its links stay where they are.

Which trials are counted

Drug pages collect trials whose intervention was tagged with the molecule's MeSH term by the National Library of Medicine, restricted to trials that actually administer a drug, biological or combination product. A trial that merely mentions a molecule — a lifestyle study, a survey — is not counted.

Disease pages collect trials tagged with the condition's MeSH term. Subtypes are included only when we listed them explicitly; a page never silently absorbs its children, because a number nobody can reproduce from the term is not worth publishing.

Company pages collect trials whose lead sponsor matches one of the spellings we mapped to that company, including subsidiaries and former names. Sponsor names are not normalized upstream, so this list is maintained by hand.

Individual trials can be force-included or force-excluded per page. That is how we correct an upstream mislabel without touching anything else.

How the numbers are computed

Once a day, a job resolves each page's trial set against the registry and recomputes every statistic on the page. Each result is stored as a dated snapshot; pages read the newest one and print its date in the footer.

Ratios are stored as their two parts and divided only at rendering time. You will always see the numerator and the denominator next to a percentage, because a rate without its base cannot be checked.

If a page's trial count moves by more than a set threshold from one day to the next — or drops to zero — that day's snapshot is flagged and the page keeps showing the previous one until a human has looked. This is deliberate: the most likely cause of a sudden change is our own mapping, not the science.

Termination reasons are grouped from the sponsor's own free-text wording by a rule-based classifier. The grouping is approximate; the original wording is always shown next to it.

What a human decides

Automatic matching never publishes itself. Candidate mappings, newly seen upstream terms and snapshot anomalies land in a review queue; nothing reaches a page until someone confirms it. Profile text is reviewed before it renders — unreviewed text is stored but never shown.

Pages generated in bulk from upstream terms exist unpublished until they are reviewed. If you can read a page, its mapping was confirmed by a person.

Known limits

Registry data reflects what sponsors filed, not what happened. Trials can be mis-tagged, updated late, or never updated at all. Results posted on the registry are not peer-reviewed. Counts on this site therefore measure the registry, and only indirectly the research.

We do not mirror study records, publish per-trial pages, or reproduce third-party article text. Every trial and every headline links out to its source.