Ehrhart · Scientific data 2021 · Database and provenance curation study · n=4,166 rare diseases (3,163 genes)

A resource to explore the discovery of rare diseases and their causative genes.

Cited 37 times in the scientific literature.

Level 5 - mechanism / opinion, no new human data

By design analogy, not clinical CEBM (database curation/resource description).

PubMed 33947870 · doi:10.1038/s41597-021-00905-y · record verified 2026-08-26

What was done

The authors assembled a curated dataset of 4,166 rare monogenic diseases linked to 3,163 causative genes, annotated with OMIM, Ensembl, and HGNC identifiers. They manually extracted publication provenance (PubMed identifiers) for the initial descriptions of the diseases and the discovery of their underlying causative genes using OMIM, PubMed, Wikipedia, whonamedit.com, and Google Scholar. The resource was published under a CC0 license as a spreadsheet, added to Wikidata, and structured as RDF using a modified DisGeNET semantic model.

What was found

The authors produced an interoperable dataset containing 4,166 monogenic rare diseases and 3,163 causative genes linked to their historical discovery publications. Using these data, they mapped the timeline of rare disease and causative gene discoveries relative to technological and methodological developments.

Why it matters

This resource creates an open, machine-readable link between monogenic rare diseases, their causal genes, and the primary literature detailing their initial discovery.

Limits

The dataset is restricted to monogenic diseases with known genetic backgrounds and relies on publications with PubMed identifiers. Provenance curation depended on existing public resources and may carry inherited omissions or historical attribution inaccuracies.

Cited by