Chandak · Scientific data 2023 · bioinformatics dataset and knowledge graph development · n=20 integrated resources

Building a knowledge graph to enable precision medicine.

Cited 493 times in the scientific literature.

Level 5 - mechanism / opinion, no new human data

Level 5 by design analogy; computational dataset integration and ontology development without human clinical trial data.

PubMed 36732524 · doi:10.1038/s41597-023-01960-3 · record verified 2026-08-26

What was done

Constructed PrimeKG, a multimodal biomedical knowledge graph for precision medicine, by integrating 20 primary databases and ontologies. The resource maps genotype-to-phenotype connections across 10 biological scales—including disease-associated protein perturbations, pathways, anatomical features, phenotypes, approved drugs, and textual clinical guideline descriptions—and provides a framework for continual updates.

What was found

The resulting knowledge graph integrates 17,080 diseases connected by 4,050,249 relationships across 10 biological scales. It specifically incorporates drug-disease edges for indications, contraindications, and off-label uses alongside structured text annotations from clinical guidelines.

Why it matters

Fragmented biological and clinical data hinder precision medicine workflows; PrimeKG provides a unified, structured network to train artificial intelligence models for drug repositioning, target discovery, and disease mechanism analysis.

Limits

The abstract describes a data integration resource rather than empirical biological or clinical trials. It does not report error rates, edge accuracy benchmarks, conflict-resolution methods between conflicting source databases, or predictive performance validation of downstream models.

Cited by