Learning disease cohorts from the record instead of hand-built rules
Can scalable disease cohorts be identified from heterogeneous longitudinal EHR data without building a new ruleset for every condition?
Phe2vec learned unsupervised embeddings from diagnoses, medications, procedures, laboratory tests, vital signs, and clinical notes across 1.9 million patients, then used those representations to rank patients for ten diseases.
Phe2vec matched or outperformed a widely used rule-based standard for nine of ten diseases in head-to-head chart review.
Cohort definition becomes a representation-learning problem—supporting scalable phenotyping beyond brittle code lists and hand-built logic.
