Precision Disease Modeling

We build predictive models that stratify patients and reveal disease subtypes.

Better stratification is the foundation of precision medicine.

We integrate longitudinal health records with multimodal clinical data to identify patient subgroups, uncover latent physiology, and predict risk.

Patient stratificationDisease subtypingPredictive modelingPrecision medicine
Layered disease trajectories emerging from a shared origin
01Patterns · 2021

Jessica K. De Freitas et al.

Phe2vec: Automated disease phenotyping based on unsupervised embeddings from electronic health records

Read the publication
EHR phenotypingUnsupervised learningCohort discovery
Computational phenotyping

Learning disease cohorts from the record instead of hand-built rules

Figure 1Phe2vec framework for medical-concept embedding, disease phenotyping, and cohort selection. Figure reproduced from the publication.
Question

Can scalable disease cohorts be identified from heterogeneous longitudinal EHR data without building a new ruleset for every condition?

Approach

Phe2vec learned unsupervised embeddings from diagnoses, medications, procedures, laboratory tests, vital signs, and clinical notes across 1.9 million patients, then used those representations to rank patients for ten diseases.

Finding

Phe2vec matched or outperformed a widely used rule-based standard for nine of ten diseases in head-to-head chart review.

Why it matters

Cohort definition becomes a representation-learning problem—supporting scalable phenotyping beyond brittle code lists and hand-built logic.

02JACC: Cardiovascular Imaging · 2022

Akhil Vaid et al.

Using Deep-Learning Algorithms to Simultaneously Identify Right and Left Ventricular Dysfunction From the Electrocardiogram

Read the publication
AI-ECGVentricular functionExternal validation
Latent physiology

Can a single ECG expose the state of both ventricles?

Central IllustrationDeep-learning workflow for identifying left- and right-ventricular dysfunction from ECGs, including performance and explainability. Figure reproduced from the publication.
Question

Can routinely collected ECG waveforms quantify both left- and right-sided cardiac dysfunction that ordinarily requires echocardiography?

Approach

Across five hospitals, deep-learning models were trained on more than 700,000 ECG–echo pairs for left-ventricular ejection fraction and more than 760,000 pairs for right-ventricular dysfunction or dilation, with external validation at a held-out hospital.

Finding

Reduced LVEF was detected with an AUROC of 0.94 in both internal and external validation; the composite right-ventricular outcome reached 0.84 in both settings.

Why it matters

A routine ECG can screen for physiology that normally requires an echocardiogram.

03Cardiovascular Digital Health Journal · 2022

Hossein Honarvar et al.

Enhancing convolutional neural network predictions of electrocardiograms with left ventricular dysfunction using a novel sub-waveform representation

Read the publication
ECG representationExplainabilityFairness
Signal representation

Changing ECG data representation, not the model

Figure 1Transformation of full ECG waveforms into aligned, multilead sub-waveform representations. Figure reproduced from the publication.
Question

Can a physiologically aligned input representation improve prediction and interpretability without changing the underlying neural-network architecture?

Approach

For 92,446 patients, ten-second ECGs were discretized into heartbeat-aligned sub-waveforms and compared with conventional full-waveform inputs for detecting left-ventricular dysfunction.

Finding

The sub-waveform representation increased AUROC by 2% and AUPRC by 10%, while reducing predictive uncertainty across demographic subgroups.

Why it matters

Model performance depends on how disease signals are represented. Better inputs can improve accuracy, interpretability, and fairness at the same time.

04European Heart Journal – Digital Health · 2022

Sulaiman S. Somani et al.

Development of a machine learning model using electrocardiogram signals to improve acute pulmonary embolism screening

Read the publication
Multimodal modelingPulmonary embolismClinical screening
Multimodal screening

Combining waveform and chart data to screen for pulmonary embolism

Graphical AbstractECG and EHR inputs, model-development framework, comparative performance, and proposed screening workflow. Figure reproduced from the publication.
Question

Can the electrical signature of the heart add useful information to clinical data when deciding which patients need CT pulmonary angiography?

Approach

The study linked 320,746 ECGs and encounter-level clinical data to 23,793 CTPAs from 21,183 patients, comparing ECG-only, EHR-only, and fused multimodal models.

Finding

The fusion model reached an AUROC of 0.81 overall and 0.84 in a clinically reviewed subset, exceeding four established screening scores.

Why it matters

Combining modalities can represent disease probability better than either waveform or chart data alone and may reduce unnecessary imaging.

05Patterns · 2021

Tingyi Wanyan et al.

Contrastive learning improves critical event prediction in COVID-19 patients

Read the publication
Contrastive learningCritical eventsClass imbalance
Robust clinical prediction

Keeping patient representations useful when outcomes are rare

Figure 1Contrastive-learning architecture, patient representation space, time windows, and event-time selection. Figure reproduced from the publication.
Question

Can representation learning improve near-term prediction when mortality, intubation, and ICU transfer are clinically important but statistically imbalanced?

Approach

Across five hospitals, contrastive loss was added to sequential EHR models predicting mortality, intubation, and ICU transfer 24 and 48 hours in advance, then tested in both full and restricted imbalanced cohorts.

Finding

In the restricted intubation cohort, contrastive learning improved AUROC from 0.80 to 0.88 and AUPRC from 0.35 to 0.45.

Why it matters

Separating clinically meaningful patient states in representation space can make models more robust when the outcomes that matter most are the least common.

Precision Disease Modeling
Better representations make better questions possible: who belongs, what is changing, and what comes next.