Research

Computational methods for spatial biology and multimodal integration

I develop computational approaches to characterize cell state and decode its drivers. By placing otherwise incompatible assays into a shared anatomical-statistical space, my models describe cells and their neighborhoods in intact tissue, learn what drives cellular decisions, and design experiments that close the loop from data to action.

USHER overview: foundation-model embeddings of out-of-distribution data are mapped back to the model's reference space
Describe

USHER

Adapting foundation models to new data without retraining.

Foundation-model embeddings drift under shifts in protocol, instrument, or imaging, and retraining is costly. USHER maps new data back into the model's established embedding space with fused Gromov-Wasserstein optimal transport and a low-complexity transform, leaving the model's weights untouched.

It removes artifactual variation while preserving biological structure, across sequencing and imaging data.

RECOMB'26 · 2026
SAME overview: space-tearing transforms align serial sections measured with different spatial assays
Describe

SAME

Aligning serial tissue sections measured with different assays.

Spatial assays for proteins, transcripts, and metabolites are usually run on neighboring tissue sections, and existing alignment methods break down when those sections tear, fold, or change anatomically. SAME introduces space-tearing transforms, which allow controlled local breaks during alignment, and uses integer linear programming to maximize cell-type matches across modalities. It improves cell-type alignment accuracy by 20% over existing methods.

In tongue tissue and lung adenocarcinoma, integrating protein and RNA revealed immune subpopulations that neither modality found alone. Integrating protein and metabolite data localized mevalonic acid upregulation to tumor-macrophage niches.

Under revision · 2025
BEELINE pipeline: simulated and experimental single-cell data, containerized GRN inference methods, and evaluation metrics
Drive

BEELINE

A community benchmark for gene regulatory network inference.

BEELINE evaluates methods that infer regulatory networks from single-cell data on synthetic networks, curated models, and experimental data, with reproducible, containerized pipelines and curated gold standards.

It has been cited in over 1,000 subsequent works and is now a standard for benchmarking regulatory network inference.

Nature Methods · 2020

Other work

Fast-SL Genome-scale search for synthetic-lethal gene sets, reformulated as an ℓ1-regularized optimization for a roughly 40x speedup. Bioinformatics · 2015 Paper ↗
CrossPlan Planning genetic crosses to validate mathematical models. Proves the optimal design problem NP-complete and gives fixed-parameter tractable algorithms for practical cases. Bioinformatics · 2018 Paper ↗
Signaling pathways Reconstructing signaling pathways using regular-language constrained paths. Bioinformatics · 2019 Paper ↗
Image-based phenotyping A review of deep learning for image-based cell phenotyping. Curr. Opin. Chem. Biol. · 2021 Paper ↗

Data Science Competitions

Multiple Myeloma DREAM Challenge Top performer, Sub-challenge 1: multiple myeloma patient risk stratification using DNA-based features. The challenge findings identified the epigenetic regulator PHF19 as a marker of aggressive disease. Leukemia · 2020 Results ↗Paper ↗
Broad Obesity ML Challenge Top 5, for designing algorithms that identify genes driving obesity and metabolic disease. CrunchDAO · 2026 Leaderboard ↗Profile ↗

I'm always happy to talk with potential collaborators and students. Get in touch.