Research

I work on interpretable machine learning for biological discovery, with a focus on drug-resistance prediction in Mycobacterium tuberculosis. My models put known biology, such as protein structure and evolutionary constraint, into the model itself, so that they stay accurate and explainable on new lineages, genes and populations.

BIG-TB benchmark · Under revision

Can sequence-based models predict resistance, and do their attributions recover the known causal loci, or just fit the phenotype?

A large, multimodal benchmark of 17,942 M. tuberculosis isolates across 11 WHO-priority antibiotics that puts DNA- and protein-based models on equal footing.

17,942isolates scored on 11 antibiotics, including lineages the models never saw in training

What I did

  • Led the protein-side benchmark: CNN and Transformer models, regression baselines and frozen ESM-2 embeddings
  • Leave-one-lineage-out splits test whether performance survives a shift in population structure instead of fitting clonal lineage signal
  • SHAP attributions scored against WHO-catalogue resistance loci (precision and recall at the top-k residues), with one evaluation harness shared across every model family
BIG-TB phenotype dataset pipeline: extracting VCFs, reconstructing DNA, and translating to protein sequence BIG-TB training and evaluation pipeline across data encodings, training data, and ML models

FARM forecasting · Under review at PNAS

Trained on one snapshot of the WHO resistance catalogue, can a model forecast which uncertain variants will later be called resistance-causing?

FARM learns only from variants with known effects in the 2021 WHO catalogue, then is tested prospectively on the “uncertain significance” variants that WHO reclassified in 2023.

80.7%of the variants later reclassified as resistant were recovered in that prospective test

What I did

  • Built 25 features per variant: 3D structural proximity, Rosetta energetics including ΔΔG, ESM-2 protein language model features and AAIndex physicochemical descriptors
  • Random forest and logistic regression classifiers, interpreted with SHAP
  • The model now scores 4,525 uncertain-significance variants to prioritize candidates for experimental follow-up (Supplementary Data 1)
FARM framework: multimodal feature integration, TB mutation resistance forecasting, and performance evaluation

Structure-aware models · Published

Does knowing where a mutation sits in 3D make resistance prediction better, and more interpretable?

Two published papers on putting protein structure into the model: Fused Ridge (ICLR MLGenX 2025) and 3D mutational clustering (eLife 2025).

94.6%F1 for 3D proximity, against 92.8% for sequence distance and 80.8% for the clustering score (eLife 2025)

What I did

  • Fused Ridge (first author): a linear model whose structure-based penalty pulls the coefficients of neighbouring residues together. Across nine resistance genes it reached a mean AUC of 0.766, against 0.755 for plain ridge and 0.603 for zero-shot ESM-2 scoring
  • 3D mutational clustering (with Anna Green, Rodrigo Vargas Jr. and Maha Farhat): I built the classifier comparing 3D proximity, 1D sequence proximity and the Getis-Ord clustering score across 641 labeled variants
Figure from Green, Tasmin et al., eLife 2025: the Getis-Ord workflow and the clustering of resistance mutations on the KatG structure

Cropped from Figure 2 of Green, Tasmin, Vargas Jr., Farhat, eLife 14:RP109450 (2025), CC BY 4.0. The full figure is in the paper.

Evolutionary augmentation · Ongoing dissertation work

When does evolutionary information genuinely help protein-level models trained on sparse data?

Multi-species protein homologs and language-model scoring, used to enrich the small training sets that resistance prediction usually has to work with.

What I do

  • Use homologs from UniProt and language-model plausibility scores to expand small training sets
  • Evaluate with leakage-aware protocols (homology-filtered test sets and nested cross-validation), so that real gains can be told apart from benchmark artifacts
  • Goal: establish when augmentation helps, not just whether it can

Other interests

  • Multi-modal integration of protein and genomic embeddings
  • Transfer learning for cross-species resistance prediction
  • Benchmark design and interpretability evaluation pipelines

For full paper details, see my publications.