Research Projects

Interpretable machine learning in population genetics

Machine-learning methods can detect complex patterns in genomic data, but their predictions are often difficult to interpret. My research combines the flexibility of machine learning with established population-genetic theory. The aim is to build models that are powerful and interpretable to better understand what they learn from genomic data.

Sampling bias in bacterial genome databases

Public bacterial genome databases contain large amounts of data, but their samples are often strongly biased towards particular strains of interest. We study how this affects estimates of bacterial diversity and adaptation and have developed a phylogeny-based method to detect and reduce oversampling.

Polygenic adaptation

Many traits are influenced by a large number of genes, making their response to natural selection difficult to predict. During my PhD, I used mathematical models and simulations to study how complex traits adapt and how this adaptation is reflected at the gene-frequency level. In particular, we investigated when adaptation occurs through large changes at a few loci or small allele-frequency shifts across many loci.