Driven by curiosity.
Powered by data.
I'm an ML Systems and Data Science Engineer building end-to-end pipelines at the
intersection of scientific computing, generative AI, and MLOps. My work ranges from
fine-tuning genomic language models on ancient DNA to benchmarking LLMs on hardware
design tasks — two lines of work now written up as arXiv preprints — always with an
eye on reproducibility and real-world deployment.
Most recently I shipped backend AI services in production — multi-language OCR,
async batch pipelines, and document Q&A — where correctness under real rate limits
and messy Unicode mattered more than benchmark numbers.
I care about the full lifecycle: research, model design, containerised deployment,
and production monitoring — a Kubernetes operator that survives a real EC2 Spot
interruption is the same discipline as a retrieval model that survives real
telescope data. If a system doesn't hold up outside a notebook, it isn't done.
Genomic ML
LLM Systems
MLOps
Benchmarking
Sim-to-Real
Scientific Computing