I help research groups, scientific organizations, and data-intensive teams build software and analytical workflows that are robust, maintainable, and reproducible. I have worked with academic research groups, clinical laboratories, and multinational companies.
My strongest work sits at the intersection of R software engineering, scientific computing, statistics, bioinformatics, and reproducible research. I am available for hourly consulting, fixed-scope projects, code review, and short-term contract engagements.
I design, build, modernize, and maintain R packages and analytical software. This includes package architecture, APIs, testing, documentation, performance improvement, dependency management, and integration with existing workflows. For performance-critical components, I also develop compiled extensions using R’s C, C++, and Fortran interfaces.
I am particularly comfortable taking research code that grew organically and turning it into a clearer, tested, documented system that other people can reliably use and maintain.
I design end-to-end analytic pipelines that make complex research easier to reproduce, audit, update, scale, and hand off. Depending on the problem, that may include data ingestion, cleaning, transformation, statistical analysis, simulation, visualization, automated reporting, dependency-aware workflow orchestration, and version-controlled research artifacts.
The implementation can range from lightweight R-based pipelines using tools such as targets and renv to large-scale, multi-step scientific workflows using systems such as Snakemake and conda. I have built genomics pipelines that move from raw sequencing data through analysis to dynamically generated publications, with figures, tables, and reported results updating directly from the underlying analysis. I also work with SQL, Linux/Bash, LaTeX/Sweave, Quarto/R Markdown, Shiny, and the broader Posit ecosystem where appropriate.
I can support open-ended technical and quantitative problems including:
I develop interactive analytical tools and lightweight applications when a research workflow needs to be usable by people who are not writing code themselves. This can include Shiny applications, internal dashboards, automated reports, and R-based APIs. For example, I have developed a Shiny application for sample tracking and quality control in a clinical laboratory setting.
I am comfortable inheriting messy or undocumented analytical codebases. Typical engagements include debugging, refactoring, restructuring, adding tests, improving documentation, reviewing analytical correctness, and making workflows reproducible enough for publication, collaboration, or production use.
I also provide technical code review and evaluation for R, statistics, scientific computing, reproducibility, and research-software projects.
I have designed and taught graduate-level R programming courses and can provide technical documentation, package vignettes, onboarding materials, reproducibility instructions, SOPs, and live remote training for research or technical teams.
Before completing my MD and PhD in Bioinformatics and Computational Biology at the University of North Carolina at Chapel Hill, I worked in computational toxicology at the U.S. EPA National Center for Computational Toxicology. I designed and implemented the initial versions of the R package and SQL-backed analytical infrastructure that became tcpl for high-throughput chemical screening data.
My consulting experience has ranged from helping a small research group strengthen an NIH grant proposal to developing scientific software and advising a multinational company on high-throughput screening workflows.
You can see representative public work on the Selected Work page and on GitHub.
If you have a software, data, reproducibility, or scientific-computing problem that may be a fit, email me with a short description of the project, the current state of the work, and any relevant timeline.