Skip to main content

Bioinformatics Scientist

Leads automation of demultiplexing, quality control, gene expression analysis, and standardized data analysis using containerized workflows (Docker, Argo Workflows, Nextflow) deployed across Kubernetes, SLURM, and cloud infrastructure. Responsible for diagnosing and resolving complex sequencing and bioinformatics failures, optimizing pipeline performance, and ensuring analyses meet internal quality standards and customer requirements.

Acts as the primary Bioinformatics scientist interface between sequencing operations, wet-lab teams, and customers—translating experimental design and data quality constraints into robust, reproduciable bioinformatics outputs.

I am also responsible for diagnosing and resolving complex sequencing and bioinformatics failures, optimizing pipeline performance, and maintaining reproduciblity across diverse study designs and platforms.

Key Contributions & Impact
  • Production-scale NGS Pipeline Ownership: Designed, productionized, and maintained multiple end-to-end NGS analysis pipelines supporting RNA-seq, whole-genome sequencing (WGS), amplicon sequencing (16S/ITS), metagenomics, and DNA methylation assays. These pipelines are used routinely for customer projects and are engineered for reproducibility, traceability, and consistent turnaround across heterogeneous datasets.

  • High-Throughput WGBS Processing at Scale: Built and operated a scalable Whole-Genome Bisulfite Sequencing (WGBS) pipeline capable of processing >200 GB per project at >30× genome coverage, with automated QC checkpoints and robust methylation calling. The workflow is optimized for concurrent production workloads and minimizes manual intervention.

  • Multi-Platform Demultiplexing & QC Automation: Implemented automated demultiplexing and run-level QC workflows compatible with Illumina (NovaSeq), PacBio, 10x Genomics, Oxford Nanopore (ONT), and AVITI. These workflows standardized index validation, sample assignment, and QC reporting, significantly reducing manual handling and improving data reliability and turnaround time.

  • RNA-seq Data Analysis and Reports Delivery: Led RNA-seq analysis for 80+ customer projects spanning bacterial, plant, fish, mouse, and human samples. Analyses were executed using STAR, HISAT2, and Salmon within containerized pipelines, ensuring consistent gene-expression quantification, compatibility across reference annotations, and reproducible downstream results.

  • Custom Variant & Mutational Analysis Pipelines: Developed targeted variant and mutational analysis workflows for plasmid and DNA-based assays using GATK, SAMtools, and bcftools, enabling accurate mutation detection, standardized filtering, and customer-ready reporting aligned with validation and QC requirements.

  • Standardization of Downstream Analysis & Reporting: Standardized preprocessing, filtering, and downstream analysis workflows using Python and R, including automated validation, normalization, and statistical analysis. This improved cross-project consistency and supported reproducible differential-expression analysis and microbiome profiling.

  • Failure Diagnosis & Root-Cause Resolution: Systematically diagnosed and resolved complex sequencing and bioinformatics failures, including index misassignment, adapter contamination, primer artifacts, low-complexity libraries, and mapping biases. These interventions reduced reprocessing costs, minimized customer delays, and improved overall data quality.

  • Customer-Facing Technical Leadership: Served as a primary technical point of contact for customers, working directly with wet-lab teams and external collaborators to convert experimental goals and sequencing constraints into clear analytical strategies, reliable deliverables, and biologically interpretable results.

Technologies Used

Python, R, Bash, Docker, Argo Workflows, Nextflow, Kubernetes, SLURM, GATK, Samtools, bcftools, STAR, HISAT2, Salmon, SPAdes, Unicycler, QUAST, Bakta, abstar, MiXCR