Skip to main content

Jan 2025 - Present
Bioinformatics Scientist
AmpSeq LLC. Maryland, USA

Bioinformatics Scientist

Leads automation of demultiplexing, quality control, gene expression analysis, and standardized data analysis using containerized workflows (Docker, Argo Workflows, Nextflow) deployed across Kubernetes, SLURM, and cloud infrastructure. Responsible for diagnosing and resolving complex sequencing and bioinformatics failures, optimizing pipeline performance, and ensuring analyses meet internal quality standards and customer requirements. Acts as the primary Bioinformatics scientist interface between sequencing operations, wet-lab teams, and customers—translating experimental design and data quality constraints into robust, reproduciable bioinformatics outputs. I am also responsible for diagnosing and resolving complex sequencing and bioinformatics failures, optimizing pipeline performance, and maintaining reproduciblity across diverse study designs and platforms. Key Contributions & Impact Production-scale NGS Pipeline Ownership: Designed, productionized, and maintained multiple end-to-end NGS analysis pipelines supporting RNA-seq, whole-genome sequencing (WGS), amplicon sequencing (16S/ITS), metagenomics, and DNA methylation assays. These pipelines are used routinely for customer projects and are engineered for reproducibility, traceability, and consistent turnaround across heterogeneous datasets. High-Throughput WGBS Processing at Scale: Built and operated a scalable Whole-Genome Bisulfite Sequencing (WGBS) pipeline capable of processing >200 GB per project at >30× genome coverage, with automated QC checkpoints and robust methylation calling. The workflow is optimized for concurrent production workloads and minimizes manual intervention. Multi-Platform Demultiplexing & QC Automation: Implemented automated demultiplexing and run-level QC workflows compatible with Illumina (NovaSeq), PacBio, 10x Genomics, Oxford Nanopore (ONT), and AVITI. These workflows standardized index validation, sample assignment, and QC reporting, significantly reducing manual handling and improving data reliability and turnaround time. RNA-seq Data Analysis and Reports Delivery: Led RNA-seq analysis for 80+ customer projects spanning bacterial, plant, fish, mouse, and human samples. Analyses were executed using STAR, HISAT2, and Salmon within containerized pipelines, ensuring consistent gene-expression quantification, compatibility across reference annotations, and reproducible downstream results. Custom Variant & Mutational Analysis Pipelines: Developed targeted variant and mutational analysis workflows for plasmid and DNA-based assays using GATK, SAMtools, and bcftools, enabling accurate mutation detection, standardized filtering, and customer-ready reporting aligned with validation and QC requirements. Standardization of Downstream Analysis & Reporting: Standardized preprocessing, filtering, and downstream analysis workflows using Python and R, including automated validation, normalization, and statistical analysis. This improved cross-project consistency and supported reproducible differential-expression analysis and microbiome profiling. Failure Diagnosis & Root-Cause Resolution: Systematically diagnosed and resolved complex sequencing and bioinformatics failures, including index misassignment, adapter contamination, primer artifacts, low-complexity libraries, and mapping biases. These interventions reduced reprocessing costs, minimized customer delays, and improved overall data quality. Customer-Facing Technical Leadership: Served as a primary technical point of contact for customers, working directly with wet-lab teams and external collaborators to convert experimental goals and sequencing constraints into clear analytical strategies, reliable deliverables, and biologically interpretable results. Technologies Used Python, R, Bash, Docker, Argo Workflows, Nextflow, Kubernetes, SLURM, GATK, Samtools, bcftools, STAR, HISAT2, Salmon, SPAdes, Unicycler, QUAST, Bakta, abstar, MiXCR
Jan 2025 - Present (Project-based consulting)
Bioinformatics Engineer
Karyon Bio. California, USA

Bioinformatics Engineer (Contract)

Project-based consulting role concurrent with full-time responsibilities, focusing on deep learning-based biomarker discovery for liver disease. Responsible for designing and implementing deep learning models for biomarker identification, integrating multi-modal data sources, and delivering reproducible analyses that support downstream modeling and research reporting. Key Contributions & Impact ML/DL Data Analysis: Conducted end-to-end analysis of bulk and single-cell RNA-seq datasets for liver disease cohorts, with heavy use of public resources such as NCBI GEO and ARCHS4. Led extensive data cleaning, normalization, gene filtering, and metadata harmonization across studies, producing reproducible analyses and clear, biologically interpretable summaries to support downstream modeling and research reporting. Single-cell and genomics-oriented modeling and evaluation: Applied scRNA-seq embedding and latent space models with batch-aware and transfer-learning strategies, conducted rigorous model evaluation through ablation and feature-importance analyses to control overfitting and data leakage, and built scalable data handling and reporting workflows using pandas, NumPy, AnnData/H5AD, CSV/Parquet formats, with results communicated through clear visualizations in matplotlib and seaborn Deep Learning–Based Biomarker Discovery: Designed, trained, and evaluated deep learning models for biomarker identification using PyTorch and Hugging Face, including transformer-based models, autoencoders, and variational autoencoders (VAE) applied to both scRNA-seq and bulk RNA-seq datasets in NASH. Baseline ML and Model Benchmarking: Implemented and benchmarked classical machine learning models (e.g., regularized regression, Random Forest, gradient boosting) alongside transformer-style architectures (NanoGPT-inspired) for cell-type prediction and disease-state classification, ensuring proper baselines, robust comparisons, and meaningful sanity checks. Feature Engineering and Interpretability: Applied feature selection, dimensionality reduction (PCA, UMAP), and embedding-based approaches to high-dimensional scRNA-seq data to improve model stability and interpretability, with insights visualized using matplotlib and seaborn. Public and Internal Data Integration: Performed systematic data mining and harmonization using NCBI GEO, ENCODE, ENSEMBL, and the UCSC Genome Browser, integrating heterogeneous public datasets with internal cohorts to support meta-analysis, cross-study validation, and improved model generalizability. Data Preprocessing Pipelines for ML/DL Workflows: Built Python-based preprocessing pipelines to clean, normalize, and restructure RNA-seq expression matrices and metadata into CSV, Parquet, and H5AD formats, creating model-ready inputs for deep learning and multi-omics workflows. Project Coordination and Research Execution: Managed timelines and deliverables for a NAFLD/NASH meta-analysis review project using Jira, supporting structured collaboration, milestone tracking, and timely manuscript development. Grant and Strategic Research Support: Contributed to the preparation of an NIH grant proposal focused on foundational AI models for precision biomarker development in fatty liver disease, supporting methodological framing, feasibility assessment, and translational relevance. Multi-Modal Data Integration (Transcriptomics + Clinical + Imaging): Integrated bulk RNA-seq features with clinical variables and imaging-derived representations to support liver disease risk modeling and biomarker discovery, enabling direct comparison between single-modality and fused models. Fusion Modeling and Ablation Evaluation: Implemented late-fusion and representation-level fusion strategies, performing ablation analyses (RNA-seq only, clinical only, imaging only, integrated) to quantify modality contributions and reduce overfitting to any single data source. End-to-End Multi-Modal Pipeline Preparation: Developed preprocessing and harmonization scripts to align sample identifiers, labels, and metadata across modalities, producing clean, audit-ready training datasets suitable for ML/DL workflows and downstream reporting. Technologies Used scikit-learn, XGBoost, classical regression models, tree-based models, feature selection methods, dimensionality reduction (PCA, UMAP), clustering algorithms, model validation, cross-validation strategies, PyTorch, Hugging Face Transformers, transformer-based architectures, autoencoders (AE), variational autoencoders (VAE), embedding-based representation learning, attention mechanisms, multimodal fusion models

Work Experience