Professional Summary
I am a Bioinformatics Scientist specializing in scalable NGS workflow automation and multi-omics data analysis. With extensive experience in designing and deploying containerized pipelines, I have worked across various sequencing platforms such as Illumina, AVITI, Nanopore and PacBio and cloud computing environments to accelerate research and clinical applications. My expertise spans from pipeline development to deep learning-based biomarker discovery, with a particular focus on liver disease research.
My career has been focused on bridging the gap between computational biology and clinical applications. I have worked across various sectors including biotechnology companies, diagnostic laboratories, and research institutions, developing solutions that accelerate research and improve patient outcomes.
Core Expertise & Skills
Computational & Technical Skills
| Category | Technologies & Tools |
|---|---|
| Programming Languages | Python, R, Bash, Linux |
| Workflow Management | Snakemake, Nextflow, Argo, Docker, Kubernetes |
| Cloud & Infrastructure | AWS (S3, EC2, Batch), SLURM, Kubernetes, Git, Bitbucket, CI/CD concepts, Data versioning, Secure data transfer (SFTP, IAM, BOX) |
| Genome Data Analysis | Bulk RNA-seq, Single-cell RNA-seq, Long-read RNA-seq, BRB-seq, Whole-genome sequencing (WGS), Whole-exome sequencing (WES), Variant analysis, Mutational analysis, DNA methylation analysis, 16S rRNA amplicon analysis, ITS amplicon analysis, Shotgun metagenomics, De novo genome assembly, Comparative genomics, Multi-omics integration, Biomarker discovery analysis, Clinical outcome and survival analysis |
| Bioinformatics Tools | STAR, HISAT2, Salmon, Kallisto, RSEM, StringTie, StringTie2, TieBrush, featureCounts, HTSeq, DESeq2, edgeR, limma-voom, Sleuth, tximport, IsoformSwitchAnalyzeR, GATK (HaplotypeCaller, Mutect2, BQSR), Samtools, bcftools, FreeBayes, Strelka, CNVkit, Manta, Delly, DeepVariant, VEP, ANNOVAR, SnpEff, SpliceAI, AlphaMissense, Minimap2, NGMLR, FLAIR, TALON, Iso-Seq pipelines (PacBio), Nanopore direct RNA-seq workflows, Cell Ranger, STARsolo, Seurat, Kraken2, Bracken, MetaPhlAn, SPAdes, MEGAHIT, Flye, CheckM, QUAST, Bakta, Prokka, FastQC, MultiQC |
| Pathway Analysis | Gene Ontology (GO) enrichment analysis (BP, MF, CC), KEGG pathway analysis, Reactome pathway analysis, WikiPathways analysis, BioCarta pathway analysis, Panther pathway analysis, MSigDB gene set analysis, Hallmark gene set analysis, Gene Set Enrichment Analysis (GSEA), preranked GSEA, single-sample GSEA (ssGSEA), Gene Set Variation Analysis (GSVA), Pathway activity scoring, Over-representation analysis (ORA), Functional class scoring, Network-based pathway analysis |
| Sequencing Platforms | AVITI, Illumina (NovaSeq), PacBio, Nanopore (ONT), 10x Genomics |
| Machine Learning & Statistical Modeling | Regression, Logistic Regression, Random Forest, XGBoost, Feedforward Neural Networks (FNN), Autoencoders, Clustering analysis, Survival analysis, Feature selection, Dimensionality reduction |
| Deep Learning | Feedforward neural networks (FNN), Autoencoders (AE), Variational autoencoders (VAE), Denoising autoencoders, Multi-layer perceptrons (MLP), Transformer architectures, Attention mechanisms, Representation learning, Latent space modeling, Multimodal fusion models, Deep survival models, Embedding-based modeling |
Key Achievements
Built production-scale NGS Pipeline: Designed, deployed, and maintained scalable, containerized Argo Workflows and Nextflow pipelines for RNA-seq, WGS/WES, amplicon (16S/ITS), and methylation workflows, supporting routine processing of large, multi-project sequencing runs with standardized QC, reproducibility, and automated delivery.
Operational Automation & Efficiency: Reduced manual effort and turnaround time by automating routine bioinformatics operations including demultiplexing validation, QC checks, re-demultiplexing, reporting, and data delivery—through workflow orchestration, scripting, and event-driven triggers, improving reliability and reducing human error in production environments.
Advanced RNA-seq & Splicing Analysis Expertise: Performed in-depth RNA-seq analyses across bulk, single-cell, and long-read datasets, including transcript-level quantification, alternative splicing detection, isoform interpretation, and systematic QC diagnostics (gene body bias, junction support, mapping artifacts), often troubleshooting complex or failed datasets.
Multi-Platform Sequencing Support: Built and validated analysis pipelines compatible across Illumina (NovaSeq), PacBio, Oxford Nanopore (ONT), 10x Genomics, and AVITI, enabling consistent downstream analysis despite platform-specific biases and data characteristics.
Multi-Omics & Biomarker Modeling: Developed RNA-seq–anchored biomarker discovery pipelines integrating transcriptomic, clinical, and imaging-derived features for liver disease (HCC / NAFLD / NASH), applying classical ML and deep learning models with rigorous validation, feature stability analysis, and biological interpretation.
Disease-Focused ML/DL Modeling: Implemented and evaluated ML and deep learning models including autoencoder-based representations and transformer-style architectures—for disease classification, risk stratification, and early-stage signal detection, emphasizing interpretability, robustness, and avoidance of data leakage.
Clinical & Microbial Genomics Applications Built and curated in-house reference databases and analysis workflows for rapid identification of clinically relevant bacterial and fungal pathogens from sequencing data, supporting diagnostic and translational use cases.
Cross-Platform Epigenomics Analysis: Conducted genome-wide DNA methylation analyses comparing Illumina and ONT platforms, assessing concordance, coverage biases, and suitability for clinical and translational applications.
Personal Technical Growth & Leadership: Transitioned from executing analyses to owning end-to-end systems—from study design review and failure-mode anticipation to pipeline architecture, automation, and stakeholder communication—while building deep specialization in RNA-seq, splicing biology, and translational data interpretation.
Current Focus
Currently, I am working on:
- RNA-seq & Splicing: Developing deep, end-to-end expertise in RNA-seq analysis with a focus on splicing, isoform quantification, and transcript-level biology across bulk, single-cell, and long-read technologies.
- Multi-omics Integration: Combining transcriptomic, genomic, and clinical data for comprehensive disease understanding.
- Deep Learning for Biomarker Discovery: Developing transformer-based models for NASH and liver disease biomarker identification.
- Scalable & Reproducible Genomics Pipeliness: Building containerized pipelines that can process large-scale sequencing data efficiently.
- Automation: Building a automation pipline for routine works to reduce a mannual work using, n8n, Python, Bash, cron, APIs, Argo and AI-assisted tools.
Research Interests
My research interests center on leveraging computational approaches to solve complex biological problems:
- Precision Medicine: Using multi-omics data to develop personalized treatment strategies
- Disease Biomarker Discovery: Applying machine learning and deep learning to identify predictive biomarkers
- Workflow Optimization: Improving the efficiency and reproducibility of bioinformatics analyses
Education & Training
- Master of Science in Bioinformatics - Hood College, Maryland, USA
- Master of Science in Pharmaceutical Technology
- Bachelor of Science in Pharmacy
My educational background in pharmaceutical sciences and bioinformatics provides a unique perspective on translating computational findings into data-driven drug discovery and therapeutic decision-making, with a clear path toward clinical application
Publications & Presentations
- Multi-Modal Fusion Framework for HCC Prediction - AASLD 2025
- Portable DNA Sequencing Technologies for Far-Forward Operations - MHSRS 2025
- Splice Junction Analysis from Public RNA-seq Data - Capstone Project, Hood College 2024
Let’s Connect
I’m always interested in discussing new research opportunities, collaborations, or innovative approaches to bioinformatics challenges. Feel free to reach out through the contact form or connect with me on LinkedIn or GitHub.
“My North Star is building scalable computational systems that integrate genomics data and machine learning to support target drug discovery, repurposing, and precision medicine decisions.”