Bioinformatics Engineer
Bioinformatics Engineer (Contract)
Project-based consulting role concurrent with full-time responsibilities, focusing on deep learning-based biomarker discovery for liver disease. Responsible for designing and implementing deep learning models for biomarker identification, integrating multi-modal data sources, and delivering reproducible analyses that support downstream modeling and research reporting.
Key Contributions & Impact
ML/DL Data Analysis: Conducted end-to-end analysis of bulk and single-cell RNA-seq datasets for liver disease cohorts, with heavy use of public resources such as NCBI GEO and ARCHS4. Led extensive data cleaning, normalization, gene filtering, and metadata harmonization across studies, producing reproducible analyses and clear, biologically interpretable summaries to support downstream modeling and research reporting.
Single-cell and genomics-oriented modeling and evaluation: Applied scRNA-seq embedding and latent space models with batch-aware and transfer-learning strategies, conducted rigorous model evaluation through ablation and feature-importance analyses to control overfitting and data leakage, and built scalable data handling and reporting workflows using pandas, NumPy, AnnData/H5AD, CSV/Parquet formats, with results communicated through clear visualizations in matplotlib and seaborn
Deep Learning–Based Biomarker Discovery: Designed, trained, and evaluated deep learning models for biomarker identification using PyTorch and Hugging Face, including transformer-based models, autoencoders, and variational autoencoders (VAE) applied to both scRNA-seq and bulk RNA-seq datasets in NASH.
Baseline ML and Model Benchmarking: Implemented and benchmarked classical machine learning models (e.g., regularized regression, Random Forest, gradient boosting) alongside transformer-style architectures (NanoGPT-inspired) for cell-type prediction and disease-state classification, ensuring proper baselines, robust comparisons, and meaningful sanity checks.
Feature Engineering and Interpretability: Applied feature selection, dimensionality reduction (PCA, UMAP), and embedding-based approaches to high-dimensional scRNA-seq data to improve model stability and interpretability, with insights visualized using matplotlib and seaborn.
Public and Internal Data Integration: Performed systematic data mining and harmonization using NCBI GEO, ENCODE, ENSEMBL, and the UCSC Genome Browser, integrating heterogeneous public datasets with internal cohorts to support meta-analysis, cross-study validation, and improved model generalizability.
Data Preprocessing Pipelines for ML/DL Workflows: Built Python-based preprocessing pipelines to clean, normalize, and restructure RNA-seq expression matrices and metadata into CSV, Parquet, and H5AD formats, creating model-ready inputs for deep learning and multi-omics workflows.
Project Coordination and Research Execution: Managed timelines and deliverables for a NAFLD/NASH meta-analysis review project using Jira, supporting structured collaboration, milestone tracking, and timely manuscript development.
Grant and Strategic Research Support: Contributed to the preparation of an NIH grant proposal focused on foundational AI models for precision biomarker development in fatty liver disease, supporting methodological framing, feasibility assessment, and translational relevance.
Multi-Modal Data Integration (Transcriptomics + Clinical + Imaging): Integrated bulk RNA-seq features with clinical variables and imaging-derived representations to support liver disease risk modeling and biomarker discovery, enabling direct comparison between single-modality and fused models.
Fusion Modeling and Ablation Evaluation: Implemented late-fusion and representation-level fusion strategies, performing ablation analyses (RNA-seq only, clinical only, imaging only, integrated) to quantify modality contributions and reduce overfitting to any single data source.
End-to-End Multi-Modal Pipeline Preparation: Developed preprocessing and harmonization scripts to align sample identifiers, labels, and metadata across modalities, producing clean, audit-ready training datasets suitable for ML/DL workflows and downstream reporting.
Technologies Used
scikit-learn, XGBoost, classical regression models, tree-based models, feature selection methods, dimensionality reduction (PCA, UMAP), clustering algorithms, model validation, cross-validation strategies, PyTorch, Hugging Face Transformers, transformer-based architectures, autoencoders (AE), variational autoencoders (VAE), embedding-based representation learning, attention mechanisms, multimodal fusion models