Bioinformatics · Cancer Genomics · Healthcare Data Engineering · Scientific Machine Learning
I am a Research Scientist, Bioinformatician, and Data Scientist working at the intersection of computational biology, cancer genomics, artificial intelligence, biostatistics, and precision medicine.
My research combines genomics, transcriptomics, multi-omics integration, machine learning, and large language models to transform complex biological and health data into evidence for:
- Biomarker discovery
- Therapeutic-target prioritization
- Precision oncology
- AI-assisted drug discovery
- Reproducible biomedical research
I currently contribute to cancer-genomics research as a Bioinformatician at the University of Arkansas for Medical Sciences (UAMS), developing scalable NGS workflows and analyzing TCGA and CPTAC datasets. I am also pursuing an MSc in Bioinformatics at Northeastern University, strengthening my expertise in translational bioinformatics, machine learning, and trustworthy biomedical AI.
Beyond biomedical research, I independently reproduced and extended machine-learning workflows for streamflow prediction and National Water Model bias correction. My HYDRO-FLOW-AI project builds on open workflows from the Alabama Water Institute’s NWM-ML project and research involving a University of West Florida researcher, while remaining an independent project with no claim of institutional affiliation.
In addition, I am conducting collaborative computational biology research with Professor Wei Hsu, whose research program at the ADA Forsyth Institute is connected with Harvard-affiliated biomedical and stem-cell research activities.
My long-term goal is to build trustworthy computational systems that connect biological evidence, clinical data, and artificial intelligence to accelerate scientific discovery and improve patient outcomes.
- 🧬 Computational biology and bioinformatics
- 🎗️ Cancer genomics and precision oncology
- 🧪 NGS analysis and variant interpretation
- 🔗 Multi-omics integration
- 💊 Computational drug discovery
- 🤖 Artificial intelligence and deep learning
- 📚 Biomedical large language models
- ⚡ Agentic AI and retrieval-augmented generation
- ☁️ Cloud-based scientific computing
- 🏥 Healthcare data engineering
- 🌊 Scientific machine learning for hydrology
| Project | Research focus | Repository |
|---|---|---|
| RTK/NRTK TNBC | Patient-level kinase alterations and drug-target prioritization in triple-negative breast cancer | View project |
| HYDRO-FLOW-AI | Extreme-aware streamflow prediction and National Water Model bias correction | View project |
| NIH Clinical Trials Lakehouse | Reproducible healthcare data engineering for clinical-trial analytics | View project |
| CDC Healthcare Streaming ETL | Streaming ingestion, validation, transformation, and public-health analytics | View project |
| TCGA–CPTAC Kafka Platform | Event-driven processing of large-scale cancer multi-omics data | View project |
| Synthetic Variant Calling Benchmark | Reproducible benchmarking of NGS variant-calling workflows | View project |
| Genomic Foundation Models | Transformer-based representation learning for genomic sequences | View project |
| USAG1 Validation | Computational evidence synthesis and therapeutic-target validation | View project |
HYDRO-FLOW-AI is an independent, extreme-aware machine-learning framework for streamflow prediction and site-specific National Water Model bias correction at USGS gauges.
- Leakage-safe temporal training, validation, and testing
- Historical USGS streamflow and climate-data integration
- Site-specific model-performance diagnostics
- Evaluation using RMSE, MAE, bias, and NSE
- Q95 and Q99 high-flow evaluation
- Peak-magnitude error analysis
- Extreme-event detection and threshold-based assessment
- XGBoost residual bias correction
- Quantile-regression uncertainty intervals
- LSTM and Transformer-based forecasting
- River-network graph neural networks
- Explainability and model-drift monitoring
I am conducting collaborative computational biology research with Professor Wei Hsu at the ADA Forsyth Institute, investigating how Wnt/β-catenin signaling relates to craniofacial development, skeletal biology, epithelial differentiation, and osteogenic regulatory programs.
The project analyzes bulk RNA-seq data from a 4-control versus 4-mutant mouse experiment using DESeq2 differential-expression results, human–mouse orthology mapping, branch-specific Wnt gene signatures, ranked gene-set analysis, pathway enrichment, and transcription-factor target-set analysis. Curated resources—including TRRUST, ChEA, Gene Ontology, WikiPathways, FunMap, and STRING—support evidence-audited regulatory and functional-association networks. Particular attention is given to canonical and non-canonical Wnt signaling, TCF/LEF transcription, planar cell polarity, epithelial junction remodeling, craniofacial morphogenesis, and osteogenesis.
To strengthen reproducibility, I implemented the analysis through a version-controlled Docker environment and equivalent Nextflow and Snakemake workflows. Automated GitHub Actions tests verify the container, execute synthetic workflow smoke tests, compare outputs across workflow engines, and confirm that private study data are excluded from public computational assets. The resulting figures, audit tables, and provenance records distinguish exploratory gene-level findings, adjusted pathway or TF enrichment, external network evidence, and hypothesis-generating biological interpretations.
The project is being developed as a reproducible computational biology framework that can support future mechanistic experiments and, following completion and collaborator approval, peer-reviewed publication.
Computational framework for rational drug-combination prioritization using complementary biological and pharmacological evidence.
Identification of compensatory kinase alteration patterns and potential drug-target combinations using TCGA cancer-genomics data.
LLM-supported retrieval, evaluation, and synthesis of biomedical evidence for research decision support.
Transformer-based representation-learning approaches for genomic sequences and downstream biological prediction.
Reproducible lakehouse and streaming architectures for clinical-trial, public-health, and biomedical data.
- AI for precision oncology
- Computational drug discovery
- Cancer multi-omics
- Biomedical large language models
- Genomic foundation models
- Agentic AI for scientific discovery
- Trustworthy and explainable AI
- Scalable scientific computing
- Extreme-aware streamflow forecasting
- MSc Bioinformatics — Northeastern University (in progress)
- MSc Molecular Biology (Bioinformatics) — Umeå University
- BSc Biotechnology and Genetic Engineering — Khulna University
- Advanced Diploma in Data Science and Data Engineering
- Graduate Certificate in Project Management
- Health Informatics — Johns Hopkins University
- Business Analytics and Data-Driven Decision-Making — University of Toronto
- Project Management
- Cloud Computing
- Machine Learning and Data Science
- Trustworthy and causal AI
- Agentic and multi-agent systems
- Biomedical large language models
- Genomic foundation models
- Scalable multi-omics analytics
---
I welcome research and open-source collaboration in:
- Computational biology and bioinformatics
- Cancer genomics and precision medicine
- Biomedical artificial intelligence
- Computational drug discovery
- Healthcare data engineering
- Scientific machine learning
- Reproducible research software
- Regenerative medicine and stem cell research

