Authors
Zachary B Rodriguez, Lindsay Guare, Lannawill Caruth, Katie M Cardone, Christopher Carson, Tess Cherlin, Stephanie Mohammed, Hritvik Gupta, Rachit Kumar, Karl Keat, Shefali S Verma, Anurag Verma
Published in
bioRxiv : the preprint server for biology. Jul 26, 2026. Epub Jul 26, 2026.
Abstract
Electronic health record (EHR)-linked biobanks generate unprecedented genomic and phenotypic datasets, but their scientific utility is constrained by data fragmentation across institutional silos and incompatible computing infrastructures, forcing researchers to rewrite ad-hoc scripts for each new environment. We present the PMBB Geno-Pheno Toolkit, a suite of modular Nextflow pipelines for biobank-scale association analyses. This note focuses on the toolkit's SAIGE family of pipelines - supporting genome-wide (GWAS), exome-wide (ExWAS), and phenome-wide (PheWAS) association testing - together with the companion GWAMA and ExWAS meta-analysis pipelines that enable cross-biobank replication. All components are containerized (Docker/Apptainer) and orchestrated with Nextflow, allowing the same workflows to run unmodified on local HPC clusters, cloud platforms, and the All of Us Research Workbench. Complementary toolkit pipelines for PLINK-based GWAS, polygenic scoring, LD-based clumping, and phenotype harmonization are also available and briefly noted.
The PMBB Geno-Pheno Toolkit is freely available at https://github.com/PMBB-Informatics-and-Genomics/pmbb-geno-pheno-toolkit under MIT open-source license.
PMID:
42539107
Bibliographic data and abstract were imported from PubMed on 01 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 8
- Comments 0