Authors
Sujin Seo, Sungho Won, Kyungtaek Park
Published in
Bioinformatics (Oxford, England). Sep 02, 2026. Epub Sep 02, 2026.
Abstract
Single-cell RNA sequencing (scRNA-seq) enables high-resolution profiling of cellular heterogeneity, yet batch effects remain a critical challenge in data integration. Existing batch correction methods often assume homogeneous batch effect across cell types, operate in reduced-dimensional space leading to potential loss of biological information, and require extensive computational resources.
Here, we introduce COBRA, a linear model-based batch correction method that explicitly adjusts cell-type-specific batch effect. By orthogonalizing batch-associated parameters with respect to biological variables, COBRA removes technical artifacts while preserving biologically meaningful transcriptional differences. When cell type annotations are unavailable, COBRA implements an iterative clustering algorithm to estimate pseudo-cell types while accounting for batch effects. COBRA retains the full gene expression matrix, ensuring seamless integration for downstream analyses. We evaluated COBRA across simulated and real-world datasets, including type 2 diabetes and COVID-19 datasets. COBRA outperformed in terms of batch mixing efficiency, preservation of biological group structure, and accuracy of differentially expressed gene detection.
COBRA is freely available at https://github.com/wonlab-healthstat/COBRA. The code to reproduce the analyses is archived at Zenodo (https://doi.org/10.5281/zenodo.19891355).
Supplementary data are available at Bioinformatics online.
PMID:
42684056
Bibliographic data and abstract were imported from PubMed on 02 Sep 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 2
- Comments 0