Authors
Ye, W., Jiang, X., Shen, F.
Abstract
In high-throughput single-cell transcriptomics (p{approx} 20,000 genes), performing double machine learning (DML) causal inference on q{approx} 5,000 target genes requires nuisance function fits that grow linearly with the number of targets (K_{f} cross-fitting folds, K_{f}=5 or 10), far exceeding feasible computational budgets, especially with deep learning. We propose a Randomized Partition Strategy (RPS): randomly divide target genes into groups, share one background compression per group, reducing deep learning model training to q/m runs (m = group size) - a factor of m savings. The cost of grouping is accuracy loss - we prove that the cumulative deviation of the estimator follows a one-dimensional drift-free symmetric random walk, with diffusion variance growing linearly with group size and mean squared displacement equaling the mean squared error, so accuracy loss is predictable: m=1 is always optimal, accuracy cost is monotonically increasing, and a small accuracy sacrifice yields m-fold compute savings. On GSE189050 SLE single-cell data (Memory B cells, n=2120), both PCA and DL methods converge to the same conclusion, confirming the random walk mechanism is method-independent; an unexpected finding is that DL diffusion growth is only 16%, far slower than PCAs 7.4 times. This work provides a quantifiable theoretical foundation for compute strategy selection in single-cell high-dimensional causal inference.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 11 Aug 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 10
- Comments 0