Authors
Wang, Y., Shu, Z., Gao, F., Cao, Z.
Abstract
Selecting stochastic gene-expression models from single-cell counts requires balancing goodness of fit against unnecessary mechanistic complexity. The probability-generating-function-based Bayesian information criterion (PGF-BIC) combines covariance-weighted fitting in generating-function space with a complexity penalty, allowing candidate models to be compared without reconstructing their full count distributions. However, its conventional zero-threshold rule does not account for sampling uncertainty in the fitted score difference and may therefore favor overly complex models in finite samples. To address this limitation, we develop an uncertainty-aware PGF-BIC rule that selects the more complex model only when its score advantage exceeds a data-driven threshold. We use Cantelli's one-sided inequality to motivate a selection margin expressed in terms of a standard deviation. To determine this scale, we use influence functions to quantify sensitivity to small perturbations in the data distribution and obtain a first-order description of sampling fluctuations. The resulting variance estimate accounts for variability in both the empirical probability generating function and the estimated covariance weights, yielding a data-driven threshold for assessing the complex model's score advantage. A Poisson versus Bursty benchmark shows that the calibrated rule reduces incorrect selection of the more complex model. The calibration requires neither resampling nor additional optimization, incorporating sampling uncertainty into model selection while retaining the computational efficiency of PGF-BIC.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 18 Sep 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 11
- Comments 0