Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Not every gene is special: Modelling scale controls the false discovery rate when analysing high-throughput sequencing data.

Created on 09 Sep 2026

Authors

Scott J Dos Santos, Andreea C Murariu, Justin D Silverman, Gregory B Gloor

Published in

PLoS computational biology. Volume 22. Issue 9. Pages e1014728. Sep 08, 2026. Epub Sep 08, 2026.

Abstract

Differential expression/abundance analyses are commonplace in studies employing high-throughput sequencing (HTS); however different tools often fail to return comparable results when applied to the same dataset. Most tools employ normalisations to attempt to correct for technical variation in the count data. Previously, we demonstrated that these normalisations are often inappropriate due to incorrect assumptions regarding the overall scale (i.e., size) of the biological system in question. In this study, we used a combination of binomial thinning and permutation of sample groupings to produce 100 analysis iterations of 11 RNA-seq and other HTS datasets in which ~5% of all features are expected to be significantly different between groups. This enabled calculation of the false discovery rate (FDR) and sensitivity across the iterations. Our simulations showed that scale misspecification results in poor control of the FDR by several commonly used tools and that, counterintuitively, FDRs increased as the modelled difference between groups increased. Implementing a scale model in ALDEx2 or ALDEx3 ameliorated unacceptably high FDRs; however, there was an inherent trade-off between satisfactory FDR control and high sensitivity- no tool offered both. We established that increasing scale uncertainty also increased the minimum difference between groups required for a feature to be reported as differentially expressed. This phenomenon was consistently observed in disparate types of HTS data and was remarkably consistent. Critically, we leveraged a 'real-world', non-permuted analysis of an RNA-seq dataset to demonstrate that the latter effect is not a result of our thinning/permutation approach. Overall, our work highlights the potentially unwitting choice between sensitivity and FDR control that all researchers are making when analysing sequencing data and provides guidance on choosing an appropriate amount of scale uncertainty for the analysis of HTS data.

PMID:
42709908
Bibliographic data and abstract were imported from PubMed on 09 Sep 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 8
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement