Authors
Yuan, H., Huang, X., Auerbach, B., Linder, J., Srivastava, D., Kelley, D. R.
Abstract
Long-context sequence-to-function models enable nucleotide-resolution prediction of regulatory variant effects, motivating comprehensive mutational interrogation across broad genomic contexts. Yet conventional in silico saturation mutagenesis (ISM) scores mutations one at a time, requiring millions of model evaluations for a single gene and billions to trillions at genome scales. Here, we introduce Multi-ISM, a scalable framework that reformulates ISM as a sparse recovery problem. Multi-ISM generates mutational maps using 45-fold fewer model evaluations than exhaustive single-variant ISM and matches or exceeds its accuracy on variant-effect benchmarks. Multi-ISM is architecture-agnostic and transfers across large-scale sequence-to-function models. We applied Multi-ISM to 5,000 protein-coding genes, including 3,317 OMIM disease genes, generating base-pair-resolution, tissue-resolved attribution maps across 500-kb windows. These maps supported enhancer--gene prioritization and identification of cell-type-specific regulatory elements. Gene-level summaries of the maps captured regulatory complexity and showed that more constrained genes had smaller predicted mutational effects. Aggregating Multi-ISM predictions into gene-level rare-variant burdens improved personalized expression prediction over a common-variant elastic net, with the largest gains at expression outliers. Multi-ISM makes nucleotide-resolution interpretation of long-context sequence models a routine computation rather than a dedicated effort, so that new architectures, functional readouts, and cellular contexts can be mapped as they appear.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 02 Oct 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 13
- Comments 0