Hiring in life sciences? Share your open positions with our professional community. Read more Close

Advertisement

Using the DNA language model, GROVER, to parse effects of sequence, chromatin and regulatory features on genome stability.

Created on 12 Aug 2026

Authors

Pierre M Joubert, Anton Vlasov, Nikola Janakievski, Melissa Sanabria, Anna R Poetsch

Published in

Bioinformatics (Oxford, England). Aug 11, 2026. Epub Aug 11, 2026.

Abstract

Genome stability is shaped by DNA sequence and chromatin context, but their relative contributions to double-strand break (DSB) sensitivity remain unclear.
We show that the DNA language model, GROVER, can infer DSB location based on sequence. DSB hotspots tend to contain GC-rich sequences that belong to promoters, genes and short interspersed nuclear elements (SINEs). Additionally, we identified several specific short sequences (tokens) that are associated with modulating DSB sensitivity. Another model using chromatin and genome regulatory features outperforms the sequence-only model, highlighting complementary and cell-type specific information. Integrating sequence and genome biological features yields the best performance, demonstrating their synergy. Analyzing this model revealed that, dependent on the sample, genome stability information encoded in H3K36me3 and DNase-seq can be learned from the sequence, but not H3K27ac or H3K9me3. Embedding chromatin data directly into the GROVER architecture enabled cell-type specific modeling with performance matching the full chromatin feature model. Our results suggest that while chromatin and regulatory context provides important information, such as cell-type specificity, much of the information shaping DSB patterns is already encoded in the DNA sequence itself. Our integrative modeling approach not only reveals DSB patterns but also provides a generalizable strategy for tracing predictions in genomic data.
Data, models, and a tutorial are available on Zenodo.
doi:10.5281/zenodo.16375950.
doi:10.5281/zenodo.21131569.
Supplementary data are available at Bioinformatics online.

PMID:
42581590
Bibliographic data and abstract were imported from PubMed on 12 Aug 2026.

Read full publication at:
Please sign in to see all details.

Advertisement

Stats

  • Community rating n/a 0 votes
  • Reviewers' rating n/a 0 votes
  • Your rating

1-terrible, 9-excellent. How would you rate this publication? Sign in in to submit your rating.

  • Recommendations n/a n/a positive of 0 vote(s)
  • Views 10
  • Comments 0

Recommended by

  • No recommendations yet.

Post a comment

You need to be signed in to post comments. You can sign in here.

Comments

There are no comments yet.

Advertisement