Authors
Kraft, L., Sackett, P. W., Renaud, G.
Abstract
DNA sequences derived from ancient samples provide insights into human history, paleoenvironments, and evolutionary biology. Advances in laboratory techniques and computational tools have established ancient DNA research as a distinct field. However, current analyses focus mainly on the DNA level, while the protein space remains under-explored. Recent progress in the de novo assembly of ancient metagenomes and the availability of protein structure prediction tools, such as AlphaFold 2, enable the reconstruction of protein structures from these degraded sequences. Here, we present a computational framework to assemble contigs, evaluate their authenticity as ancient sequences, predict open reading frames, and fold ancient protein structures directly from highly damaged metagenomic data. Applying this pipeline to two-million-year-old datasets from the Kap Kobenhavn Formation, we successfully rescued ancient proteins involved in methane metabolism. By generating structural models with AlphaFold 2 and comparing them to modern predicted reference structures, we demonstrate that these ancient proteins can be reconstructed and aligned with high confidence. We showcase this by analyzing an archaeal V/A-type ATP synthase protein recovered from the 2M-year-old Greenlandic data. Ultimately, our work proves that ancient proteins can be reliably recovered from highly degraded palaeogenomic material, establishing a new computational avenue for evolutionary and biochemical research.
Preprint server:
bioRxiv
The authors list and abstract were imported from bioRxiv on 23 Aug 2026.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 10
- Comments 0