Authors
Feiquan Wang, Noor-Ul Áin, Fang Wang, Shengcheng Zhang, Weilong Kong, Yutao Shi, Hua Feng, Bo Zhang, Xingtan Zhang
Published in
GigaScience. Aug 21, 2026. Epub Aug 21, 2026.
Abstract
Tea (Camellia sinensis L.), a major global economic crop in Asia, poses challenges for genetic identification because its highly heterozygous, repetitive genome reduces the efficacy of conventional single-nucleotide polymorphism (SNP) and microsatellite markers, and interspecific hybridization further complicates the situation. To address these issues, CamK-DB was developed as a reference-free Camellia fingerprinting database built on MIKE MinHash sketches. We curated 418 candidate resequencing datasets, and built a database using standardized 5× genome-coverage fingerprints. Each accession is stored as a MIKE .jac fingerprint generated with k=21 and recommended sketch/pre_cnt=2000. CamK-DB provides a command-line interface for data management and a custom C++ query engine that computes top-10 matches using Jaccard similarity, complemented by a QT-based graphical interface for interactive analysis. This resource offers a robust and scalable framework for precise and routine germplasm identification, genomic phylogenetic inference, and strategic breeding program design. CamK-DB (database and code) is publicly available at (https://github.com/sc-zhang/CamK-DB). CamK-DB binaries are provided for Windows 10/11 and Linux (x86_64, glibc ≥ 2.27).
PMID:
42627355
Bibliographic data and abstract were imported from PubMed on 21 Aug 2026.
Read full publication at:
Please sign in
to see all details.
Advertisement
Stats
- Recommendations n/a n/a positive of 0 vote(s)
- Views 6
- Comments 0