Projects per year
Abstract
genomeRxiv is a newly-funded US-UK collaboration to provide a public, web-accessible database of public genome sequences, accurately catalogued and classified by whole-genome similarity independent of their taxonomic affiliation. Our goal is to supply the basic and applied research community with rapid, precise and accurate identification of unknown isolates based on genome sequence alone, and with molecular tools for environmental analysis.
The DNA sequencing revolution enabled the use of cultured and uncultured microorganism genomes for fast and precise identification. However, precise identification is impossible without
1. reference databases that precisely circumscribe classes of microorganisms, and label these with their uniquely-shared characteristics
2. fast algorithms that can handle the volumes of genome data
Our approach integrates the highly-resolved classification framework of Life Identification Numbers (LINs) with the speed and computational efficiency of sourmash and k-mer hashing algorithms, and the precision and filtering of average nucleotide identity (ANI). We aim to construct a single genome-based indexing scheme that extends from phylum to strain, enabling the unique and consistent placement of any sequenced prokaryote genome.
genomeRxiv includes protocols for confidentiality, allowing groups to identify and announce the identities of newly-sequenced organisms without sharing genome data directly. This protects communities working with commercially- and ethically-sensitive organisms (e.g. production engineering strains, potential bioweapons, and to enable benefit sharing with indigenous communities).
genomeRxiv will also provide online capability to design molecular diagnostic tools for metabarcoding and qPCR, to enable tracking of specific groupings of bacteria directly in the environment.
The DNA sequencing revolution enabled the use of cultured and uncultured microorganism genomes for fast and precise identification. However, precise identification is impossible without
1. reference databases that precisely circumscribe classes of microorganisms, and label these with their uniquely-shared characteristics
2. fast algorithms that can handle the volumes of genome data
Our approach integrates the highly-resolved classification framework of Life Identification Numbers (LINs) with the speed and computational efficiency of sourmash and k-mer hashing algorithms, and the precision and filtering of average nucleotide identity (ANI). We aim to construct a single genome-based indexing scheme that extends from phylum to strain, enabling the unique and consistent placement of any sequenced prokaryote genome.
genomeRxiv includes protocols for confidentiality, allowing groups to identify and announce the identities of newly-sequenced organisms without sharing genome data directly. This protects communities working with commercially- and ethically-sensitive organisms (e.g. production engineering strains, potential bioweapons, and to enable benefit sharing with indigenous communities).
genomeRxiv will also provide online capability to design molecular diagnostic tools for metabarcoding and qPCR, to enable tracking of specific groupings of bacteria directly in the environment.
| Original language | English |
|---|---|
| Number of pages | 1 |
| DOIs | |
| Publication status | Published - 6 Apr 2021 |
| Event | Microbiology Society Annual Conference 2021 - Online Duration: 26 Apr 2021 → 30 Apr 2021 https://microbiologysociety.org/event/annual-conference/annual-conference-online-2021.html |
Conference
| Conference | Microbiology Society Annual Conference 2021 |
|---|---|
| Period | 26/04/21 → 30/04/21 |
| Internet address |
Keywords
- genomeRxiv
- microbiology
- whole-genome database
- DNA sequencing
- microorganisms
- genomes
- qPCR
- metabarcoding
Fingerprint
Dive into the research topics of 'genomeRxiv: a microbial whole-genome database for classification, identification, and data sharing'. Together they form a unique fingerprint.Projects
- 1 Finished
-
BBSRC-NSF/BIO: Collaborative Proposal: genomeRxiv: a microbial whole-genome database and diagnostic marker design resource for classification, identification, and data sharing
Pritchard, L. (Principal Investigator)
BBSRC (Biotech & Biological Sciences Research Council)
8/03/21 → 28/03/25
Project: Research
-
pyani-plus: a multitool for average nucleotide identity estimation
Pritchard, L., Kiepas, A. & Cock, P., 13 Apr 2026. 1 p.Research output: Contribution to conference › Poster
Open AccessFile11 Downloads (Pure) -
Objective benchmarking and recommendations for Overall Genome Relatedness Index (OGRI) methods
Pritchard, L., 13 Jul 2025. 1 p.Research output: Contribution to conference › Poster
Open AccessFile8 Downloads (Pure) -
LINgroups as a robust principled approach to compare and integrate multiple bacterial taxonomies
Mazloom, R., Pierce-Ward, N. T., Sharma, P., Pritchard, L., Brown, C. T., Vinatzer, B. A. & Heath, L. S., Nov 2024, In: IEEE/ACM Transactions on Computational Biology and Bioinformatics . 21, 6, p. 2304-2314 11 p.Research output: Contribution to journal › Article › peer-review
Open AccessFile3 Link opens in a new tab Citations (Scopus)30 Downloads (Pure)
Activities
- 2 Media Participation
-
Bacterial Taxonomy: what is a species, what is a strain? - Part 2
Pritchard, L. (Interviewee)
9 Dec 2021Activity: Public Engagement and Outreach › Media Participation
-
Microbinfie podcast, episode 67 - Bacterial Taxonomy: what is a species, what is a strain?
Pritchard, L. (Interviewee)
25 Nov 2021Activity: Public Engagement and Outreach › Media Participation
Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver