Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Standardized metadata for human pathogen/vector genomic sequences.

Vivien G Dugan | Scott J Emrich | Gloria I Giraldo-Calderón | Omar S Harb | Ruchi M Newman | Brett E Pickett | Lynn M Schriml | Timothy B Stockwell | Christian J Stoeckert | Dan E Sullivan | Indresh Singh | Doyle V Ward | Alison Yao | Jie Zheng | Tanya Barrett | Bruce Birren | Lauren Brinkac | Vincent M Bruno | Elizabet Caler | Sinéad Chapman | Frank H Collins | Christina A Cuomo | Valentina Di Francesco | Scott Durkin | Mark Eppinger | Michael Feldgarden | Claire Fraser | W Florian Fricke | Maria Giovanni | Matthew R Henn | Erin Hine | Julie Dunning Hotopp | Ilene Karsch-Mizrachi | Jessica C Kissinger | Eun Mi Lee | Punam Mathur | Emmanuel F Mongodin | Cheryl I Murphy | Garry Myers | Daniel E Neafsey | Karen E Nelson | William C Nierman | Julia Puzak | David Rasko | David S Roos | Lisa Sadzewicz | Joana C Silva | Bruno Sobral | R Burke Squires | Rick L Stevens | Luke Tallon | Herve Tettelin | David Wentworth | Owen White | Rebecca Will | Jennifer Wortman | Yun Zhang | Richard H Scheuermann
PloS one | 2014

High throughput sequencing has accelerated the determination of genome sequences for thousands of human infectious disease pathogens and dozens of their vectors. The scale and scope of these data are enabling genotype-phenotype association studies to identify genetic determinants of pathogen virulence and drug/insecticide resistance, and phylogenetic studies to track the origin and spread of disease outbreaks. To maximize the utility of genomic sequences for these purposes, it is essential that metadata about the pathogen/vector isolate characteristics be collected and made available in organized, clear, and consistent formats. Here we report the development of the GSCID/BRC Project and Sample Application Standard, developed by representatives of the Genome Sequencing Centers for Infectious Diseases (GSCIDs), the Bioinformatics Resource Centers (BRCs) for Infectious Diseases, and the U.S. National Institute of Allergy and Infectious Diseases (NIAID), part of the National Institutes of Health (NIH), informed by interactions with numerous collaborating scientists. It includes mapping to terms from other data standards initiatives, including the Genomic Standards Consortium's minimal information (MIxS) and NCBI's BioSample/BioProjects checklists and the Ontology for Biomedical Investigations (OBI). The standard includes data fields about characteristics of the organism or environmental source of the specimen, spatial-temporal information about the specimen isolation event, phenotypic characteristics of the pathogen/vector isolated, and project leadership and support. By modeling metadata fields into an ontology-based semantic framework and reusing existing ontologies and minimum information checklists, the application standard can be extended to support additional project-specific data fields and integrated with other data represented with comparable standards. The use of this metadata standard by all ongoing and future GSCID sequencing projects will provide a consistent representation of these data in the BRC resources and other repositories that leverage these data, allowing investigators to identify relevant genomic sequences and perform comparative genomics analyses that are both statistically meaningful and biologically relevant.

Pubmed ID: 24936976

Research resources used in this publication

None found

Antibodies used in this publication

None found

Associated grants

  • Agency: NIAID NIH HHS, United States
    Id: HHSN272200900018C
  • Agency: NIAID NIH HHS, United States
    Id: HHSN272200900038C
  • Agency: NIAID NIH HHS, United States
    Id: U19 AI110820
  • Agency: NIAID NIH HHS, United States
    Id: HHSN272200900040C
  • Agency: NIGMS NIH HHS, United States
    Id: R01GM093132
  • Agency: Intramural NIH HHS, United States
  • Agency: NIAID NIH HHS, United States
    Id: HHSN272200900039C
  • Agency: NIAID NIH HHS, United States
    Id: U19 AI110818
  • Agency: PHS HHS, United States
    Id: HHSN266200400041C
  • Agency: NIAID NIH HHS, United States
    Id: HHSN272200900041C
  • Agency: NIGMS NIH HHS, United States
    Id: R01 GM093132
  • Agency: NIAID NIH HHS, United States
    Id: HHSN272200900007C
  • Agency: NIAID NIH HHS, United States
    Id: HHSN272200900009C

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


GitHub (tool)

RRID:SCR_002630

A web-based hosting service for software development projects that use the Git revision control system offering powerful collaboration, code review, and code management. It offers both paid plans for private repositories, and free accounts for open source projects. Large or small, every repository comes with the same powerful tools. These tools are open to the community for public projects and secure for private projects. Features include: * Integrated issue tracking * Collaborative code review * Easily manage teams within organizations * Text entry with understated power * A growing list of programming languages and data formats * On the desktop and in your pocket - Android app and mobile web views let you keep track of your projects on the go.

View all literature mentions

Broad Institute (tool)

RRID:SCR_007073

Biomedical and genomic research center located in Cambridge, Massachusetts, United States. Nonprofit research organization under the name Broad Institute Inc., and is partners with Massachusetts Institute of Technology, Harvard University, and the five Harvard teaching hospitals. Dedicated to advance understanding of biology and treatment of human disease to improve human health.

View all literature mentions

Broad Terra cloud commons for pathogen surveillance (tool)

RRID:SCR_018278

Broad Terra cloud workspace for best practices with COVID-19 genomics data. Raw COVID-19 sequencing data from NCBI Sequence Read Archive. Workflows for genome assembly, quality control, metagenomic classification, and aggregate statistics.

View all literature mentions