Searching across hundreds of databases

Our searching services are busy right now. Your search will reload in five seconds.

X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

Tracking and coordinating an international curation effort for the CCDS Project.

Database : the journal of biological databases and curation | 2012

The Consensus Coding Sequence (CCDS) collaboration involves curators at multiple centers with a goal of producing a conservative set of high quality, protein-coding region annotations for the human and mouse reference genome assemblies. The CCDS data set reflects a 'gold standard' definition of best supported protein annotations, and corresponding genes, which pass a standard series of quality assurance checks and are supported by manual curation. This data set supports use of genome annotation information by human and mouse researchers for effective experimental design, analysis and interpretation. The CCDS project consists of analysis of automated whole-genome annotation builds to identify identical CDS annotations, quality assurance testing and manual curation support. Identical CDS annotations are tracked with a CCDS identifier (ID) and any future change to the annotated CDS structure must be agreed upon by the collaborating members. CCDS curation guidelines were developed to address some aspects of curation in order to improve initial annotation consistency and to reduce time spent in discussing proposed annotation updates. Here, we present the current status of the CCDS database and details on our procedures to track and coordinate our efforts. We also present the relevant background and reasoning behind the curation standards that we have developed for CCDS database treatment of transcripts that are nonsense-mediated decay (NMD) candidates, for transcripts containing upstream open reading frames, for identifying the most likely translation start codons and for the annotation of readthrough transcripts. Examples are provided to illustrate the application of these guidelines. DATABASE URL: http://www.ncbi.nlm.nih.gov/CCDS/CcdsBrowse.cgi.

Pubmed ID: 22434842 RIS Download

Research resources used in this publication

None found

Antibodies used in this publication

None found

Associated grants

  • Agency: NHGRI NIH HHS, United States
    Id: 5U54HG00455-04
  • Agency: Wellcome Trust, United Kingdom
    Id: WT077198
  • Agency: Intramural NIH HHS, United States
  • Agency: Wellcome Trust, United Kingdom
    Id: WT062023
  • Agency: Wellcome Trust, United Kingdom
    Id: 095908
  • Agency: NHGRI NIH HHS, United States
    Id: 5U54 HG004555

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


NCBI (tool)

RRID:SCR_006472

A portal to biomedical and genomic information. NCBI creates public databases, conducts research in computational biology, develops software tools for analyzing genome data, and disseminates biomedical information for the better understanding of molecular processes affecting human health and disease.

View all literature mentions

Genome Reference Consortium (tool)

RRID:SCR_006553

Consortium that puts sequences into a chromosome context and provides the best possible reference assembly for human, mouse, and zebrafish via FTP. Tools to facilitate the curation of genome assemblies based on the sequence overlaps of long, high quality sequences.

View all literature mentions

INSDC (tool)

RRID:SCR_011967

International collaboration of the International Nucleotide Sequence Databases (INSD), DDBJ, ENA, and GenBank, maintained for over 18 years. Individuals submitting data to the international sequence databases should be aware of INSDC policy.

View all literature mentions

GWAS: Catalog of Published Genome-Wide Association Studies (tool)

RRID:SCR_012745

Catalog of published genome-wide association studies. Genome-wide set of genetic variants in different individuals to see if any variant is associated with trait and disease. Database of genome-wide association study (GWAS) publications including only those attempting to assay single nucleotide polymorphisms (SNPs). Publications are organized from most to least recent date of publication. Studies are identified through weekly PubMed literature searches, daily NIH-distributed compilations of news and media reports, and occasional comparisons with an existing database of GWAS literature (HuGE Navigator). Works with HANCESTRO ancestry representation.

View all literature mentions