Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Multiple evidence strands suggest that there may be as few as 19,000 human protein-coding genes.

Iakes Ezkurdia | David Juan | Jose Manuel Rodriguez | Adam Frankish | Mark Diekhans | Jennifer Harrow | Jesus Vazquez | Alfonso Valencia | Michael L Tress
Human molecular genetics | 2014

Determining the full complement of protein-coding genes is a key goal of genome annotation. The most powerful approach for confirming protein-coding potential is the detection of cellular protein expression through peptide mass spectrometry (MS) experiments. Here, we mapped peptides detected in seven large-scale proteomics studies to almost 60% of the protein-coding genes in the GENCODE annotation of the human genome. We found a strong relationship between detection in proteomics experiments and both gene family age and cross-species conservation. Most of the genes for which we detected peptides were highly conserved. We found peptides for >96% of genes that evolved before bilateria. At the opposite end of the scale, we identified almost no peptides for genes that have appeared since primates, for genes that did not have any protein-like features or for genes with poor cross-species conservation. These results motivated us to describe a set of 2001 potential non-coding genes based on features such as weak conservation, a lack of protein features, or ambiguous annotations from major databases, all of which correlated with low peptide detection across the seven experiments. We identified peptides for just 3% of these genes. We show that many of these genes behave more like non-coding genes than protein-coding genes and suggest that most are unlikely to code for proteins under normal circumstances. We believe that their inclusion in the human protein-coding gene catalogue should be revised as part of the ongoing human genome annotation effort.

Pubmed ID: 24939910

Research resources used in this publication

None found

Antibodies used in this publication

None found

Associated grants

  • Agency: NHGRI NIH HHS, United States
    Id: U41 HG007234

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


UniGene (tool)

RRID:SCR_004405

THIS RESOURCE IS NO LONGER IN SERVICE. Documented on January 11, 2023. Web tool for an organized view of the transcriptome. Collection of the computationally identified transcripts from the same locus. Information on protein similarities, gene expression, cDNA clones, and genomic location. System for automatically partitioning GenBank sequences into a non redundant set of gene oriented clusters.

View all literature mentions

PeptideAtlas (tool)

RRID:SCR_006783

Multi-organism, publicly accessible compendium of peptides identified in a large set of tandem mass spectrometry proteomics experiments. Mass spectrometer output files are collected for human, mouse, yeast, and several other organisms, and searched using the latest search engines and protein sequences. All results of sequence and spectral library searching are subsequently processed through the Trans Proteomic Pipeline to derive a probability of correct identification for all results in a uniform manner to insure a high quality database, along with false discovery rates at the whole atlas level. The raw data, search results, and full builds can be downloaded for other uses. All results of sequence searching are processed through PeptideProphet to derive a probability of correct identification for all results in a uniform manner ensuring a high quality database. All peptides are mapped to Ensembl and can be viewed as custom tracks on the Ensembl genome browser. The long term goal of the project is full annotation of eukaryotic genomes through a thorough validation of expressed proteins. The PeptideAtlas provides a method and a framework to accommodate proteome information coming from high-throughput proteomics technologies. The online database administers experimental data in the public domain. You are encouraged to contribute to the database.

View all literature mentions

GENCODE (tool)

RRID:SCR_014966

Human and mouse genome annotation project which aims to identify all gene features in the human genome using computational analysis, manual annotation, and experimental validation.

View all literature mentions

SignalP (tool)

RRID:SCR_015644

Web application for prediction of the presence and location of signal peptide cleavage sites in amino acid sequences from different organisms. The method incorporates a prediction of cleavage sites and a signal peptide/non-signal peptide prediction based on a combination of several artificial neural networks.

View all literature mentions

Conservation (tool)

RRID:SCR_016064

Software for scoring protein sequence conservation using the Jensen-Shannon divergence. It can be used to predict catalytic sites and residues near bound ligands.

View all literature mentions

Instituto Nacional de Bioinformtica (tool)

RRID:SCR_008586

INB is a technological plattform of Genoma Espaa. A National Network for coordination, integration and development of Spanish Bioinformatics Resources in genomics and proteomics projects. The INB (Instituto Nacional de Bioinformtica) is a community-driven technology platform of Genoma Espana founded in 2003 with the mission to consolidate Bioinformatics as a scientific discipline and to generate and apply solutions for the development and execution of projects related to genomics and proteomics, providing technical support in Bioinformatics to laboratories, institutions and companies throughout the Spanish territory.

View all literature mentions