Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

PRIDE: quality control in a proteomics data repository.

Attila Csordas | David Ovelleiro | Rui Wang | Joseph M Foster | Daniel Ríos | Juan Antonio Vizcaíno | Henning Hermjakob
Database : the journal of biological databases and curation | 2012

The PRoteomics IDEntifications (PRIDE) database is a large public proteomics data repository, containing over 270 million mass spectra (by November 2011). PRIDE is an archival database, providing the proteomics data supporting specific scientific publications in a computationally accessible manner. While PRIDE faces rapid increases in data deposition size as well as number of depositions, the major challenge is to ensure a high quality of data depositions in the context of highly diverse proteomics work flows and data representations. Here, we describe the PRIDE curation pipeline and its practical application in quality control of complex data depositions. DATABASE URL: http://www.ebi.ac.uk/pride/.

Pubmed ID: 22434838

Research resources used in this publication

None found

Antibodies used in this publication

None found

Associated grants

  • Agency: Wellcome Trust, United Kingdom
    Id: WT085949MA

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


HUPO Proteomics Standards Initiative (tool)

RRID:SCR_003158

Initiative to define community standards for data representation in proteomics to facilitate data comparison, exchange and verification. The main organizational unit is the work group, with a Gel Electrophoresis (GEL) work group, a Mass Spectrometry (MS) work group, a Molecular Interactions (MI) work group, a Protein Modifications (MOD) work group, a Proteomics Informatics (PI) work group, and a Sample Processing (SP) work group. The Gel Electrophoresis (GEL) work group aims to develop reporting requirements that supplement the Minimum Information About a Proteomics Experiment (MIAPE) parent document, describing the minimum information that should be reported about gel-based experimental techniques used in proteomics. The group will also develop data formats for capturing MIAPE-compliant data about gel electrophoresis and informatics performed on gel images. The Mass Spectrometry Standards Working Group defines community data formats and controlled vocabulary terms facilitating data exchange and archiving in the field of proteomics mass spectrometry. A past achievement is the mzData standard, which captures mass spectrometry output data. mzData's aim is to unite the large number of current formats (pkl's, dta's, mgf's, .....) into a single format. mzData has been released but is now deprecated in favor of mzML. The Molecular Interactions workgroup is concentrating on improving the annotation and representation of molecular interaction data wherever it is published, be this in journal articles, authors web-sites or public domain databases; and improving the accessibility of molecular interaction data to the user community. By using a common standard data can be downloaded from multiple sources and easily combined using a single parser. The protein modification workgroup focuses on developing a consensus nomenclature and provide an ontology reconciling in a hierarchical representation the complementary descriptions of residue modifications. The protein modification ontology (PSI-MOD) is available in OBO format or in OBO.xml. A spreadsheet containing the mapping of the descriptive labels used in various databases and search engines, the consensus list of proposed short name for protein modifications established by collaborative effort of mass spectrometry community, and the proposed rules and recommendations for this nomenclature are available. These short names are included in the ontology as synonyms of the corresponding terms. The Proteomics Informatics Standards Group (PSI-PI) goals are to provide a set of minimal reporting requirements which augment the MIAPE reporting guidelines with respect to analysis of data derived from proteomics experiments; to provide vendor-neutral and standard formats for representing results of analyzing and processing experimental data; to foster adoption of the format by highlighting efforts made by vendors and individuals that utilize the format in their products. The remit of the Sample Processing Working Group is to produce reporting guidelines, data exchange formats and controlled vocabulary covering all separation techniques not considered to be "classical" one- or two-dimensional gel electrophoresis (cf. the Gel WG home page), along with other kinds of sample handling and processing (for example, "tagging" proteins or peptides, splitting, combining and storing samples). Where possible we seek to develop our products in collaboration with all proteomics stakeholders and, where relevant, developers from other standards communities, most notably metabolomics. * Minimum reporting requirements: The evolving Minimum Information About a Proteomics Experiment (MIAPE) documents offer guidelines on how to adequately report a proteomics experiment. It is expected that these documents will be published, and that the requirements within will be enforced by journals, compliant repositories and funders (cf. MIAME). * XML formats for data exchange: Derived from the FuGE general object model, the formats developed by this workgroup are designed to function both as standalone files and as part of a "parent" FuGE-ML document. These formats will facilitate data exchange between researchers, and submission to repositories or journals. * Controlled vocabularies (CVs) and ontology: Lists of clearly defined terms are crucial for the construction of unambiguously worded data files. In addition to providing supporting CVs for the individual data capture formats as part of the integrated PSI CV, the Sample Processing WG will contribute terms to the Functional Genomics Ontology (FuGO).

View all literature mentions

Proteomics Identifications (PRIDE) (tool)

RRID:SCR_003411

Centralized, standards compliant, public data repository for proteomics data, including protein and peptide identifications, post-translational modifications and supporting spectral evidence. Originally it was developed to provide a common data exchange format and repository to support proteomics literature publications. This remit has grown with PRIDE, with the hope that PRIDE will provide a reference set of tissue-based identifications for use by the community. The future development of PRIDE has become closely linked to HUPO PSI. PRIDE encourages and welcomes direct user submissions of protein and peptide identification data to be published in peer-reviewed publications. Users may Browse public datasets, use PRIDE BioMart for custom queries, or download the data directly from the FTP site. PRIDE has been developed through a collaboration of the EMBL-EBI, Ghent University in Belgium, and the University of Manchester.

View all literature mentions

Proteome Commons Tranche repository (tool)

RRID:SCR_003441

A distributed file storage system that you can upload files to and download files from. All files uploaded to the repository are replicated several times to protect against their accidental loss. Files uploaded to the repository can be of any size, can be of any file type, and can be encrypted with a passphrase of your choosing. The Proteome Commons Tranche repository is the first instance of a Tranche repository. Tranche, was created so that anybody can take it and make their own Tranche repository. This is the first implementation of the Tranche software, and is useful as a test bed for the software. This repository relies on educational institutions to provide the hardware and facilities for Tranche servers. While we maintain a set of servers, the continued growth of this public resource will rely on the generosity of the institutions that use the repository most.

View all literature mentions

European Bioinformatics Institute (tool)

RRID:SCR_004727

Non-profit academic organization for research and services in bioinformatics. Provides freely available data from life science experiments, performs basic research in computational biology, and offers user training programme, manages databases of biological data including nucleic acid, protein sequences, and macromolecular structures. Part of EMBL.

View all literature mentions

Google Code (tool)

RRID:SCR_005786

Developer tools, APIs and resources. Search developers.google.com and code.google.com.

View all literature mentions