Searching the RRID Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes
Resource Name
PDF Report How to cite
Protein Classification Benchmark Collection (RRID:SCR_007561)
Copy Citation Copied
Resource Information

URL: http://net.icgeb.org/benchmark/

Proper Citation: Protein Classification Benchmark Collection (RRID:SCR_007561)

Description: It was created in order to create standard datasets on which the performance of machine learning methods can be compared. The collection contains datasets of sequences and structures, each subdivided into positive/negative training/test sets. Such a subdivision is called a classification task. Typical tasks include the classification of structural domains in the SCOP and CATH databases based on their sequences, as fell as various functional and taxonomic classification tasks. Running a performance evaluation test on an entire database can include many different classification tasks. These ensembles of classification tasks are encoded in a simple matrix format - called the cast matrix or membership table - that specifies the role of each sequence (or structure) in the different calculations. Each column of this matrix is a subdivision of the objects (rows) into positive/negative training/test sets. Typically, a database record contains such an ensemble of classification tasks, encoded in a single cast matrix. In addition, there is a collection of distance matrices that contain an all vs. all comparison of the datasets using methods as BLAST, Smith-Waterman, 3D-comparisons etc. Evaluation of a method on a given database consists of calculating a performance measure such as a receiver operating curve (ROC) AUC value. Results of evaluation are deposited along with the data, each dataset is evaluated at least by one classification method, such as 1NN (nearest neighbour) or SVM (support vector machines), ANN (artificial neural networks), RF (random forests) etc.. There are small datasets meant for program developers, as well as downloadable programs for various classification algorithms.

Synonyms: Benchmark

Resource Type: data or information resource, database

Keywords: classificaiton, machine learning, sequence, standard, standard dataset, structure, technology

Expand All
Usage and Citation Metrics

We found {{ ctrl2.mentions.all_count }} mentions in open access literature.

We have not found any literature mentions for this resource.

We are searching literature mentions for this resource.

Most recent articles:

{{ mention._source.dc.creators[0].familyName }} {{ mention._source.dc.creators[0].initials }}, et al. ({{ mention._source.dc.publicationYear }}) {{ mention._source.dc.title }} {{ mention._source.dc.publishers[0].name }}, {{ mention._source.dc.publishers[0].volume }}({{ mention._source.dc.publishers[0].issue }}), {{ mention._source.dc.publishers[0].pagination }}. (PMID:{{ mention._id.replace('PMID:', '') }})

Checkfor all resource mentions.

Collaborator Network

A list of researchers who have used the resource and an author search tool

Find mentions based on location


{{ ctrl2.mentions.errors.location }}

A list of researchers who have used the resource and an author search tool. This is available for resources that have literature mentions.

Ratings and Alerts

No rating or validation information has been found for Protein Classification Benchmark Collection.

No alerts have been found for Protein Classification Benchmark Collection.

Data and Source Information