Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://net.icgeb.org/benchmark/
It was created in order to create standard datasets on which the performance of machine learning methods can be compared. The collection contains datasets of sequences and structures, each subdivided into positive/negative training/test sets. Such a subdivision is called a classification task. Typical tasks include the classification of structural domains in the SCOP and CATH databases based on their sequences, as fell as various functional and taxonomic classification tasks. Running a performance evaluation test on an entire database can include many different classification tasks. These ensembles of classification tasks are encoded in a simple matrix format - called the cast matrix or membership table - that specifies the role of each sequence (or structure) in the different calculations. Each column of this matrix is a subdivision of the objects (rows) into positive/negative training/test sets. Typically, a database record contains such an ensemble of classification tasks, encoded in a single cast matrix. In addition, there is a collection of distance matrices that contain an all vs. all comparison of the datasets using methods as BLAST, Smith-Waterman, 3D-comparisons etc. Evaluation of a method on a given database consists of calculating a performance measure such as a receiver operating curve (ROC) AUC value. Results of evaluation are deposited along with the data, each dataset is evaluated at least by one classification method, such as 1NN (nearest neighbour) or SVM (support vector machines), ANN (artificial neural networks), RF (random forests) etc.. There are small datasets meant for program developers, as well as downloadable programs for various classification algorithms.
Proper citation: Protein Classification Benchmark Collection (RRID:SCR_007561) Copy
http://probeexplorer.cicancer.org/principal.php
Probe Explorer is an open access web-based bioinformatics application designed to show the association between microarray oligonucleotide probes and transcripts in the genomic context, but flexible enough to serve as a simplified genome and transcriptome browser. Coordinates and sequences of the genomic entities (loci, exons, transcripts), including vector graphics outputs, are provided for fifteen metazoa organisms and two yeasts. Alignment tools are used to built the associations between Affymetrix microarrays probe sequences and the transcriptomes (for human, mouse, rat and yeasts). Search by keywords is available and user searches and alignments on the genomes can also be done using any DNA or protein sequence query. Platform: Online tool
Proper citation: ProbeExplorer (RRID:SCR_007116) Copy
http://cmckb.cellmigration.org
It is a database of keys facts about proteins, families, and complexes involved in cell migration. This ongoing project provides a large amount of automated and curated data, collected from numerous online resources that are updated monthly. These data include names, synonyms, sequence information, summaries, CMC research data, reagents, structures, as well as protein family and complex details. CMKB''s ultimate goal is to create a database that will enable the cell migration community to conveniently access significant information about molecules of interest. This will also serve as a stepping stone to pathway analysis and demonstrate how these molecules coordinate with one another during cell adhesion and movement. Sponsors: This resource is supported by the Cell Migration Consortium.
Proper citation: CMKB (RRID:SCR_007229) Copy
http://www.broadinstitute.org/annotation/tetraodon/
This database have been funded by the National Human Genome Research Institute (NHGRI) to produce shotgun sequence of the Tetraodon nigriviridis genome. The strategy involves Whole Genome Shotgun (WGS) sequencing, in which sequence from the entire genome is generated. Whole genome shotgun libraries were prepared from Tetraodon genomic DNA obtained from the laboratory of Jean Weissenbach at Genoscope. Additional sequence data of approximately 2.5X coverage of Tetraodon has also been generated by Genoscope in plasmid and BAC end reads. Broad and Genoscope intend to pool their data and generate whole genome assemblies. Tetraodon nigroviridis is a freshwater pufferfish of the order Tetraodontiformes and lives in the rivers and estuaries of Indonesia, Malaysia and India. This species is 20-30 million years distant from Fugu rubripes, a marine pufferfish from the same family. The gene repertoire of T. nigroviridis is very similar to that of other vertebrates. However, its relatively small genome of 385 Mb is eight times more compact than that of human, mostly because intergenic and intronic sequences are reduced in size compared to other vertebrate genomes. These genome characteristics along with the large evolutionary distance between bony fish and mammals make Tetraodon a compact vertebrate reference genome - a powerful tool for comparative genetics and for quick and reliable identification of human genes.
Proper citation: Tetraodon nigroviridis Database (RRID:SCR_007123) Copy
Alternative splicing essentially increases the diversity of the transcriptome and has important implications for physiology, development and the genesis of diseases. This resource uses a different approach to investigate alternative splicing (instead of the conventional case-by case fashion) and integrates all transcripts derived from a gene into a single splicing graph. ASG is a database of splicing graphs for human genes, using transcript information from various major sources (Ensembl, RefSeq, STACK, TIGR and UniGene). Each transcript corresponds to a path in the graph, and alternative splicing is displayed by bifurcations. This representation preserves the relationships between different splicing variants and allows us to investigate systematically all possible putative transcripts. Web interface allows users to display the splicing graphs, to interactively assemble transcripts and to access their sequences as well as neighboring genomic regions. ASG also provide for each gene, an exhaustive pre-computed catalog of putative transcriptsin total more than 1.2 million sequences. It has found that ~65 of the investigated genes show evidence for alternative splicing, and in 5 of the cases, a single gene might produce over 100 transcripts.
Proper citation: Alternate splicing gallery (RRID:SCR_008129) Copy
http://pbil.univ-lyon1.fr/acuts/ACUTS.html
THIS RESOURCE IS NO LONGER IN SERVICE, Documented on August 12, 2014. Database that identifies new regulatory elements in untranslated regions of protein-coding genes (5 prime flanks, 5 prime UTRs, introns, 3 prime UTRs and 3 prime flanks). The analyses is focused on genes from metazoan species (essentially vertebrates, insects and nematodes). Information on highly conserved regions (sequences, alignments, annotations, bibliographic references) are compiled. Currently 176 out of 326 detected highly conserved regions (HCRs) have been analyzed and incorporated in the database. You can also access the list of annotated conserved elements and the list of conserved elements that remain to be processed. Their approach is based on comparative sequence analysis, for the identification of phylogenetic footprints.
Proper citation: Ancient conserved untranslated sequences (RRID:SCR_008130) Copy
http://animal.dna.affrc.go.jp/agp/index.html
Database of comparative gene mapping between species to assist the mapping of the genes related to phenotypic traits in livestock. The linkage maps, cytogenetic maps, polymerase chain reaction primers of pig, cattle, mouse and human, and their references have been included in the database, and the correspondence among species have been stipulated in the database. AGP is an animal genome database developed on a Unix workstation and maintained by a relational database management system. It is a joint project of National Institute of Agrobiological Sciences (NIAS) and Institute of the Society for Techno-innovation of Agriculture, Forestry and Fisheries (STAFF-Institute), under cooperation with other related research institutes. AGP also contains the Pig Expression Data Explorer (PEDE), a database of porcine EST collections derived from full-length cDNA libraries and full-length sequences of the cDNA clones picked from the EST collection. The EST sequences have been clustered and assembled, and their similarity to sequences in RefSeq, and UniGene determined. The PEDE database system was constructed to store sequences and similarity data of swine full-length cDNA libraries and to make them available to users. It provides interfaces for keyword and ID searches of BLAST results and enables users to obtain sequence data and names of clones of interest. Putative SNPs in EST assemblies have been classified according to breed specificity and their effect on coding amino acids, and the assemblies are equipped with an SNP search interface. The database contains porcine nucleotide sequences and cDNA clones that are ready for analyses such as expression in mammalian cells, because of their high likelihood of containing full-length CDS. PEDE will be useful for researchers who want to explore genes that may be responsible for traits such as disease susceptibility. The database also offers information regarding major and minor porcine-specific antigens, which might be investigated in regard to the use of pigs as models in various medical research applications.
Proper citation: Animal Genome Database (RRID:SCR_008165) Copy
Offer biorepository services to public and private research institutes, to the highest standards of quality and safety with the aim of contributing to the advancement of medical research and scientific discovery. The BioRep Cell Repository establishes, maintains and distributes cell line cultures as well as DNA derived from these cultures. The scientific and business affiliation between BioRep and Coriell allows access to more than a million types of cell vials, stored in liquid nitrogen. Cells that have been stored for nearly 50 years, are still viable and available for research purposes today. Thanks to an exclusive agreement with the Coriell Institute for Medical Research, the oldest and largest biorepository of the world, BioRep is specialized in cell lines preparation, in nucleic acid extraction and long term storage in liquid nitrose (-196 degrees C) and in refrigerators (-80 degrees C) of any kind of biosamples, using procedures and standards developed by the Coriell in over 50 years of activity. BioRep and Coriell together constitute one of the few Global Biorepository able to serve the pharmaceutical industries for world wide clinical trials. BioRep facility is specifically designed to give the utmost efficiency and security by implementing Coriell procedures and standards. The BioRep Tissue Repository provides safe and secure storage of tissue specimens as required for medical research and scientific investigation. All tissues are preserved with the most current preservation techniques and processes. In addition to the storage service, BioRep provides Cell Biology, Molecular Biology, Microbiology services developed in ISO 9001:2008 certified laboratories.
Proper citation: BioRep (RRID:SCR_004907) Copy
NIH initiative project to provide full-length open reading frame (FL-ORF) clones for human, mouse, and rat genes, cow. MGC cDNA clones were obtained by screening of cDNA libraries, by transcript-specific RT-PCR cloning, and by DNA synthesis of cDNA inserts. All MGC sequences are deposited in GenBank and clones can be purchased from distributors of IMAGE consortium. With conclusion of MGC project in March 2009, GenBank records of MGC sequences will be frozen, without further updates. Since definition of what constitutes full-length coding region for some of genes and transcripts for which they have MGC clones will likely change in future, users planning to order MGC clones will need to monitor for these changes. Users can make use of genome browsers and gene-specific databases, such as the UCSC Genome browser, NCBI's Map Viewer, and Entrez Gene, to view relevant regions of genome (browsers) or gene-related information (Entrez Gene).
Proper citation: Mammalian Gene Collection (RRID:SCR_007024) Copy
http://www.ebi.ac.uk/Tools/blast2/index.html
It is used to compare a novel sequence with those contained in nucleotide and protein databases by aligning the novel sequence with previously characterized genes.
Proper citation: Washington University Basic Local Alignment Search Tool (RRID:SCR_008285) Copy
http://www.uwstructuralgenomics.org/
It is a specialized research center supported by the Protein Structure Initiative (PSI) of the National Institute of General Medical Sciences (NIGMS), one of the National Institutes of Health (NIH). PSI is a federal, university, and industry effort aimed at dramatically reducing the costs and lessening the time it takes to determine a three-dimensional protein structure. The long-range goal of PSI is to solve 10,000 protein structures in 10 years and to make the three-dimensional atomic-level structures of most proteins easily obtainable from knowledge of their corresponding DNA sequences. CESG is located within the Department of Biochemistry at the University of Wisconsin-Madison (Madison, WI) and the Department of Biochemistry at the Medical College of Wisconsin (Milwaukee, WI). CESG develops new methods and technologies to address unique eukaryotic bottlenecks and disseminates its methodologies and experimental results to the scientific community worldwide through: :- Cell-Free Protein Production Workshops :- Plasmids at PSI Materials Repository :- Posters Presented at Scientific Meetings :- Publications in PubMed / PubMed Central :- Sesame (LIMS) Available for Researchers :- Solved Structures in the Protein Data Bank :- Technology Dissemination Reports They have welcomed requests by researchers to solve eukaryotic protein structures, particularly medically relevant proteins, through our Online Structure Request System for Researchers. They have solved many community-nominated targets and deposited information about these targets in public databases and published on our investigations and findings. Sponsors: CESG is supported by NIH / NIGMS Protein Structure Initiative grant numbers U54 GM074901 and P50 GM064598.
Proper citation: CESG (RRID:SCR_008451) Copy
Center for Computational Biology as a joint research center in the McKusick-Nathans Institute of Genetic Medicine, spanning the School of Medicine, the Whiting School of Engineering, the Bloomberg School of Public Health, and the Krieger School of Arts & Sciences. Multidisciplinary center dedicated to research on genomics, genetics, DNA sequencing technology, and computational methods for DNA and RNA sequence analysis.
Proper citation: Center for Computational Biology at JHU (RRID:SCR_016680) Copy
Ratings or validation data are available for this resource
Human and mouse genome annotation project which aims to identify all gene features in the human genome using computational analysis, manual annotation, and experimental validation.
Proper citation: GENCODE (RRID:SCR_014966) Copy
https://www.ncbi.nlm.nih.gov/genbank/tbl2asn2/
Software tool as a command-line program that automates the creation of sequence records for submission to GenBank. Records need no additional manual editing before submission.
Proper citation: tbl2asn (RRID:SCR_016636) Copy
https://github.com/asdcid/Gene-conservation-informed-contig-alignment
Software tool for separation haplotigs from genome assembly. Method to separate haplotigs based on sequence similarity.
Proper citation: Gene-conservation-informed-contig-alignment (RRID:SCR_017617) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on August 18,2025. Sequence analysis software for molecular biologists.
Proper citation: VectorFriends (RRID:SCR_001230) Copy
http://www.broad.mit.edu/mpr/lung
Data set of a molecular taxonomy of lung carcinoma, the leading cause of cancer death in the United States and worldwide. Using oligonucleotide microarrays, researchers analyzed mRNA expression levels corresponding to 12,600 transcript sequences in 186 lung tumor samples, including 139 adenocarcinomas resected from the lung. Hierarchical and probabilistic clustering of expression data defined distinct sub-classes of lung adenocarcinoma. Among these were tumors with high relative expression of neuroendocrine genes and of type II pneumocyte genes, respectively. Retrospective analysis revealed a less favorable outcome for the adenocarcinomas with neuroendocrine gene expression. The diagnostic potential of expression profiling is emphasized by its ability to discriminate primary lung adenocarcinomas from metastases of extra-pulmonary origin. These results suggest that integration of expression profile data with clinical parameters could aid in diagnosis of lung cancer patients.
Proper citation: Classification of Human Lung Carcinomas by mRNA Expression Profiling Reveals Distinct Adenocarcinoma Sub-classes (RRID:SCR_003010) Copy
http://www.uniprot.org/program/Chordata
Data set of manually annotated chordata-specific proteins as well as those that are widely conserved. The program keeps existing human entries up-to-date and broadens the manual annotation to other vertebrate species, especially model organisms, including great apes, cow, mouse, rat, chicken, zebrafish, as well as Xenopus laevis and Xenopus tropicalis. A draft of the complete human proteome is available in UniProtKB/Swiss-Prot and one of the current priorities of the Chordata protein annotation program is to improve the quality of human sequences provided. To this aim, they are updating sequences which show discrepancies with those predicted from the genome sequence. Dubious isoforms, sequences based on experimental artifacts and protein products derived from erroneous gene model predictions are also revisited. This work is in part done in collaboration with the Hinxton Sequence Forum (HSF), which allows active exchange between UniProt, HAVANA, Ensembl and HGNC groups, as well as with RefSeq database. UniProt is a member of the Consensus CDS project and thye are in the process of reviewing their records to support convergence towards a standard set of protein annotation. They also continuously update human entries with functional annotation, including novel structural, post-translational modification, interaction and enzymatic activity data. In order to identify candidates for re-annotation, they use, among others, information extraction tools such as the STRING database. In addition, they regularly add new sequence variants and maintain disease information. Indeed, this annotation program includes the Variation Annotation Program, the goal of which is to annotate all known human genetic diseases and disease-linked protein variants, as well as neutral polymorphisms.
Proper citation: UniProt Chordata protein annotation program (RRID:SCR_007071) Copy
http://www.tsl.ac.uk/groups/bioinformatics/
Core develops tools for high throughput sequence data to study non reference, non model organisms.
Proper citation: Sainsbury Laboratory Bioinformatics Core Facility (RRID:SCR_017185) Copy
https://sdrc.stanford.edu/sdrc-research-cores/dgac/home/
Core facility that offers library preparation and sequencing services on a variety of platforms - Illumina HiSeq 4000, MiSeq, HiSeq 2500 and PacBio Sequel - as well as bioinformatics analysis. It can sequence a variety of commercial sample preparation kits as well as custom workflows. DGAC provides access to high throughput sequencing and analysis to researchers at the Stanford Diabetes Research Center.
Proper citation: Stanford Diabetes Research Center Diabetes Genomics Analysis Core (RRID:SCR_016213) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the ASWG Resources search. From here you can search through a compilation of resources used by ASWG and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that ASWG has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on ASWG then you can log in from here to get additional features in ASWG such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into ASWG you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within ASWG that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.