Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://goblet.molgen.mpg.de/cgi-bin/goblet2008/goblet.cgi
Tool that performs annotation based on GO and pathway terms for anonymous cDNA or protein sequences. It uses the species independent GO structure and vocabulary together with a series of protein databases collected from various sites, to perform a detailed GO annotation by sequence similarity searches. The sensitivity and the reference protein sets can be selected by the user. GOblet runs automatically and is available as a public service on our web server. GOblet expects query sequences to be in FASTA-Format (with header-lines). Protein and nucleotide sequences are accepted. Total size of all sequences submitted per request should not be larger than 50kb currently. For security reasons: Larger post's will be rejected. Due to limited capacities the queries may be processed in batches depending on the server load. The output of the BLAST job is filtered automatically and the relevant hits are displayed. In addition, the respective GO-terms are shown together with the complete GO-hierarchy of parent terms., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: GOblet (RRID:SCR_006998) Copy
http://yetfasco.ccbr.utoronto.ca/
Collection of all available transcription factor (TF) specificities for the yeast Saccharomyces cerevisiae in Position Frequency Matrix (PFM) or Position Weight Matrix (PWM) formats. The specificities are evaluated for quality using several metrics. With this website, you can scan sequences with the motifs to find where potential binding sites lie, inspect precomputed genome-wide binding sites, find which TFs have similar motifs to one you have found, and download the collection of motifs. Submissions are welcome.
Proper citation: YeTFaSCo (RRID:SCR_006893) Copy
This service offers a gateway to well-benchmarked protein structure and function prediction methods. Structural models collected from the prediction servers are assessed using the powerful 3D-jury consensus approach. The Structure Prediction Meta Server provides access to various fold recognition, function prediction and local structure prediction methods. The Server takes the amino acid sequence of the query protein, the reference name for the prediction job, and the E-mail address as input. The E-mail address is used only for notification about errors during the execution of the job. The query sequence and the reference name are placed in the process queue. The Meta Server accepts only sequences, which have not been submitted before. In case of duplicate sequences the second user will be notified with a link to the previous submission. Sequences longer than 800 amino acids are not accepted by some services. The internal SQL database offers the possibility to find any previous jobs processed by the Meta Server using regular expressions addressing field like E-mail, Job Name and the host name, from which the job was initiated. Each server has its own process queuing system managed by the Meta Server. All results of fold recognition servers are translated into uniform formats. The information extracted from the raw output of the servers includes the PDB codes of the hits, the alignments and the similarity (reliability) scores specific for every server. Mapping of the hits to the SCOP and FSSP classifications are made either using known PDB representatives or alignment of the template sequence with the databases of proteins in both classifications. The secondary structure assignments for all hits are taken from the mapped FSSP (red for helices and blue for strands). Underscored amino acids indicate the first residue after an insertion in the template sequence. The Meta server provides translation of the alignments in standard formats like FASTA, PDB or CASP. The Meta Server is coupled to consensus servers. They provide jury predictions based on the results collected from other services. Not all fold recognition servers are used by the jury system. The data stored on the meta server is available through http://meta.bioinfo.pl/data/JOBID/. Jobs older than 2 months are not shown. The Meta Server is only a set of programs aimed to process and manage biological data, while the predictive power of the service comes from (mostly) remote prediction providers. Sponsors: This resource is supported by The BioInfoBank Institute.
Proper citation: BioInfoBank Meta Server (RRID:SCR_007181) Copy
Integrative database of germ-line V genes from the immunoglobulin loci of human and mouse. It presents V gene sequences extracted from the EMBL nucleotide sequence database and Ensembl together with links to the respective source sequences. Based on the properties of the source sequences, V genes are classified into 3 different classes: * Class 1: genomic and rearranged evidence * Class 2: genomic evidence only * Class 3: rearranged evidence only This allows careful sequence quality validation by the user. References to other immunological databases ( KABAT, IMGT/LIGM and VBASE ) are given to provide all public annotation data for each V gene. The VBASE2 database can be accessed either by the Direct Query interface or by the DNAPLOT Query interface. The Sequences given by the user are aligned with DNAPLOT against the VBASE2 database. Direct Query allows to enter sequence IDs and names (Field 1), choose species, locus, V gene family and class (Field 2) or search for 100% sequences (Field 3). At the DNAPLOT Query, the sequences given by the user are aligned with DNAPLOT against the VBASE2 database. The DNAPLOT program offers V gene nucleotide sequence alignment referring to the IMGT V gene unique numbering. The Quick Search can be used either for Direct Query to search for sequence IDs and V gene names or for DNAPLOT Query for up to 5 sequences. The new Fab Analysis allows you to align Fab, scFab, scAb or scFv sequences with DNAPLOT against the VBASE2 database, where both heavy and light chain are analyzed.
Proper citation: VBASE2 (RRID:SCR_007082) Copy
http://research-pub.gene.com/gmap/
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 29, 2016. A software program for mapping and aligning cDNA sequences to a genome. The program maps and aligns a single sequence with minimal startup time and memory requirements, and provides fast batch processing of large sequence sets. The program generates accurate gene structures, even in the presence of substantial polymorphisms and sequence errors, without using probabilistic splice site models. Methodology underlying the program includes a minimal sampling strategy for genomic mapping, oligomer chaining for approximate alignment, sandwich DP for splice site detection, and microexon identification with statistical significance testing.
Proper citation: GMAP (RRID:SCR_008992) Copy
Web-based tool that allows users to view comparisons of genetic and physical maps. The package also includes tools for curating map data. (entry from Genetic Analysis Software)
Proper citation: CMAP (RRID:SCR_009034) Copy
http://www.evocontology.org/site/Main/EvocOntologyDotOrg
THIS RESOURCE IS NO LONGER IN SERVICE, documented May 10, 2017. A pilot effort that has developed a centralized, web-based biospecimen locator that presents biospecimens collected and stored at participating Arizona hospitals and biospecimen banks, which are available for acquisition and use by researchers. Researchers may use this site to browse, search and request biospecimens to use in qualified studies. The development of the ABL was guided by the Arizona Biospecimen Consortium (ABC), a consortium of hospitals and medical centers in the Phoenix area, and is now being piloted by this Consortium under the direction of ABRC. You may browse by type (cells, fluid, molecular, tissue) or disease. Common data elements decided by the ABC Standards Committee, based on data elements on the National Cancer Institute''s (NCI''s) Common Biorepository Model (CBM), are displayed. These describe the minimum set of data elements that the NCI determined were most important for a researcher to see about a biospecimen. The ABL currently does not display information on whether or not clinical data is available to accompany the biospecimens. However, a requester has the ability to solicit clinical data in the request. Once a request is approved, the biospecimen provider will contact the requester to discuss the request (and the requester''s questions) before finalizing the invoice and shipment. The ABL is available to the public to browse. In order to request biospecimens from the ABL, the researcher will be required to submit the requested required information. Upon submission of the information, shipment of the requested biospecimen(s) will be dependent on the scientific and institutional review approval. Account required. Registration is open to everyone., documented September 6, 2016. Set of orthogonal controlled vocabularies that unifies gene expression data by facilitating a link between the genome sequence and expression phenotype information. The system associates labelled target cDNAs for microarray experiments, or cDNA libraries and their associated transcripts with controlled terms in a set of hierarchical vocabularies. eVOC consists of four orthogonal controlled vocabularies suitable for describing the domains of human gene expression data including Anatomical System, Cell Type, Pathology and Developmental Stage. The four core eVOC ontologies provide an appropriate set of detailed human terms that describe the sample source of human experimental material such as cDNA and SAGE libraries. These expression terms are linked to libraries and transcripts allowing the assessment of tissue expression profiles, differential gene expression levels and the physical distribution of expression across the genome. Analysis is currently possible using EST and SAGE data, with microarray data being incorporated. The eVOC data is increasingly being accepted as a standard for describing gene expression and eVOC ontologies are integrated with the Ensembl EnsMart database, the Alternate Transcript Diversity Project and the UniProt Knowledgebase. Several groups are currently working to provide shared development of this resource such that it is of maximum use in unifying transcript expression information.
Proper citation: eVOC (RRID:SCR_010704) Copy
Software tool to identify known and novel miRNA genes in seven animal clades by analyzing sequenced RNAs. Used for discovering known and novel miRNAs from small RNA sequencing data.
Proper citation: miRDeep (RRID:SCR_010829) Copy
http://www.mbio.ncsu.edu/BioEdit/bioedit.html
Software tool as biological sequence alignment editor written for Windows 95/98/NT/2000/XP/7 and sequence analysis program. Provides sequence manipulation and analysis options and links to external analysis programs to view and manipulate sequences with simple point and click operations., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: BioEdit (RRID:SCR_007361) Copy
http://senselab.med.yale.edu/ordb/
Database of vertebrate olfactory receptors genes and proteins. It supports sequencing and analysis of these receptors by providing a comprehensive archive with search tools for this expanding family. The database also incorporates a broad range of chemosensory genes and proteins, including the taste papilla receptors (TPRs), vomeronasal organ receptors (VNRs), insect olfaction receptors (IORs), Caenorhabditis elegans chemosensory receptors (CeCRs), and fungal pheromone receptors (FPRs). ORDB currently houses chemosensory receptors for more than 50 organisms. ORDB contains public and private sections which provide tools for investigators to analyze the functions of these very large gene families of G protein-coupled receptors. It also provides links to a local cluster of databases of related information in SenseLab, and to other relevant databases worldwide. The database aims to house all of the known olfactory receptor and chemoreceptor sequences in both nucleotide and amino acid form and serves four main purposes: * It is a repository of olfactory receptor sequences. * It provides tools for sequence analysis. * It supports similarity searches (screens) which reduces duplicate work. * It provides links to other types of receptor information, e.g. 3D models. The database is accessible to two classes of users: * General public www users have full access to all the public sequences, models and resources in the database. * Source laboratories are the laboratories that clone olfactory receptors and submit sequences in the private or public database. They can search any sequence they deposited to the database against any private or public sequence in the database. This user level is suited for laboratories that are actively cloning olfactory receptors.
Proper citation: Olfactory Receptor DataBase (RRID:SCR_007830) Copy
http://gene3d.biochem.ucl.ac.uk/Gene3D/
A large database of CATH protein domain assignments for ENSEMBL genomes and Uniprot sequences. Gene3D is a resource of form studying proteins and the component domains. Gene3D takes CATH domains from Protein Databank (PDB) structures and assigns them to the millions of protein sequences with no PDB structures using Hidden Markov models. Assigning a CATH superfamily to a region of a protein sequence gives information on the gross 3D structure of that region of the protein. CATH superfamilies have a limited set of functions and so the domain assignment provides some functional insights. Furthermore most proteins have several different domains in a specific order, so looking for proteins with a similar domain organization provides further functional insights. Strict confidence cut-offs are used to ensure the reliability of the domain assignments. Gene3D imports functional information from sources such as UNIPROT, and KEGG. They also import experimental datasets on request to help researchers integrate there data with the corpus of the literature. The website allows users to view descriptions for both single proteins and genes and large protein sets, such as superfamilies or genomes. Subsets can then be selected for detailed investigation or associated functions and interactions can be used to expand explorations to new proteins. The Gene3D web services provide programmatic access to the CATH-Gene3D annotation resources and in-house software tools. These services include Gene3DScan for identifying structural domains within protein sequences, access to pre-calculated annotations for the major sequence databases, and linked functional annotation from UniProt, GO and KEGG., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: Gene3D (RRID:SCR_007672) Copy
http://organelledb.lsi.umich.edu/
Database of organelle proteins, and subcellular structures / complexes from compiled protein localization data from organisms spanning the eukaryotic kingdom. All data may be downloaded as a tab-delimited text file and new localization data (and localization images, etc) for any organism relevant to the data sets currently contained in Organelle DB is welcomed. The data sets in Organelle DB encompass 138 organisms with emphasis on the major model systems: S. cerevisiae, A. thaliana, D. melanogaster, C. elegans, M. musculus, and human proteins as well. In particular, Organelle DB is a central repository of yeast protein localization data, incorporating results from both previous and current (ongoing) large-scale studies of protein localization in Saccharomyces cerevisiae. In addition, we have manually curated several recent subcellular proteomic studies for incorporation in Organelle DB. In total, Organelle DB is a singular resource consolidating our knowledge of the protein composition of eukaryotic organelles and subcellular structures. When available, we have included terms from the Gene Ontologies: the cellular component, molecular function, and biological process fields are discussed more fully in GO. Additionally, when available, we have included fluorescent micrographs (principally of yeast cells) visualizing the described protein localization. Organelle View is a visualization tool for yeast protein localization. It is a visually engaging way for high school and undergraduate students to learn about genetics or for visually-inclined researchers to explore Organelle DB. By revealing the data through a colorful, dimensional model, we believe that different kinds of information will come to light.
Proper citation: Organelle DB (RRID:SCR_007837) Copy
It provides information on natural and artificial mutants, including random and site-directed ones, for all proteins except members of the globin and immunoglobulin families. The PMD is based on literature, and each entry in the database corresponds to one article which may describe one, several or a number of protein mutants. Each database entry is identified by a serial number and is defined as either natural or artificial, depending on the type of the mutation. For each entry the following are recorded : JOURNAL, TITLE, CROSS-REFERENCE, PROTEIN, N-TERMINAL, CHANGE, FUNCTION, STRUCTURE, STABILITY, etc. CROSS-REFERENCE indicates the code names of the protein given in other databases such as Protein Identification Resources (2). N-TERMINAL shows the N-terminal sequence of five amino acids which may help to show the unambiguous numbering of th e sequence. CHANGE indicates the position and kind of mutations, such as amino acid substitution, insertion and deletion, denoted with a specific notation. Any functional or structural features (FUNCTION, STRUCTURE, STABILITY,etc) observed in the mutant are described immediately after ''CHANGE''. Relative differences in activity and/or stability, in comparison with the wild-type protein, are indicated with symbols (- -),(-),(=),(+) or (+ +). Complete loss of activity is denoted as (0). Data Submission A data submission system was newly prepared in the PMD. We welcome the authors of articles published in academic journals to submit their own mutant data to the PMD. After checking the contents, we will register the data with a unique accession number.
Proper citation: Protein Mutant Database (RRID:SCR_007878) Copy
Database of information about restriction enzymes and related proteins containing published and unpublished references, recognition and cleavage sites, isoschizomers, commercial availability, methylation sensitivity, crystal, genome, and sequence data. DNA methyltransferases, homing endonucleases, nicking enzymes, specificity subunits and control proteins are also included. Several tools are available including REBsites, BLAST against REBASE, NEBcutter and REBpredictor. Putative DNA methyltransferases and restriction enzymes, as predicted from analysis of genomic sequences, are also listed. REBASE is updated daily and is constantly expanding. Users may submit new enzyme and/or sequence information, recommend references, or send them corrections to existing data. The contents of REBASE may be browsed from the web and selected compilations can be downloaded by ftp (ftp.neb.com). Additionally, monthly updates can be requested via email.,
Proper citation: REBASE (RRID:SCR_007886) Copy
http://locus.jouy.inra.fr/cgi-bin/bovmap/intro.pl
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 22, 2016. Database containing information on the cattle genome comprising loci list, phenes list, homology query, cattle maps, gene list, and chromosome homology. The objective of BovMap is to develop a set of anchored loci for the cattle genome map. In total, 58 clones were hybridized with chromosomes and identified loci on 22 of the 31 different bovine chromosomes. Three clones contained satellite DNA. Two or more markers were placed on 12 chromosomes. Sequencing of the microsatellites and flanking regions was performed directly from 43 cosmids, as previously reported. Primers were developed for 39 markers and used to describe the polymorphism associated with the corresponding loci. Users are also allowed to summit their own data for Bovmap. An integrated cytogenetic and meiotic map of the bovine genome has also been developed around the Bovmap database. One objective that Bovmap uses as the mapping strategy for the bovine genome uses large insert clones as a tool for physical mapping and as a source of highly polymorphic microsatellites for genetic typing.
Proper citation: BovMap Database (RRID:SCR_008145) Copy
The E. coli Genome Project has the goal of completely sequencing the E. coli and human genomes. They began isolation of an overlapping lambda clonebank of E. coli K-12 strain MG1655. Those clones served as the starting material in our initial efforts to sequence the whole genome. Improvements in sequencing technology have since reached the point where whole-genome sequencing of microbial genomes is routine, and the human genome has in fact been completed. They initiated additional sequencing efforts, concentrating on pathogenic members of the family Enterobacteriaceae -- to which E. coli belongs. They also began a systematic functional characterization of E. coli K-12 genes and their regulation, using the whole genome sequence to address how the over 4000 genes of this organism act together to enable its survival in a wide range of environments.
Proper citation: E. coli Genome project (RRID:SCR_008139) Copy
MitoRes, is a comprehensive and reliable resource for massive extraction of sequences and sub-sequences of nuclear genes and encoded products targeting mitochondria in metazoa. It has been developed for supporting high-throughput in-silico analyses aimed to studies of functional genomics related to mitochondrial biogenesis, metabolism and to their pathological dysfunctions. It integrates information from the most accredited world-wide databases to bring together gene, transcript and encoded protein sequences associated to annotations on species name and taxonomic classification, gene name, functional product, organelle localization, protein tissue specificity, Enzyme Classification (EC), Gene Ontology (GO) classification and links to other related public databases. The section Cluster, has been dedicated to the collection of data on protein clustering of the entire catalogue of MitoRes protein sequences based on all versus all global pair-wise alignments for assessing putative intra- and inter-species functional relationships. The current version of MitoRes is based on the UniProt release 4 and contains 64 different metazoan species. The incredible explosion of knowledge production in Biology in the past two decades has created a critical need for bioinformatic instruments able to manage data and facilitate their retrieval and analysis. Hundreds of biological databases have been produced and the integration of biological data from these different resources is very important when we want to focus our efforts towards the study of a particular layer of biological knowledge. MitoRes is a completely rebuilt edition of MitoNuc database, which has been extensively modified to deal successfully with the challenges of the post genomic era. Its goal is to represent a comprehensive and reliable resource supporting high-quality in-silico analyses aimed to the functional characterization of gene, transcript and amino acid sequences, encoded by the nuclear genome and involved in mitochondrial biogenesis, metabolism and pathological dysfunctions in metazoa. The central features of MitoRes are: # an integrated catalogue of protein, transcript and gene sequences and sub-sequences # a Web-based application composed of a wide spectrum of search/retrieval facilities # a sequence export manager allowing massive extraction of bio-sequences (genes, introns, exons, gene flanking regions, transcripts, UTRs, CDS, proteins and signal peptides) in FASTA, EMBL and GenBank formats. It is an interconnected knowledge management system based on a MySQL relational database, which ensures data consistency and integrity, and on a Web Graphical User Interface (GUI), built in Seagull PHP Framework, offering a wide range of search and sequence extraction facilities. The database is compiled extracting and integrating information from public resources and data generated by the MitoRes team. The MitoRes database consists of comprehensive sequence entries whose core data are protein, transcript and gene sequences and taxonomic information describing the biological source of the protein. Additional information include: bio-sequences structure and location, biological function of protein product and dynamic links to both, external public databases used as data resources and public databases reporting complementary information. The core entity of the MitoRes database is represented by the protein so that each MitoRes entry is generated for each protein reported in the UniProt database as a nuclear encoded protein involved in mitochondrial biogenesis and function. Sponsors: MitoRes has been supported by Ministero Universit e Ricerca Scientifica, Italy (PRIN, Programma Biotecnologie legge 95/95-MURST 5, Proiect MURST Cluster C03/2000, CEGBA). Currently it is supported by operating grants from the Ministero dellIstruzione, dellUniversit e della Ricerca (MIUR), Italy (PNR 2001-2003 (FIRB art.8) D.M. 199, Strategic Program: Post-genome, grant 31-063933 and Project n.2, Cluster C03 L. 488/929).
Proper citation: MitoRes (RRID:SCR_008208) Copy
http://www.ebi.ac.uk/parasites/parasite-genome.html
This website contains information about the genomic sequence of parasites. It also contains multiple search engines to search six frame translations of parasite nucleotide databases for motifs, parasite protein databases for motifs, and parasite protein databases for keywords and text terms. * Guide to Internet Access to Parasite Genome Information * Guide to web-based analysis tools * Parasite Genome BLAST Server: Search a range of parasite specific nucleotide sequence databases with your own sequence. * Parasite Proteome Keyword Search Facility: Search parasite protein databases for keywords and text terms * Parasite Proteome Motif Search Facility: Search parasite protein databases for motifs * Parasite Six Frame Translation Motif Search Facility: Search six frame translations of parasite nucleotide databases for motifs * Genome computing resources: A list of ftp and gopher sites where genome computing applications and other resources can be found.
Proper citation: Parasite genome databases and genome research resources (RRID:SCR_008150) Copy
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 23, 2016. PDBfun is a web server for structural and functional analysis of proteins at the residue level. pdbFun gives fast access to the whole Protein Data Bank (PDB) organized as a database of annotated residues. The available data (features) range from solvent exposure to ligand binding ability, location in a protein cavity, secondary structure, residue type, sequence functional pattern, protein domain and catalytic activity. PDBfun is an integrated web tool for querying the PDB at the residue level and for local structural comparison. It integrates knowledge on single residues in protein structures coming from other databases or calculated with available or in-house developed instruments for structural analysis. Each set of different annotations represents a feature. Features are listed in PDBfun main page in orange. Features can be used for building residues selections.
Proper citation: Protein Databank Fun (RRID:SCR_008226) Copy
http://mafft.cbrc.jp/alignment/server/
Software package as multiple alignment program for amino acid or nucleotide sequences. Can align up to 500 sequences or maximum file size of 1 MB. First version of MAFFT used algorithm based on progressive alignment, in which sequences were clustered with help of Fast Fourier Transform. Subsequent versions have added other algorithms and modes of operation, including options for faster alignment of large numbers of sequences, higher accuracy alignments, alignment of non-coding RNA sequences, and addition of new sequences to existing alignments.
Proper citation: MAFFT (RRID:SCR_011811) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the SPARC SAWG Resources search. From here you can search through a compilation of resources used by SPARC SAWG and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that SPARC SAWG has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on SPARC SAWG then you can log in from here to get additional features in SPARC SAWG such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into SPARC SAWG you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within SPARC SAWG that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.