Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
A database of three-dimensional structural information about nucleic acids and their complexes. In addition to primary data, it contains derived geometric data, classifications of structures and motifs, standards for describing nucleic acid features, as well as tools and software for the analysis of nucleic acids. A variety of search capabilities are available, as are many different types of reports. NDB maintains the macromolecular Crystallographic Information File (mmCIF).
Proper citation: Nucleic Acid Database (RRID:SCR_003255) Copy
http://www.ncbi.nlm.nih.gov/taxonomy/
Database for a curated classification and nomenclature that contains the names of all organisms that are represented in the public sequence databases with at least one nucleotide or protein sequence. Data provided encompasses archaea, bacteria, eukaryota, viroids and viruses. The NCBI taxonomy database is not a primary source for taxonomic or phylogenetic information. Furthermore, the database does not follow a single taxonomic treatise but rather attempts to incorporate phylogenetic and taxonomic knowledge from a variety of sources, including the published literature, web-based databases, and the advice of sequence submitters and outside taxonomy experts. Consequently, the NCBI taxonomy database is not a phylogenetic or taxonomic authority and should not be cited as such.
Proper citation: NCBI Taxonomy (RRID:SCR_003256) Copy
A database of human mitochondrial genomes containing mtDNA sequences, polymorphic sites, and the ability to search for specific variants. It contains 1865 complete sequences and 839 coding region sequences.
Proper citation: mtDB - Human Mitochondrial Genome Database (RRID:SCR_002945) Copy
http://bioinfo.mbi.ucla.edu/ASAP/
THIS RESOURCE IS NO LONGER IN SERVICE, documented on 8/12/13. Database to access and mine alternative splicing information coming from genomics and proteomics based on genome-wide analyses of alternative splicing in human (30 793 alternative splice relationships found) from detailed alignment of expressed sequences onto the genomic sequence. ASAP provides precise gene exon-intron structure, alternative splicing, tissue specificity of alternative splice forms, and protein isoform sequences resulting from alternative splicing. They developed an automated method for discovering human tissue-specific regulation of alternative splicing through a genome-wide analysis of expressed sequence tags (ESTs), which involves classifying human EST libraries according to tissue categories and Bayesian statistical analysis. They use the UniGene clusters of human Expressed Sequence Tags (ESTs) to identify splices. The UniGene EST's are clustered so that a single cluster roughly corresponds to a gene (or at least a part of a gene). A single EST represents a portion of a processed (already spliced) mRNA. A given cluster contains many ESTs, each representing an outcome of a series of splicing events. The ESTs in UniGene contain the different mRNA isoforms transcribed from an alternatively spliced gene. They are not predicting alternative splicing, but locating it based on EST analysis. The discovered splices are further analyzed to determine alternative splicing events. They have identified 6201 alternative splice relationships in human genes, through a genome-wide analysis of expressed sequence tags (ESTs). Starting with 2.1 million human mRNA and EST sequences, they mapped expressed sequences onto the draft human genome sequence and only accepted splices that obeyed the standard splice site consensus. After constructing a tissue list of 46 human tissues with 2 million human ESTs, they generated a database of novel human alternative splices that is four times larger than our previous report, and used Bayesian statistics to compare the relative abundance of every pair of alternative splices in these tissues. Using several statistical criteria for tissue specificity, they have identified 667 tissue-specific alternative splicing relationships and analyzed their distribution in human tissues. They have validated our results by comparison with independent studies. This genome-wide analysis of tissue specificity of alternative splicing will provide a useful resource to study the tissue-specific functions of transcripts and the association of tissue-specific variants with human diseases.
Proper citation: ASAP: the Alternative Splicing Annotation Project (RRID:SCR_003415) Copy
Database of polymorphisms and mutations of the human mitochondrial DNA. It reports published and unpublished data on human mitochondrial DNA variation. All data is curated by hand. If you would like to submit published articles to be included in mitomap, please send them the citation and a pdf.
Proper citation: MITOMAP - A human mitochondrial genome database (RRID:SCR_002996) Copy
http://caps.ncbs.res.in/3dswap/index.html
Curated knowledegbase of protein structures that are reported to be involved in 3-dimensional domain swapping. 3DSwap provides literature curated information and structure related information about 3D domain swapping in proteins. Information about swapping, hinge region, swapped region, extent of swapping, etc. are extracted from original research publications after extensive literature curation.
Proper citation: 3DSwap (RRID:SCR_004133) Copy
http://www.hgsc.bcm.tmc.edu/content/hapmap-3-and-encode-3
Draft release 3 for genome-wide SNP genotyping and targeted sequencing in DNA samples from a variety of human populations (sometimes referred to as the HapMap 3 samples). This release contains the following data: * SNP genotype data generated from 1184 samples, collected using two platforms: the Illumina Human1M (by the Wellcome Trust Sanger Institute) and the Affymetrix SNP 6.0 (by the Broad Institute). Data from the two platforms have been merged for this release. * PCR-based resequencing data (by Baylor College of Medicine Human Genome Sequencing Center) across ten 100-kb regions (collectively referred to as ENCODE 3) in 712 samples. Since this is a draft release, please check this site regularly for updates and new releases. The HapMap 3 sample collection comprises 1,301 samples (including the original 270 samples used in Phase I and II of the International HapMap Project) from 11 populations, listed below alphabetically by their 3-letter labels. Five of the ten ENCODE 3 regions overlap with the HapMap-ENCODE regions; the other five are regions selected at random from the ENCODE target regions (excluding the 10 HapMap-ENCODE regions). All ENCODE 3 regions are 100-kb in size, and are centered within each respective ENCODE region. The HapMap 3 and ENCORE 3 data are downloadable from the ftp site.
Proper citation: HapMap 3 and ENCODE 3 (RRID:SCR_004563) Copy
A database of protein families, each represented by multiple sequence alignments and hidden Markov models (HMMs). Users can analyze protein sequences for Pfam matches, view Pfam family annotation and alignments, see groups of related families, look at the domain organization of a protein sequence, find the domains on a PDB structure, and query Pfam by keywords. There are two components to Pfam: Pfam-A and Pfam-B. Pfam-A entries are high quality, manually curated families that may automatically generate a supplement using the ADDA database. These automatically generated entries are called Pfam-B. Although of lower quality, Pfam-B families can be useful for identifying functionally conserved regions when no Pfam-A entries are found. Pfam also generates higher-level groupings of related families, known as clans (collections of Pfam-A entries which are related by similarity of sequence, structure or profile-HMM).
Proper citation: Pfam (RRID:SCR_004726) Copy
http://www.ncbi.nlm.nih.gov/mapview/map_search.cgi?taxid=7165
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on January 11, 2023. A database for the Anopheles gambiae str. PEST genome that was sequenced using a whole genome shotgun approach. The database aims to contribute to the understanding of mosquito genome structure and organization and will assist the development of malaria control strategies and improved anti-malarial drugs and vaccines. Sequences were generated and assembled into contigs for submission to GenBank.
Proper citation: Anopheles gambiae (African malaria mosquito) genome view (RRID:SCR_004402) Copy
http://www.uniprot.org/taxonomy/
NEWT is the taxonomy database maintained by the UniProt group. It integrates taxonomy data compiled in the NCBI database and data specific to the UniProt Knowledgebase. Browse by hierarchy, List all, or Complete proteomes. Organisms are classified in a hierarchical tree structure. Our taxonomy database contains every node (taxon) of the tree. UniProtKB taxonomy data is manually curated: next to manually verified organism names, we provide a selection of external links, organism strains and viral host information. Species with protein sequences stored in the UniProt Knowledgebase are named according to UniProt nomenclature. We endeavour to maintain a list of manually curated species names for which protein sequence data is available. In particular, we have adopted a systematic convention for naming viral and bacterial strains and isolates. Links to external sites are chosen by the UniProt taxonomy team and show pictures and various scientific data of interest (taxonomy, biology, physiology,...).
Proper citation: NEWT (RRID:SCR_004477) Copy
http://amphoranet.pitgroup.org/
Webserver implementation of the AMPHORA2 workflow for phylogenetic analysis of metagenomic shotgun sequencing data. It is capable of assigning a probability-weighted taxonomic group for each phylogenetic marker gene found in the input metagenomic sample.
Proper citation: AmphoraNet (RRID:SCR_005009) Copy
http://www.ebi.ac.uk/biosamples/
Database that aggregates sample information for reference samples (e.g. Coriell Cell lines) and samples for which data exist in one of the EBI''''s assay databases such as ArrayExpress, the European Nucleotide Archive or PRoteomics Identificates DatabasE. It provides links to assays for specific samples, and accepts direct submissions of sample information. The goals of the BioSample Database include: # recording and linking of sample information consistently within EBI databases such as ENA, ArrayExpress and PRIDE; # minimizing data entry efforts for EBI database submitters by enabling submitting sample descriptions once and referencing them later in data submissions to assay databases and # supporting cross database queries by sample characteristics. The database includes a growing set of reference samples, such as cell lines, which are repeatedly used in experiments and can be easily referenced from any database by their accession numbers. Accession numbers for the reference samples will be exchanged with a similar database at NCBI. The samples in the database can be queried by their attributes, such as sample types, disease names or sample providers. A simple tab-delimited format facilitates submissions of sample information to the database, initially via email to biosamples (at) ebi.ac.uk. Current data sources: * European Nucleotide Archive (424,811 samples) * PRIDE (17,001 samples) * ArrayExpress (1,187,884 samples) * ENCODE cell lines (119 samples) * CORIELL cell lines (27,002 samples) * Thousand Genome (2,628 samples) * HapMap (1,417 samples) * IMSR (248,660 samples)
Proper citation: BioSample Database at EBI (RRID:SCR_004856) Copy
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on January 11,2023. SuperCAT hosts typing databases for the Bacillus cereus group of bacteria. The databases contain MultiLocus Sequence Typing (MLST), MultiLocus Enzyme Electrophoresis (MLEE), and Amplified Fragment Length Polymorphism (AFLP) phylogenetic data. multilocus, sequence, Bacillus cereus, bacteria, Genomics, non-vertebrate, taxonomy, identification
Proper citation: SuperCAT (RRID:SCR_004882) Copy
A clade oriented, community curated database containing genomic, genetic, phenotypic and taxonomic information for plant genomes. Genomic information is presented in a comparative format and tied to important plant model species such as Arabidopsis. SGN provides tools such as: BLAST searches, the SolCyc biochemical pathways database, a CAPS experiment designer, an intron detection tool, an advanced Alignment Analyzer, and a browser for phylogenetic trees. The SGN code and database are developed as an open source project, and is based on database schemas developed by the GMOD project and SGN-specific extensions.
Proper citation: SGN (RRID:SCR_004933) Copy
http://cgi-www.daimi.au.dk/cgi-chili/datfap/frontdoor.py
A database of transcription factors from 13 plant species, and PCR primers for around 90% of them.
Proper citation: DATFAP (RRID:SCR_005413) Copy
A publicly available database of Transposed elements (TEs) which are located within protein-coding genes of 7 organisms: human, mouse, chicken, zebrafish, fruilt fly, nematode and sea squirt. Using TranspoGene the user can learn about the many aspects of the effect these TEs have on their hosting genes, such as: exonization events (including alternative splicing-related data), insertion of TEs into introns, exons, and promoters, specific location of the TE over the gene, evolutionary divergence of the TE from its consensus sequence and involvement in diseases. TranspoGene database is quickly searchable through its website, enables many kinds of searches and is available for download. TranspoGene contains information regarding specific type and family of the TEs, genomic and mRNA location, sequence, supporting transcript accession and alignment to the TE consensus sequence. The database also contains host gene specific data: gene name, genomic location, Swiss-Prot and RefSeq accessions, diseases associated with the gene and splicing pattern. The TranspoGene and microTranspoGene databases can be used by researchers interested in the effect of TE insertion on the eukaryotic transcriptome.
Proper citation: TranspoGene (RRID:SCR_005634) Copy
http://www.gene-regulation.com/pub/databases.html#transfac
Manually curated database of eukaryotic transcription factors, their genomic binding sites and DNA binding profiles. Used to predict potential transcription factor binding sites.
Proper citation: TRANSFAC (RRID:SCR_005620) Copy
http://edwardslab.bmcb.georgetown.edu/downloads/
The Peptide Sequence Database contains putative peptide sequences from human, mouse, rat, and zebrafish. Compressed to eliminate redundancy, these are about 40 fold smaller than a brute force enumeration. Current and old releases are available for download. Each species'' peptide sequence database comprises peptide sequence data from releveant species specific UniGene and IPI clusters, plus all sequences from their consituent EST, mRNA and protein sequence databases, namely RefSeq proteins and mRNAs, UniProt''s SwissProt and TrEMBL, GenBank mRNA, ESTs, and high-throughput cDNAs, HInv-DB, VEGA, EMBL, IPI protein sequences, plus the enumeration of all combinations of UniProt sequence variants, Met loss PTM, and signal peptide cleavages. The README file contains some information about the non amino-acid symbols O (digest site corresponding to a protein N- or C-terminus) and J (no digest sequence join) used in these peptide sequence databases and information about how to configure various search engines to use them. Some search engines handle (very) long sequences badly and in some cases must be patched to use these peptide sequence databases. All search engines supported by the PepArML meta-search engine can (or can be patched to) successfully search these peptide sequence databases.
Proper citation: Peptide Sequence Database (RRID:SCR_005764) Copy
http://indel.bioinfo.sdu.edu.cn/gridsphere/gridsphere
THIS RESOURCE IS NO LONGER IN SERVCE, documented September 2, 2016. Indel Flanking Region Database is an online resource for indels and the flanking regions of proteins in SCOP superfamilies, including amino acid sequences, lengths, locations, secondary structure constitutions, hydrophilicity / hydrophobicity, domain information, 3D structures and so on. It aims at providing a comprehensive dataset for analyzing the qualities of amino acid insertion/deletions(indels), substitutions and the relationship between them. The indels were obtained through the pairwise alignment of homologous structures in SCOP superfamilies. The IndelFR database contains 2,925,017 indels with flanking regions extracted from 373,402 structural alignment pairs of 12,573 non-redundant domains from 1053 superfamilies. IndelFR has already been used for molecular evolution studies and may help to promote future functional studies of indels and their flanking regions.
Proper citation: IndelFR - Indel Flanking Region Database (RRID:SCR_006050) Copy
http://www.ebi.ac.uk/thornton-srv/databases/FunTree/
FunTree provides a range of data resources to detect the evolution of enzyme function within distant structurally related clusters within domain super families as determined by CATH. To access the resource enter a specific CATH superfamily code or search for a structure / sequence / function (either via a EC code or KEGG ligand / reaction ID, PDB ID or UniProtKB ID). Or browse the resource via superfamily / function / structure / metabolites & reactions via the menu on the left panel. FunTree is a new resource that brings together sequence, structure, phylogenetic, chemical and mechanistic information for structurally defined enzyme superfamilies. Gathering together this range of data into a single resource allows the investigation of how novel enzyme functions have evolved within a structurally defined superfamily as well as providing a means to analyse trends across many superfamilies. This is done not only within the context of an enzyme''''s sequence and structure but also the relationships of their reactions. Developed in tandem with the CATH database, it currently comprises 276 superfamilies covering 1800 (70%) of sequence assigned enzyme reactions. Central to the resource are phylogenetic trees generated from structurally informed multiple sequence alignments using both domain structural alignments supplemented with domain sequences and whole sequence alignments based on commonality of multi-domain architectures. These trees are decorated with functional annotations such as metabolite similarity as well as annotations from manually curated resources such the catalytic site atlas and MACiE for enzyme mechanisms.
Proper citation: FunTree (RRID:SCR_006014) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the ASWG Resources search. From here you can search through a compilation of resources used by ASWG and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that ASWG has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on ASWG then you can log in from here to get additional features in ASWG such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into ASWG you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within ASWG that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.