Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://animal.dna.affrc.go.jp/agp/index.html
Database of comparative gene mapping between species to assist the mapping of the genes related to phenotypic traits in livestock. The linkage maps, cytogenetic maps, polymerase chain reaction primers of pig, cattle, mouse and human, and their references have been included in the database, and the correspondence among species have been stipulated in the database. AGP is an animal genome database developed on a Unix workstation and maintained by a relational database management system. It is a joint project of National Institute of Agrobiological Sciences (NIAS) and Institute of the Society for Techno-innovation of Agriculture, Forestry and Fisheries (STAFF-Institute), under cooperation with other related research institutes. AGP also contains the Pig Expression Data Explorer (PEDE), a database of porcine EST collections derived from full-length cDNA libraries and full-length sequences of the cDNA clones picked from the EST collection. The EST sequences have been clustered and assembled, and their similarity to sequences in RefSeq, and UniGene determined. The PEDE database system was constructed to store sequences and similarity data of swine full-length cDNA libraries and to make them available to users. It provides interfaces for keyword and ID searches of BLAST results and enables users to obtain sequence data and names of clones of interest. Putative SNPs in EST assemblies have been classified according to breed specificity and their effect on coding amino acids, and the assemblies are equipped with an SNP search interface. The database contains porcine nucleotide sequences and cDNA clones that are ready for analyses such as expression in mammalian cells, because of their high likelihood of containing full-length CDS. PEDE will be useful for researchers who want to explore genes that may be responsible for traits such as disease susceptibility. The database also offers information regarding major and minor porcine-specific antigens, which might be investigated in regard to the use of pigs as models in various medical research applications.
Proper citation: Animal Genome Database (RRID:SCR_008165) Copy
THIS RESOURCE IS NO LONGER IN SERVICE, documented on August 20,2019.The COG-database has become a powerful tool in the field of comparative genomics. The construction of this data-base is based on sequence homologies of proteins from different completely sequenced genomes. Highly homologous proteins are assigned to clusters of orthologous groups. The updated collection of orthologous protein sets for prokaryotes and eukaryotes is expected to be a useful platform for functional annotation of newly sequenced genomes, including those of complex eukaryotes, and genome-wide evolutionary studies. The availability of multiple, essentially complete genome sequences of prokaryotes and eukaryotes spurred both the demand and the opportunity for the construction of an evolutionary classification of genes from these genomes. Such a classification system based on orthologous relationships between genes appears to be a natural framework for comparative genomics and should facilitate both functional annotation of genomes and large-scale evolutionary studies. Here is a major update of the previously developed system for delineation of Clusters of Orthologous Groups of proteins (COGs) from the sequenced genomes of prokaryotes and unicellular eukaryotes and the construction of clusters of predicted orthologs for 7 eukaryotic genomes, which we named KOGs after eukaryotic orthologous groups. The COG collection currently consists of 138,458 proteins, which form 4873 COGs and comprise 75% of the 185,505 (predicted) proteins encoded in 66 genomes of unicellular organisms. The eukaryotic orthologous groups (KOGs) include proteins from 7 eukaryotic genomes: three animals (the nematode Caenorhabditis elegans, the fruit fly Drosophila melanogaster and Homo sapiens), one plant, Arabidopsis thaliana, two fungi (Saccharomyces cerevisiae and Schizosaccharomyces pombe), and the intracellular microsporidian parasite Encephalitozoon cuniculi. The current KOG set consists of 4852 clusters of orthologs, which include 59,838 proteins, or approximately 54% of the analyzed eukaryotic 110,655 gene products. Compared to the coverage of the prokaryotic genomes with COGs, a considerably smaller fraction of eukaryotic genes could be included into the KOGs; addition of new eukaryotic genomes is expected to result in substantial increase in the coverage of eukaryotic genomes with KOGs. Examination of the phyletic patterns of KOGs reveals a conserved core represented in all analyzed species and consisting of approximately 20% of the KOG set. This conserved portion of the KOG set is much greater than the ubiquitous portion of the COG set (approximately 1% of the COGs). In part, this difference is probably due to the small number of included eukaryotic genomes, but it could also reflect the relative compactness of eukaryotes as a clade and the greater evolutionary stability of eukaryotic genomes.
Proper citation: Phylogenetic Clusters of Orthologous Groups Ranking (RRID:SCR_008223) Copy
http://www.nisc.nih.gov/projects/comp_seq.html
Generates data for use in developing and refining computational tools for comparing genomic sequence from multiple species. The NISC Comparative Sequencing Program's goal is to establish a data resource consisting of sequences for the same set of targeted genomic regions derived from multiple animal species. The broader program includes plans for a diverse set of analytical studies using the generated sequence and the publication of a series of papers describing the results of those analysis in peer-reviewed journals in a timely fashion. Experimentally, this project involves the shotgun sequencing of mapped BAC clones. For each BAC, an assembly is first performed when a sufficient number of sequence reads have been generated to provide full shotgun coverage of the clone. At that time, the assembled sequence is submitted to the HTGS division of GenBank. Subsequent refinements of the sequence, including the generation of higher-accuracy finished sequence, results in the updating of the sequence record in GenBank. By immediately submitting our BAC-derived sequences to GenBank, it makes their data available as a public service to allow colleagues to speed up their research, consistent with the now well-established routine of sequencing centers participating in the Human Genome Project. However, at the same time, it has made considerable investment in acquiring these mapping and sequence data, including sizable efforts of graduate students, postdoctoral fellows, and other trainees. Furthermore, in most cases, large data sets involving multiple BAC sequences from multiple species must first be generated, often taking many months to accumulate, before the planned analysis can be performed and the resulting papers written and submitted for publication.
Proper citation: Comparative Vertebrate Sequencing (RRID:SCR_008213) Copy
http://genome.jgi.doe.gov/programs/metagenomes/index.jsf
Portal providing access to metagenomics projects, data and tools supported by the DOE Joint Genome Institute (JGI). A primary motivation for metagenomics is that most microbes found in nature exist in complex, interdependent communities and cannot readily be grown in isolation in the laboratory. One can, however, isolate DNA or RNA from the community as a whole, and studies of such communities have revealed a diversity of microbes far beyond those found in culture collections. It is suspected that these uncultivated organisms must harbor considerable as-yet undiscovered genomic, functional, and metabolic features and capabilities. Thus to fully explore microbial genomics, it is imperative that we access the genomes of these elusive players.
Proper citation: Metagenomics Program at JGI (RRID:SCR_008804) Copy
http://fungi.ensembl.org/index.html
The Ensembl Genomes project produces genome databases for important species from across the taxonomic range, using the Ensembl software system. Five sites are now available, one of which is Ensembl Fungi, which houses fungal species. Sponsors: EnsembFungi is a project run by EMBL - EBI to maintain annotation on selected genomes, based on the software developed in the Ensembl project developed jointly by the EBI and the Wellcome Trust Sanger Institute.
Proper citation: Ensembl Fungi (RRID:SCR_008681) Copy
http://plants.ensembl.org/index.html
Ensembl Genomes project produces genome databases for important species from across the taxonomic range, using the Ensembl software system. Five sites are now available, one of which is Ensembl Plants, which houses plant species. Sponsors: EnsembPlants is a project run by EMBL - EBI to maintain annotation on selected genomes, based on the software developed in the Ensembl project developed jointly by the EBI and the Wellcome Trust Sanger Institute.
Proper citation: Ensembl Plants (RRID:SCR_008680) Copy
https://bbgre.brc.iop.kcl.ac.uk
A database and associated tools for investigating the genetic basis of neurodisability. It combines phenotype information from patients with neurodevelopmental and behavioral problems with clinical genetic data, and displays this information on the human genome map. Basic access to genetic information (deletions, duplications) relating to participants with neurodevelopmental disorders is provided without an account; access to the full dataset requires an account. The genetic information that is available to view comprises potentially pathogenic copy number variation across the genome, detected by array comparative genome hybridization (aCGH) using a customized 44K oligonucleotide array.
Proper citation: Brain and Body Genetic Resource Exchange (RRID:SCR_008959) Copy
http://bacteria.ensembl.org/index.html
The Ensembl Genomes project produces genome databases for important species from across the taxonomic range, using the Ensembl software system. Five sites are now available, one of which is Ensembl Bacteria, which houses bacterial species. All bacterial collections in Ensembl Bacteria have been updated with the latest data from ENA and UniProtKB. New genomes have been added to Escherichia/Shigella (3 additional genomes) and Staphylococcus (3 additional genomes). The mapping of array probes has been expanded to all genomes in the Escherichia/Shigella and Staphylococcus collections. Ensembl Bacteria also now features improved interfaces for selecting regions of circular molecules a new visualisation allowing the large scale comparison of multiple genomes. In multi-synteny view, users can select multiple genomes and observe the syntenic relationships between them. Sponsors: EnsembBacteria is a project run by EMBL - EBI to maintain annotation on selected genomes, based on the software developed in the Ensembl project developed jointly by the EBI and the Wellcome Trust Sanger Institute.
Proper citation: Ensembl Bacteria (RRID:SCR_008679) Copy
http://cmr.jcvi.org/cgi-bin/CMR/shared/GenomePropertiesHomePage.cgi
The Genome Properties system consists of a suite of Properties which are carefully defined attributes of prokaryotic organisms whose status can be described by numerical values or controlled vocabulary terms for individual completely sequenced genomes. The system has been designed to capture the widest possible range of attributes and currently encompasses taxonomic terms, genometric calculations, metabolic pathways, systems of interacting macromolecular components and quantitative and descriptive experimental observations (phenotypes) from the literature. You may search the Genome Properties Database in 1 of 3 ways: * Search For Predicted Properties in the CMR: The Genome Property Search allows you to search the Genome Property database for state information for selected genomes and properties. * Perform a Keyword Search for a Specific Property: Lists all Genome Properties that match a specific text string. You can choose to search All Fields within a genome property or the Property Name. * Browse Top Level Genome Properties: Click on the properties to see the specific genome property report page. The Genome Properties system presents key aspects of prokaryotic biology using standardized computational methods and controlled vocabularies. Properties reflect gene content, phenotype, phylogeny and computational analyses. The results of searches using hidden Markov models allow many properties to be deduced automatically, especially for families of proteins (equivalogs) conserved in function since their last common ancestor. Additional properties are derived from curation, published reports and other forms of evidence. Genome Properties system was applied to 156 complete prokaryotic genomes, and is easily mined to find differences between species, correlations between metabolic features and families of uncharacterized proteins, or relationships among properties.
Proper citation: JCVI GenProp (RRID:SCR_004592) Copy
http://www.jbldesign.com/jmogil/enter.html
Database of genes regulated by pain derived from published manuscripts describing results of pain-relevant knockout studies. The database has two levels of exploration: across-gene and within-gene. The across-gene level, the PainGenesdbSelector, is encountered first. All genes in the database can be accessed and sorted by their gene name, protein name, common names and acronyms, or genomic position (by navigating a graphic representation of the mouse genome). The gene and protein names can be selected from an alphabetical list, or by typing a text string into a search box.
Proper citation: Pain Genes database (RRID:SCR_004771) Copy
http://www.ncbi.nlm.nih.gov/bioproject
Database of biological data related to a single initiative, originating from a single organization or from a consortium. A BioProject record provides users a single place to find links to the diverse data types generated for that project. It is a searchable collection of complete and incomplete (in-progress) large-scale sequencing, assembly, annotation, and mapping projects for cellular organisms. Submissions are supported by a web-based Submission Portal. The database facilitates organization and classification of project data submitted to NCBI, EBI and DDBJ databases that captures descriptive information about research projects that result in high volume submissions to archival databases, ties together related data across multiple archives and serves as a central portal by which to inform users of data availability. BioProject records link to corresponding data stored in archival repositories. The BioProject resource is a redesigned, expanded, replacement of the NCBI Genome Project resource. The redesign adds tracking of several data elements including more precise information about a project''''s scope, material, and objectives. Genome Project identifiers are retained in the BioProject as the ID value for a record, and an Accession number has been added. Database content is exchanged with other members of the International Nucleotide Sequence Database Collaboration (INSDC). BioProject is accessible via FTP.
Proper citation: NCBI BioProject (RRID:SCR_004801) Copy
http://genome.jgi.doe.gov/programs/plants/index.jsf
The goal of the DOE JGI Plant Genome Program is to shed light on the fundamental biology of photosynthesis and transduction of solar to chemical energy. Other areas of interest include characterizing: * Ecosystems and the role of terrestrial plants and oceanic phytoplankton-in carbon sequestration. * The role of plants in coping with toxic pollutants in soils by hyper-accumulation and detoxification. * Feedstocks for biofuels, e.g., biodiesel from soybean; cellulosic ethanol from perennial grasses. * The ability to respond to environmental change (e.g., loss of diversity from monoculture produces vulnerabilities; nitrogen fixing nodules in legumes reduce fertilizer need). * The generation of useful secondary metabolites (produced largely for disease resistance)- for positive/negative control in agriculture, with attendant influence on global carbon cycle. The Plant Genome Program accomplishes the above through the following activities: # Sequence. Produce genome sequences of key plant (and algal) species to accelerate biofuel development and understand response to climate change. # Function. Develop datasets (and synthetic biology tools) to elucidate functional elements in plant genomes, with special focus on handful of flagship genomes. # Variation. Characterize natural genomic variation in plants (and their associated microbiomes), and relate to biofuel sustainability and adaptation to climate change. # Integration. Provide a centralized hub for the retrieval and deep integrated analysis of plant genome datasets.
Proper citation: Plant Genome Resource at JGI (RRID:SCR_005315) Copy
http://deepbase.sysu.edu.cn/chipbase/
A database for decoding transcription factor binding maps, expression profiles and transcriptional regulation of long non-coding RNAs (lncRNAs, lincRNAs), microRNAs, other ncRNAs (snoRNAs, tRNAs, snRNAs, etc.) and protein-coding genes from ChIP-Seq data. ChIPBase currently includes millions of transcription factor binding sites (TFBSs) among 6 species. ChIPBase provides several web-based tools and browsers to explore TF-lncRNA, TF-miRNA, TF-mRNA, TF-ncRNA and TF-miRNA-mRNA regulatory networks., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: ChIPBase (RRID:SCR_005404) Copy
Collects mammalian cis- and trans-regulatory elements together with experimental evidence. Regulatory elements were mapped on to assembled genomes. Resource for gene regulation and function studies. Users can retrieve primers, search TF target genes, retrieve TF motifs, search Gene Regulatory Networks and orthologs, and make use of sequence analysis tools. Uses databases such as Genbank, EPD and DBTSS, and employ promoter finding program FirstEF combined with mRNA/EST information and cross-species comparisons. Manually curated.
Proper citation: Transcriptional Regulatory Element Database (RRID:SCR_005661) Copy
Database of computationally predicted Transcription Factors and binding sites in gamma-proteobacterial genomes. The user may browse a map containing all known E. coli transcription factors and regulatory interactions that connect them, and retrieve information on the conservation of each regulatory interaction across the 30 organisms included in the database. Downloading the information is straightforward, and navigation tabs added to dynamic pages ease navigation between the five interfaces of the database. The original prediction approach, based on the representation of binding sites through statistical models was complemented by a new approach that uses known E. coli regulatory sites as the basis for a pattern matching search of regulatory sites. The use of both approaches together resulted in a more intensive exploration of the sequence space of each regulator's binding site. These data should aid researchers in the design of microarray experiments and the interpretation of their results. They should also facilitate studies of Comparative Genomics of the regulatory networks of this group of organisms.
Proper citation: Tractor db (RRID:SCR_005610) Copy
Database providing integrated access to genome sequence, expression data and literature curation for Tuberculosis (TB) that houses genome assemblies for numerous strains of Mycobacterium tuberculosis (MTB) as well assemblies for over 20 strains related to MTB and useful for comparative analysis. TBDB stores pre- and post-publication gene-expression data from M. tuberculosis and its close relatives, including over 3000 MTB microarrays, 95 RT-PCR datasets, 2700 microarrays for human and mouse TB related experiments, and 260 arrays for Streptomyces coelicolor. (July 2010) To enable wide use of these data, TBDB provides a suite of tools for searching, browsing, analyzing, and downloading the data.
Proper citation: Tuberculosis Database (RRID:SCR_006619) Copy
http://www.youtube.com/ncbinlm
Videos from the National Center for Biotechnology Information including presentations and tutorials about NCBI biomolecular and biomedical literature databases and tools.
Proper citation: NCBI YouTube Channel (RRID:SCR_006084) Copy
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 22, 2016. Database for corrected read counts and genome mapping on NCBI's Short Read Archive. The corrected count was done using RECOUNT and the mapping with LAST. We also provide information of reference genome to which we aligned the short reads. We focus on transcriptomic data, specifically TSS-Seq and RNA-Seq. Because this is the type of data for which sequence count correction is most important. Hence we do not include the genomic reads. The current version contains 2,265 entries from 45 organisms, with read lengths from 17 to 100bp. Via a searchable and browseable interface users can obtain corrected data in formats useful for transcriptomic analysis. We provide the data grouped according to the genome, type of studies and submitter in TAB , PSL and BAM format. They contain the mapping position and annotation of reads observed and corrected counts.
Proper citation: RecountDB (RRID:SCR_006117) Copy
http://operons.ibt.unam.mx/OperonPredictor/
The Prokaryotic Operon DataBase (ProOpDB) constitutes one of the most precise and complete repository of operon predictions in our days. Using our novel and highly accurate operon algorithm, we have predicted the operon structures of more than 1,200 prokaryotic genomes. ProOpDB offers diverse alternatives by which a set of operon predictions can be retrieved including: i) organism name, ii) metabolic pathways, as defined by the KEGG database, iii) gene orthology, as defined by the COG database, iv) conserved protein motifs, as defined by the Pfam database, v) reference gene, vi) reference operon, among others. In order to limit the operon output to non-redundant organisms, ProOpDB offers an efficient protocol to select the more representative organisms based on a precompiled phylogenetic distances matrix. In addition, the ProOpDB operon predictions are used directly as the input data of our Gene Context Tool (GeConT) to visualize their genomic context and retrieve the sequence of their corresponding 5�� regulatory regions, as well as the nucleotide or amino acid sequences of their genes. The prediction algorithm The algorithm is a multilayer perceptron neural network (MLP) classifier, that used as input the intergenic distances of contiguous genes and the functional relationship scores of the STRING database between the different groups of orthologous proteins, as defined in the COG database. Nevertheless, the operon prediction of our method is not restricted to only those genes with a COG assignation, since we successfully defined new groups of orthologous genes and obtained, by extrapolation, a set of equivalent STRING-like scores based on conserved gene pairs on different genomes. Since the STRING functional relationships scores are determined in an un-bias manner and efficiently integrates a large amount of information coming from different sources and kind of evidences, the prediction made by our MLP are considerably less influenced by the bias imposed in the training procedure using one specific organism.
Proper citation: ProOpDB (RRID:SCR_006111) Copy
Database storing and integrating genomic data of diamondback moth (DBM), Plutella xylostella (L.). It provides comprehensive search tools and downloadable datasets for scientists to study comparative genomics, biological interpretation and gene annotation of this insect pest. DBM-DB contains assembled transcriptome datasets from multiple DBM strains and developmental stages, and the annotated genome of P. xylostella (version 2). They have also integrated publically available ESTs from NCBI and a putative gene set from a second DBM genome (KONAGbase) to enable users to compare different gene models. DBM-DB was developed with the capacity to incorporate future data resources, and will serve as a long-term and open-access database that can be conveniently used for research on the biology, distribution and evolution of DBM. This resource aims to help reduce the impact DBM has on agriculture using genomic and molecular tools.
Proper citation: DBM-DB (RRID:SCR_006258) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the T1D Resources search. From here you can search through a compilation of resources used by T1D and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that T1D has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on T1D then you can log in from here to get additional features in T1D such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into T1D you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within T1D that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.