Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
http://operons.ibt.unam.mx/OperonPredictor/
The Prokaryotic Operon DataBase (ProOpDB) constitutes one of the most precise and complete repository of operon predictions in our days. Using our novel and highly accurate operon algorithm, we have predicted the operon structures of more than 1,200 prokaryotic genomes. ProOpDB offers diverse alternatives by which a set of operon predictions can be retrieved including: i) organism name, ii) metabolic pathways, as defined by the KEGG database, iii) gene orthology, as defined by the COG database, iv) conserved protein motifs, as defined by the Pfam database, v) reference gene, vi) reference operon, among others. In order to limit the operon output to non-redundant organisms, ProOpDB offers an efficient protocol to select the more representative organisms based on a precompiled phylogenetic distances matrix. In addition, the ProOpDB operon predictions are used directly as the input data of our Gene Context Tool (GeConT) to visualize their genomic context and retrieve the sequence of their corresponding 5�� regulatory regions, as well as the nucleotide or amino acid sequences of their genes. The prediction algorithm The algorithm is a multilayer perceptron neural network (MLP) classifier, that used as input the intergenic distances of contiguous genes and the functional relationship scores of the STRING database between the different groups of orthologous proteins, as defined in the COG database. Nevertheless, the operon prediction of our method is not restricted to only those genes with a COG assignation, since we successfully defined new groups of orthologous genes and obtained, by extrapolation, a set of equivalent STRING-like scores based on conserved gene pairs on different genomes. Since the STRING functional relationships scores are determined in an un-bias manner and efficiently integrates a large amount of information coming from different sources and kind of evidences, the prediction made by our MLP are considerably less influenced by the bias imposed in the training procedure using one specific organism.
Proper citation: ProOpDB (RRID:SCR_006111) Copy
Database providing integrated access to genome sequence, expression data and literature curation for Tuberculosis (TB) that houses genome assemblies for numerous strains of Mycobacterium tuberculosis (MTB) as well assemblies for over 20 strains related to MTB and useful for comparative analysis. TBDB stores pre- and post-publication gene-expression data from M. tuberculosis and its close relatives, including over 3000 MTB microarrays, 95 RT-PCR datasets, 2700 microarrays for human and mouse TB related experiments, and 260 arrays for Streptomyces coelicolor. (July 2010) To enable wide use of these data, TBDB provides a suite of tools for searching, browsing, analyzing, and downloading the data.
Proper citation: Tuberculosis Database (RRID:SCR_006619) Copy
THIS RESOURCE IS NO LONGER IN SERVICE, documented August 22, 2016. Database for corrected read counts and genome mapping on NCBI's Short Read Archive. The corrected count was done using RECOUNT and the mapping with LAST. We also provide information of reference genome to which we aligned the short reads. We focus on transcriptomic data, specifically TSS-Seq and RNA-Seq. Because this is the type of data for which sequence count correction is most important. Hence we do not include the genomic reads. The current version contains 2,265 entries from 45 organisms, with read lengths from 17 to 100bp. Via a searchable and browseable interface users can obtain corrected data in formats useful for transcriptomic analysis. We provide the data grouped according to the genome, type of studies and submitter in TAB , PSL and BAM format. They contain the mapping position and annotation of reads observed and corrected counts.
Proper citation: RecountDB (RRID:SCR_006117) Copy
Database storing and integrating genomic data of diamondback moth (DBM), Plutella xylostella (L.). It provides comprehensive search tools and downloadable datasets for scientists to study comparative genomics, biological interpretation and gene annotation of this insect pest. DBM-DB contains assembled transcriptome datasets from multiple DBM strains and developmental stages, and the annotated genome of P. xylostella (version 2). They have also integrated publically available ESTs from NCBI and a putative gene set from a second DBM genome (KONAGbase) to enable users to compare different gene models. DBM-DB was developed with the capacity to incorporate future data resources, and will serve as a long-term and open-access database that can be conveniently used for research on the biology, distribution and evolution of DBM. This resource aims to help reduce the impact DBM has on agriculture using genomic and molecular tools.
Proper citation: DBM-DB (RRID:SCR_006258) Copy
ViralZone is a SIB Swiss Institute of Bioinformatics web-resource for all viral genus and families, providing general molecular and epidemiological information, along with virion and genome figures. Each virus or family page gives an easy access to UniProtKB/Swiss-Prot viral protein entries. ViralZone project is handled by the virus program of SwissProt group. Proteins popups were developed in collaboration with Prof. Christian von Mering and Andrea Franceschini, Bioinformatics Group , Institute of Molecular Life Sciences, University of Zurich, Winterthurerstrasse 190, CH-8057 Zurich, Switzerland, funded in part by the SIB Swiss Institute of bioinformatics. All pictures in ViralZone are copyright of the SIB Swiss Institute of Bioinformatics.
Proper citation: ViralZone (RRID:SCR_006563) Copy
Interactive database which incorporates a suite of tools designed to aid the interpretation of submicroscopic chromosomal imbalance. Used to enhance clinical diagnosis by retrieving information from bioinformatics resources relevant to the imbalance found in the patient. Contributing to the DECIPHER database is a Consortium, comprising an international community of academic departments of clinical genetics. Each center maintains control of its own patient data (which are password protected within the center''''s own DECIPHER project) until patient consent is given to allow anonymous genomic and phenotypic data to become freely viewable within Ensembl and other genome browsers. Once data are shared, consortium members are able to gain access to the patient report and contact each other to discuss patients of mutual interest, thus facilitating the delineation of new microdeletion and microduplication syndromes.
Proper citation: DECIPHER (RRID:SCR_006552) Copy
http://software.broadinstitute.org/gsea/msigdb/index.jsp
Collection of annotated gene sets for use with Gene Set Enrichment Analysis (GSEA) software.
Proper citation: Molecular Signatures Database (RRID:SCR_016863) Copy
Project exploring the spectrum of genomic changes involved in more than 20 types of human cancer that provides a platform for researchers to search, download, and analyze data sets generated. As a pilot project it confirmed that an atlas of changes could be created for specific cancer types. It also showed that a national network of research and technology teams working on distinct but related projects could pool the results of their efforts, create an economy of scale and develop an infrastructure for making the data publicly accessible. Its success committed resources to collect and characterize more than 20 additional tumor types. Components of the TCGA Research Network: * Biospecimen Core Resource (BCR); Tissue samples are carefully cataloged, processed, checked for quality and stored, complete with important medical information about the patient. * Genome Characterization Centers (GCCs); Several technologies will be used to analyze genomic changes involved in cancer. The genomic changes that are identified will be further studied by the Genome Sequencing Centers. * Genome Sequencing Centers (GSCs); High-throughput Genome Sequencing Centers will identify the changes in DNA sequences that are associated with specific types of cancer. * Proteome Characterization Centers (PCCs); The centers, a component of NCI's Clinical Proteomic Tumor Analysis Consortium, will ascertain and analyze the total proteomic content of a subset of TCGA samples. * Data Coordinating Center (DCC); The information that is generated by TCGA will be centrally managed at the DCC and entered into the TCGA Data Portal and Cancer Genomics Hub as it becomes available. Centralization of data facilitates data transfer between the network and the research community, and makes data analysis more efficient. The DCC manages the TCGA Data Portal. * Cancer Genomics Hub (CGHub); Lower level sequence data will be deposited into a secure repository. This database stores cancer genome sequences and alignments. * Genome Data Analysis Centers (GDACs) - Immense amounts of data from array and second-generation sequencing technologies must be integrated across thousands of samples. These centers will provide novel informatics tools to the entire research community to facilitate broader use of TCGA data. TCGA is actively developing a network of collaborators who are able to provide samples that are collected retrospectively (tissues that had already been collected and stored) or prospectively (tissues that will be collected in the future).
Proper citation: The Cancer Genome Atlas (RRID:SCR_003193) Copy
http://ki.se/en/meb/twingene-and-genomeeutwin
In collaboration with GenomeEUtwin, the TwinGene project investigates the importance of quantitative trait loci and environmental factors for cardiovascular disease. It is well known that genetic factors are of considerable importance for some familial lipid syndromes and that Type A Behavior pattern and increased lipid levels infer increased risk for cardiovascular disease. It is furthermore known that genetic factors are of importance levels of blood lipid biomarkers. The interplay of genetic and environmental effects for these risk factors in a normal population is less well understood and virtually unknown for the elderly. In the TwinGene project twins born before 1958 are contacted to participate. Health and medication data are collected from self-reported questionnaires, and blood sampling material is mailed to the subject who then contacts a local health care center for blood sampling and a health check-up. In the simple health check-up, height, weight, circumference of waist and hip, and blood pressure are measured. Blood is sampled for DNA extraction, serum collection and clinical chemistry tests of C-reactive protein, total cholesterol, triglycerides, HDL and LDL cholesterol, apolipo��protein A1 and B, glucose and HbA1C. The TwinGene cohort contains more than 10000 of the expected final number of 16000 individuals. Molecular genetic techniques are being used to identify Quantitative Trait Loci (QTLs) for cardiovascular disease and biomarkers in the TwinGene participants. Genome-wide linkage and association studies are ongoing. DZ twins have been genome-scanned with 1000 STS markers and a subset of 300 MZ twins have been genome-scanned with Illumina 317K SNP platform. Association of positional candidate SNPs arising from these genomscans are planned. The TwinGene project is associated with the large European collaboration denoted GenomEUtwin (www.genomeutwin.org, see below) which since 2002 has aimed at gathering genetic data on twins in Europe and setting up the infrastructure needed to enable pooling of data and joint analyses. It has been the funding source for obtaining the genome scan data. Types of samples: * EDTA whole blood * DNA * Serum Number of sample donors: 12 044 (sample collection completed)
Proper citation: KI Biobank - TwinGene (RRID:SCR_006006) Copy
https://bbgre.brc.iop.kcl.ac.uk
A database and associated tools for investigating the genetic basis of neurodisability. It combines phenotype information from patients with neurodevelopmental and behavioral problems with clinical genetic data, and displays this information on the human genome map. Basic access to genetic information (deletions, duplications) relating to participants with neurodevelopmental disorders is provided without an account; access to the full dataset requires an account. The genetic information that is available to view comprises potentially pathogenic copy number variation across the genome, detected by array comparative genome hybridization (aCGH) using a customized 44K oligonucleotide array.
Proper citation: Brain and Body Genetic Resource Exchange (RRID:SCR_008959) Copy
http://www.ncbi.nlm.nih.gov/projects/gap/cgi-bin/study.cgi?study_id=phs000674.v1.p1
Human genetics data from an immense (78,000) and ethnically diverse population available for secondary analysis to qualified researchers through the database of Genotypes and Phenotypes (dbGaP). It offers the opportunity to identify potential genetic risks and influences on a broad range of health conditions, particularly those related to aging. The GERA cohort is part of the Research Program on Genes, Environment, and Health (RPGEH), which includes more than 430,000 adult members of the Kaiser Permanente Northern California system. Data from this larger cohort include electronic medical records, behavioral and demographic information from surveys, and saliva samples from 200,000 participants obtained with informed consent for genomic and other analyses. The RPGEH database was made possible largely through early support from the Robert Wood Johnson Foundation to accelerate such health research. The genetic information in the GERA cohort translates into more than 55 billion bits of genetic data. Using newly developed techniques, the researchers conducted genome-wide scans to rapidly identify single nucleotide polymorphisms (SNPs) in the genomes of the people in the GERA cohort. These data will form the basis of genome-wide association studies (GWAS) that can look at hundreds of thousands to millions of SNPs at the same time. The RPGEH then combined the genetic data with information derived from Kaiser Permanente''s comprehensive longitudinal electronic medical records, as well as extensive survey data on participants'' health habits and backgrounds, providing researchers with an unparalleled research resource. As information is added to the Kaiser-UCSF database, the dbGaP database will also be updated.
Proper citation: Resource for Genetic Epidemiology Research on Adult Health and Aging (RRID:SCR_010472) Copy
http://www.genomicus.biologie.ens.fr/genomicus-72.01/cgi-bin/search.pl
A genome browser that enables users to navigate in genomes in several dimensions: linearly along chromosome axes, transversaly across different species, and chronologicaly along evolutionary time.
Proper citation: Genomicus (RRID:SCR_011791) Copy
http://www.informatics.jax.org/
Community model organism database for laboratory mouse and authoritative source for phenotype and functional annotations of mouse genes. MGD includes complete catalog of mouse genes and genome features with integrated access to genetic, genomic and phenotypic information, all serving to further the use of the mouse as a model system for studying human biology and disease. MGD is a major component of the Mouse Genome Informatics.Contains standardized descriptions of mouse phenotypes, associations between mouse models and human genetic diseases, extensive integration of DNA and protein sequence data, normalized representation of genome and genome variant information. Data are obtained and integrated via manual curation of the biomedical literature, direct contributions from individual investigators and downloads from major informatics resource centers. MGD collaborates with the bioinformatics community on the development and use of biomedical ontologies such as the Gene Ontology (GO) and the Mammalian Phenotype (MP) Ontology.
Proper citation: Mouse Genome Database (RRID:SCR_012953) Copy
THIS RESOURCE IS NO LONGER IN SERVICE.Documented on April 14,2022. Database of comprehensive information on the approximately 600 prokaryote species that are present in the human oral cavity. The majority of these species are uncultivated and unnamed, recognized primarily by their 16S rRNA sequences. The HOMD presents a provisional naming scheme for the currently unnamed species so that strain, clone, and probe data from any laboratory can be directly linked to a stably named reference entity. The HOMD links sequence data with phenotypic, phylogenetic, clinical, and bibliographic information. Full and partial oral bacterial genome sequences determined as part of this project and the Human Microbiome Project, are being added to the HOMD as they become available. HOMD offers easy to use tools for viewing all publicly available oral bacterial genomes. Data is also downloadable.
Proper citation: HOMD (RRID:SCR_012770) Copy
The Ensembl Genomes project produces genome databases for important species from across the taxonomic range, using the Ensembl software system. Five sites are now available, one of which is Ensembl Protists, which houses protists species. Sponsors: EnsembProtists is a project run by EMBL - EBI to maintain annotation on selected genomes, based on the software developed in the Ensembl project developed jointly by the EBI and the Wellcome Trust Sanger Institute.
Proper citation: Ensembl Protists (RRID:SCR_013154) Copy
GiardiaDB is a resource for information on Giardia lamblia. It contains gene information, including genomic attributes, protein expression patterns, evolution, and EST sequence information. The website provides tools for BLASTing, sequence retrieval, graphic visualization, and PubMed information.
Proper citation: GiardiaDB (RRID:SCR_013377) Copy
http://epsf.bmad.bii.a-star.edu.sg/cube/db/html/home.html
Cube-DB is a database of pre-evaluated conservation and specialization scores for residues in paralogous proteins belonging to multi-member families of human proteins. Protein family classification follows (largely) the classification suggested by HUGO Gene Nomenclature Committee. Sets of orhtologous protein sequences were generated by mutual-best-hit strategy using full vertebrate genomes available in Ensembl. The scores, described on documentation page, are assigned to each individual residue in a protein, and presented in the form of a table (html or downloadable xls formats) and mapped, when appropriate, onto the related structure (Jmol, Pymol, Chimera).
Proper citation: Cube-DB (RRID:SCR_013233) Copy
An international collaborative effort to develop and enrich new and existing reference ontologies for plants, improve ontology use and cross-references, and to develop data annotation standards. Users can search for ontology terms and bioentities and submit the ontology-related term requests by visiting the following GitHub request trackers.
Proper citation: Planteome (RRID:SCR_014411) Copy
http://signal.salk.edu/cgi-bin/RiceGE
Gene database for Japonica rice. RiceGE is associated with SIGnAL at the Salk Institute.
Proper citation: RiceGE (RRID:SCR_015061) Copy
http://hb.flatironinstitute.org/
Formerly known as GIANT (Genome-scale Integrated Analysis of gene Networks in Tissues), HumanBase applies machine learning algorithms to learn biological associations from massive genomic data collections. These integrative analyses reach beyond existing "biological knowledge" represented in the literature to identify novel, data-driven associations.
Proper citation: HumanBase (RRID:SCR_016145) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the T1D Resources search. From here you can search through a compilation of resources used by T1D and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that T1D has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on T1D then you can log in from here to get additional features in T1D such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into T1D you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within T1D that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.