Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
SciCrunch Registry is a curated repository of scientific resources, with a focus on biomedical resources, including tools, databases, and core facilities - visit SciCrunch to register your resource.
International collaboration producing an extensive public catalog of human genetic variation, including SNPs and structural variants, and their haplotype contexts, in an effort to provide a foundation for investigating the relationship between genotype and phenotype. The genomes of about 2500 unidentified people from about 25 populations around the world were sequenced using next-generation sequencing technologies. Redundant sequencing on various platforms and by different groups of scientists of the same samples can be compared. The results of the study are freely and publicly accessible to researchers worldwide. The consortium identified the following populations whose DNA will be sequenced: Yoruba in Ibadan, Nigeria; Japanese in Tokyo; Chinese in Beijing; Utah residents with ancestry from northern and western Europe; Luhya in Webuye, Kenya; Maasai in Kinyawa, Kenya; Toscani in Italy; Gujarati Indians in Houston; Chinese in metropolitan Denver; people of Mexican ancestry in Los Angeles; and people of African ancestry in the southwestern United States. The goal Project is to find most genetic variants that have frequencies of at least 1% in the populations studied. Sequencing is still too expensive to deeply sequence the many samples being studied for this project. However, any particular region of the genome generally contains a limited number of haplotypes. Data can be combined across many samples to allow efficient detection of most of the variants in a region. The Project currently plans to sequence each sample to about 4X coverage; at this depth sequencing cannot provide the complete genotype of each sample, but should allow the detection of most variants with frequencies as low as 1%. Combining the data from 2500 samples should allow highly accurate estimation (imputation) of the variants and genotypes for each sample that were not seen directly by the light sequencing. All samples from the 1000 genomes are available as lymphoblastoid cell lines (LCLs) and LCL derived DNA from the Coriell Cell Repository as part of the NHGRI Catalog. The sequence and alignment data generated by the 1000genomes project is made available as quickly as possible via their mirrored ftp sites. ftp://ftp.1000genomes.ebi.ac.uk ftp://ftp-trace.ncbi.nlm.nih.gov/1000genomes
Proper citation: 1000 Genomes: A Deep Catalog of Human Genetic Variation (RRID:SCR_006828) Copy
http://ftp://ftp.ebi.ac.uk/pub/databases/taxonomy
The taxonomy database of the International Sequence Database Collaboration contains the names of all organisms that are represented in the sequence databases with at least one nucleotide or protein sequence. The database is available via the EBI SRS server of ftp.
Proper citation: Taxonomy (RRID:SCR_004299) Copy
Repository where scientists and organizations can share, store, manipulate, and publish biological network data. Users can also run their own copies of NDEx Server software in cases where stored networks must be kept in highly secure environment (such as for HIPAA compliance) or where high application load is incompatible with shared public resource. Open source software system that is part of Cytoscape family. Project of Cytoscape Consortium in conjunction with Ideker lab at UCSD School of Medicine. Public forum where biologists can exchange and publish computable network models in many types and formats. NDEx is based on REST web API which can be accessed by any application, including NDEx website and NDEx Cytoscape App. NDEx networks are assigned stable, globally unique URIs and so can be referenced by publications, by other networks, and by analytic applications.
Proper citation: Network Data Exchange (NDEx) (RRID:SCR_003943) Copy
http://www.mycancergenome.org/
A freely available online personalized cancer medicine knowledge resource for physicians, patients, caregivers and researchers that gives up-to-date information on what mutations make cancers grow and related therapeutic implications, including available clinical trials. It is a one-stop tool that matches tumor mutations to therapies, making information accessible and convenient for busy clinicians.
Proper citation: My Cancer Genome (RRID:SCR_004140) Copy
http://tools.niehs.nih.gov/polg/
Database that lists all known mutations in the coding region of the POLG gene and describes the associated disease. Human DNA polymerase is composed of two subunits, a 140 kDa catalytic subunit encoded by the POLG on chromosome 15q25, and a 55kDa accessory subunit encoded by the POLG2 gene on chromosome 17q23-24. A number of mutations have been mapped to the gene for the catalytic subunit of DNA polymerase, POLG, and found to be associated with mitochondrial diseases. The nucleotide changes are numbered from the initiation Methionine codon and are based on the cDNA (accession U60325.1) and gene sequence (accession AF497906.1).
Proper citation: Human DNA Polymerase Gamma Mutation Database (RRID:SCR_004722) Copy
https://www.stanleygenomics.org/
The Stanley Online Genomics Database uses samples from the Stanley Medical Research Institute (SMRI) Brain Bank. These samples were processed and run on gene expression arrays by a variety of researchers in collaboration with the SMRI. These researchers have performed analyses on their respective studies using a range of analytic approaches. All of the genomic data have been aggregated in this online database, and a consistent set of analyses have been applied to each study. Additionally, a comprehensive set of cross-study analyses have been performed. A thorough collection of gene expression summaries are provided, inclusive of patient demographics, disease subclasses, regulated biological pathways, and functional classifications. Raw data is also available to download. The database is derived from two sets of brain samples, the Stanley Array collection and the Stanley Consortium collection. The Stanley Array collection contains 105 patients, and the Stanley Consortium collection contains 60 patients. Multiple genomic studies have been conducted using these brain samples. From these studies, twelve were selected for inclusion in the database on the basis of number of patients studied, genomic platform used, and data quality. The Consortium collection studies have fewer patients but more diversity in brain regions and array platforms, while the Array collection studies are more homogenous. There are tradeoffs, the Consortium results will be more variable, but findings may be more broadly representative. The collections contain brain samples from subjects in four main groups: Bipolar Schizophrenia, Depression, and Controls Brain regions used in the studies include: Broadman Area 6, Broadman Area 8/9, Broadman Area 10, Broadman Area 46, Cerebellum The 12 studies encompass a range of microarray platforms: Affymetrix HG-U95Av2, Affymetrix HG-U133A, Affymetrix HG-U133 2.0+, Codelink Human 20K, Agilent Human I, Custom cDNA Publications based on any of the clinical or genomic data should credit the Stanley Medical Research Institute, as well as any individual SMRI collaborators whose data is being used. Publications which make use of analytic results/methods in the database should additionally cite Dr. Michael Elashoff. Registration is required to access the data.
Proper citation: Stanley Medical Research Institute Online Genomics Database (RRID:SCR_004859) Copy
A database of three-dimensional protein models calculated by comparative modeling. ModBase is organized into datasets, which are either available to the public, to the academic community, or to specific users. 20 unique amidohydrolase and 41 unique enolase structures have been determined have been included in the database.
Proper citation: ModBase (RRID:SCR_004642) Copy
United Network for Organ Sharing (UNOS) is the private, non-profit organization that manages the nation''s organ transplant system under contract with the federal government. UNOS is involved in many aspects of the organ transplant and donation process: * Managing the national transplant waiting list, matching donors to recipients 24 hours a day, 365 days a year. * Maintaining the database that contains all organ transplant data for every transplant event that occurs in the U.S. * Bring together members to develop policies that make the best use of the limited supply of organs and give all patients a fair chance at receiving the organ they need, regardless of age, sex, ethnicity, religion, lifestyle or financial/social status. * Monitoring every organ match to ensure organ allocation policies are followed. * Provides assistance to patients, family members and friends. * Educates transplant professionals about their important role in the donation and transplant processes. * Educating the public about the importance of organ donation. UNOS was first awarded the national Organ Procurement and Transplantation Network (OPTN) contract in 1986 by the U.S. Department of Health and Human Services. UNOS continues as the only organization ever to operate the OPTN. As part of the OPTN contract, UNOS has: * established an organ sharing system that maximizes the efficient use of deceased organs through equitable and timely allocation * established a system to collect, store, analyze and publish data pertaining to the patient waiting list, organ matching, and transplants * informed, consulted and guided persons and organizations concerned with human organ transplantation in order to increase the number of organs available for transplantation
Proper citation: UNOS - United Network for Organ Sharing (RRID:SCR_004976) Copy
Database of known and predicted protein interactions. The interactions include direct (physical) and indirect (functional) associations and are derived from four sources: Genomic Context, High-throughput experiments, (Conserved) Coexpression, and previous knowledge. STRING quantitatively integrates interaction data from these sources for a large number of organisms, and transfers information between these organisms where applicable. The database currently covers 5''214''234 proteins from 1133 organisms. (2013)
Proper citation: STRING (RRID:SCR_005223) Copy
A comprehensive collection of human transcription factor binding sites models. DNA sequences of TF binding regions obtained by both pregenomic and high-throughput methods were collected from existing databases and other public data. The ChIPMunk software was used to construct positional weight matrices. Four motif discovery strategies were tested based on different motif shape priors including flat and periodic priors associated with DNA helix pitch. A quality rating was manually assigned to each model based on known binding preferences. An appropriate TFBS model was selected for each TF, with similar models selected for related TFs. In any case only one model per TF was selected unless there was additional evidence for two distinct binding models or different stable modes of dimerization. All TFBS models and initial binding segments data used for motif discovery were mapped to UniPROT IDs.
Proper citation: HOCOMOCO (RRID:SCR_005409) Copy
http://www.membranetransport.org
TransportDB is a relational database describing the predicted cytoplasmic membrane transport protein complement for organisms whose complete genome sequence are available. For each organism, its complete membrane transport complement was identified, classified into protein families according to the TC classification system, and functional predictions are provided.For each organism, a summary page is available, overviewing the whole transporter system, including transporter types and individual transporter families. For individual transporter types, a detailed list of transporters with their possible substrates is shown with links to individual protein page which contains protein sequence and annotation information. You can also compare the transporter system from two or more different organisms. A search engine is set up for easy search in our transporter database for transporter type, family, individual proteins and their substrates. You can also blast search your protein sequence against our transporter database.With the rapid development of genomic sequencing both in TIGR and in other institutes, more and more genomes are available for the analysis of their transporter system. We will keep updating this site with the newly published genomes. If you have any suggestions, corrections, or comments on our site, please contact us. We are currently working on providing additional functionality for this database.
Proper citation: TransportDB (RRID:SCR_005643) Copy
http://deepbase.sysu.edu.cn/chipbase/
A database for decoding transcription factor binding maps, expression profiles and transcriptional regulation of long non-coding RNAs (lncRNAs, lincRNAs), microRNAs, other ncRNAs (snoRNAs, tRNAs, snRNAs, etc.) and protein-coding genes from ChIP-Seq data. ChIPBase currently includes millions of transcription factor binding sites (TFBSs) among 6 species. ChIPBase provides several web-based tools and browsers to explore TF-lncRNA, TF-miRNA, TF-mRNA, TF-ncRNA and TF-miRNA-mRNA regulatory networks., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Proper citation: ChIPBase (RRID:SCR_005404) Copy
Collects mammalian cis- and trans-regulatory elements together with experimental evidence. Regulatory elements were mapped on to assembled genomes. Resource for gene regulation and function studies. Users can retrieve primers, search TF target genes, retrieve TF motifs, search Gene Regulatory Networks and orthologs, and make use of sequence analysis tools. Uses databases such as Genbank, EPD and DBTSS, and employ promoter finding program FirstEF combined with mRNA/EST information and cross-species comparisons. Manually curated.
Proper citation: Transcriptional Regulatory Element Database (RRID:SCR_005661) Copy
An integrated database of human maladies and their annotations, modeled on the architecture and richness of the popular GeneCards database of human genes. The database contains 17,705 diseases, consolidated from 28 sources.
Proper citation: MalaCards (RRID:SCR_005817) Copy
Database providing integrated access to genome sequence, expression data and literature curation for Tuberculosis (TB) that houses genome assemblies for numerous strains of Mycobacterium tuberculosis (MTB) as well assemblies for over 20 strains related to MTB and useful for comparative analysis. TBDB stores pre- and post-publication gene-expression data from M. tuberculosis and its close relatives, including over 3000 MTB microarrays, 95 RT-PCR datasets, 2700 microarrays for human and mouse TB related experiments, and 260 arrays for Streptomyces coelicolor. (July 2010) To enable wide use of these data, TBDB provides a suite of tools for searching, browsing, analyzing, and downloading the data.
Proper citation: Tuberculosis Database (RRID:SCR_006619) Copy
A curated repository of more than 206000 regulatory associations between transcription factors (TF) and target genes in Saccharomyces cerevisiae, based on more than 1300 bibliographic references. It also includes the description of 326 specific DNA binding sites shared among 113 characterized TFs. Further information about each Yeast gene has been extracted from the Saccharomyces Genome Database (SGD). For each gene the associated Gene Ontology (GO) terms and their hierarchy in GO was obtained from the GO consortium. Currently, YEASTRACT maintains a total of 7130 terms from GO. The nucleotide sequences of the promoter and coding regions for Yeast genes were obtained from Regulatory Sequence Analysis Tools (RSAT). All the information in YEASTRACT is updated regularly to match the latest data from SGD, GO consortium, RSA Tools and recent literature on yeast regulatory networks. YEASTRACT includes DISCOVERER, a set of tools that can be used to identify complex motifs found to be over-represented in the promoter regions of co-regulated genes. DISCOVERER is based on the MUSA algorithm. These algorithms take as input a list of genes and identify over-represented motifs, which can then be compared with transcription factor binding sites described in the YEASTRACT database.
Proper citation: Yeast Search for Transcriptional Regulators And Consensus Tracking (RRID:SCR_006076) Copy
http://operons.ibt.unam.mx/OperonPredictor/
The Prokaryotic Operon DataBase (ProOpDB) constitutes one of the most precise and complete repository of operon predictions in our days. Using our novel and highly accurate operon algorithm, we have predicted the operon structures of more than 1,200 prokaryotic genomes. ProOpDB offers diverse alternatives by which a set of operon predictions can be retrieved including: i) organism name, ii) metabolic pathways, as defined by the KEGG database, iii) gene orthology, as defined by the COG database, iv) conserved protein motifs, as defined by the Pfam database, v) reference gene, vi) reference operon, among others. In order to limit the operon output to non-redundant organisms, ProOpDB offers an efficient protocol to select the more representative organisms based on a precompiled phylogenetic distances matrix. In addition, the ProOpDB operon predictions are used directly as the input data of our Gene Context Tool (GeConT) to visualize their genomic context and retrieve the sequence of their corresponding 5�� regulatory regions, as well as the nucleotide or amino acid sequences of their genes. The prediction algorithm The algorithm is a multilayer perceptron neural network (MLP) classifier, that used as input the intergenic distances of contiguous genes and the functional relationship scores of the STRING database between the different groups of orthologous proteins, as defined in the COG database. Nevertheless, the operon prediction of our method is not restricted to only those genes with a COG assignation, since we successfully defined new groups of orthologous genes and obtained, by extrapolation, a set of equivalent STRING-like scores based on conserved gene pairs on different genomes. Since the STRING functional relationships scores are determined in an un-bias manner and efficiently integrates a large amount of information coming from different sources and kind of evidences, the prediction made by our MLP are considerably less influenced by the bias imposed in the training procedure using one specific organism.
Proper citation: ProOpDB (RRID:SCR_006111) Copy
The PDGene database aims to provide a comprehensive, unbiased and regularly updated collection of genetic association studies performed on Parkinson's disease (PD) phenotypes. Eligible publications are identified following systematic searches of scientific literature databases, as well as the table of contents of journals in genetics, neurology, and psychiatry. The database can be searched either by a variety of dropdown menus or by specific keywords. For each gene, summary overviews are provided displaying key characteristics for each publication, including links to genotype distributions of the polymorphisms studied, random-effects allelic meta-analyses, and funnel plots for an assessment of publication bias. The PDGene database, developed by Massachusetts General Hospital/Harvard Medical School, The Michael J. Fox Foundation and the Alzheimer Research Forum, is supported by a grant from The Michael J. Fox Foundation in partnership with the Alzheimer Research Forum.
Proper citation: PDGene - A database for Parkinsons disease genetic association studies (RRID:SCR_006666) Copy
wFleaBase provides gene and genomic information for species of the genus Daphnia - commonly known as the water flea. It contains the genome of Daphnia pulex and other species, including bulk data files, and all gene pages, plus genomics tools including microsatellites, cDNA, Cosmid and BAC libraries, GSS and ESTs, and microarrays. It also contains maps of the Daphnia genome, and genome annotation tools. The freshwater crustacean Daphnia is a model system for ecology, evolution and the environmental sciences. The rapidly growing genomic data for this organism is stimulating interdisciplinary research to understand the complex interplay between genome structure, gene expression, individual fitness, and population-level responses to chemical contaminants and environmental change.wFleaBase includes data from all species of the genus, yet the primary species are D. pulex and D. magna, because of the broad set of genomic tools that have already been developed for these animals. A complete sequence for Daphnia pulex is now available at this site. Please observe this Data release policy. The data is a first characterization of the crustacean genome, which was made possible by the U.S. Department of Energy (DOE) Joint Genome Institute (JGI) in collaboration with the Daphnia Genomics Consortium (DGC) whose members were funded by the National Science Foundation. Category: Genomics Databases (non-vertebrate) Subcategory: Invertebrate genome databases
Proper citation: wFleaBase (RRID:SCR_006018) Copy
Database storing and integrating genomic data of diamondback moth (DBM), Plutella xylostella (L.). It provides comprehensive search tools and downloadable datasets for scientists to study comparative genomics, biological interpretation and gene annotation of this insect pest. DBM-DB contains assembled transcriptome datasets from multiple DBM strains and developmental stages, and the annotated genome of P. xylostella (version 2). They have also integrated publically available ESTs from NCBI and a putative gene set from a second DBM genome (KONAGbase) to enable users to compare different gene models. DBM-DB was developed with the capacity to incorporate future data resources, and will serve as a long-term and open-access database that can be conveniently used for research on the biology, distribution and evolution of DBM. This resource aims to help reduce the impact DBM has on agriculture using genomic and molecular tools.
Proper citation: DBM-DB (RRID:SCR_006258) Copy
Can't find your Tool?
We recommend that you click next to the search bar to check some helpful tips on searches and refine your search firstly. Alternatively, please register your tool with the SciCrunch Registry by adding a little information to a web form, logging in will enable users to create a provisional RRID, but it not required to submit.
Welcome to the kravitz2 Resources search. From here you can search through a compilation of resources used by kravitz2 and see how data is organized within our community.
You are currently on the Community Resources tab looking through categories and sources that kravitz2 has compiled. You can navigate through those categories from here or change to a different tab to execute your search through. Each tab gives a different perspective on data.
If you have an account on kravitz2 then you can log in from here to get additional features in kravitz2 such as Collections, Saved Searches, and managing Resources.
Here is the search term that is being executed, you can type in anything you want to search for. Some tips to help searching:
You can save any searches you perform for quick access to later from here.
We recognized your search term and included synonyms and inferred terms along side your term to help get the data you are looking for.
If you are logged into kravitz2 you can add data records to your collections to create custom spreadsheets across multiple sources of data.
Here are the sources that were queried against in your search that you can investigate further.
Here are the categories present within kravitz2 that you can filter your data on
Here are the subcategories present within this category that you can filter your data on
If you have any further questions please check out our FAQs Page to ask questions and see our tutorials. Click this button to view this tutorial again.