We support boolean queries, use +,-,<,>,~,* to alter the weighting of terms
Software R package for analysis of case-control studies in genetic epidemiology.
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 23,2022. Software package to facilitate the testing of Copy Number Variant data for genetic association, typically in case-control studies.
Software for classes and statistical methods for large single nucleotide polymorphism (SNP) association studies.
Software program for counting k-mers in DNA sequence data. It identifies all the k-mers that occur more than once in a DNA sequence data set using a Bloom filter, a probabilistic data structure that stores all the observed k-mers implicitly in memory with greatly reduced memory requirements.
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 23,2022. Software for DNA sequencing analysis that integates with Sanger Sequencing files generated by Applied Biosystems Genetic Analyzers, MegaBACE, and Beckman CEQ electrophoresis systems. It can be used to find single nucleotide polymorphisms (SNPs), insertions and deletions (INDELS), and somatic mutations in direct sequencing, PCR sequencing, mitochondrial DNA sequencing, and resequencing projects.
A k-mer counting software that can count k-mers of large Illumina datasets on laptops and desktop computers.
Software utility for counting k-mers (sequences of consecutive k symbols) in a set of reads from genome sequencing projects. It scans the raw reads and produces a compact representation of all non-unique reads accompanied with number of their occurrences. The algorithm implemented makes use mostly of disk space rather than RAM, which allows to use KMC even on rather typical personal computers.
A collection of flexible and memory-efficient software programs for k-mer counting and indexing of large sequence sets. It is based on enhanced suffix arrays which gives a much larger flexibility concerning the choice of the k-mer size. It can process large data sizes of several billion bases.
A web-based genome analysis platform that integrates proprietary functional genomic data, metabolic reconstructions, expression profiling, and biochemical and microbiological data with publicly available information. Focused on microbial genomics, it provides better and faster identification of gene function across all organisms. Building upon a comprehensive genomic database integrated with a collection of microbial metabolic and non-metabolic pathways and using proprietary algorithms, it assigns functions to genes, integrates genes into pathways, and identifies previously unknown or mischaracterized genes, cryptic pathways and gene products. . * Automated and manual annotation of genes and genomes * Analysis of metabolic and non-metabolic pathways to understand organism physiology * Comparison of multiple genomes to identify shared and unique features and SNPs * Functional analysis of gene expression microarray data * Data-mining for target gene discovery * In silico metabolic engineering and strain improvement
Software to detect short insertion / deletion variants (and SNPs) from population sequence data, i.e. sequence reads generated from a population of individuals. It uses a probabilistic model to utilize sequence reads from a population of individuals to automatically account for context-specific sequencing errors associated with indels. piCALL is implemented in C for use on Linux platforms and can be applied to sequence data from different sequencing platforms. However, the method requires each individual in a dataset to be sequenced using the same platform. The reads for each individual should be aligned to the same reference genome sequence. Note that the program will not be able to call indels from individual sequence datasets or data from a small number of individuals.
A software basecaller for Illumina sequencers with calibrated quality scores.
Software for tracking and quantifying DNA damage patterns among ancient DNA sequencing reads generated by Next-Generation Sequencing platforms.
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 23,2022. Software program for deduplicating sequence fragments. It minimises memory usage by compressing sequences and using compact memory allocation techniques. A built-in parser allows a variety of input file formats and a simple specification language allows flexible output file formats. It can be made aware of paired-end reads, and it can handle degenerate sequence inserts intended to reveal amplification biases. Tally comes with reaper, a program for demultiplexing, trimming and filtering short read sequencing data.
Knowledge engineering software for reasoning with scientific observations and interpretations. The software has three parts: (a) the KEfED model editor - a design editor for creating KEfED models by drawing a flow diagram of an experimental protocol; (b) the KEfED data interface - a spreadsheet-like tool that permits users to enter experimental data pertaining to a specific model; (c) a "neural connection matrix" interface that presents neural connectivity as a table of ordinal connection strengths representing the interpretations of tract-tracing data. This tool also allows the user to view experimental evidence pertaining to a specific connection. The KEfED model is designed to provide a lightweight representation for scientific knowledge that is (a) generalizable, (b) a suitable target for text-mining approaches, (c) relatively semantically simple, and (d) is based on the way that scientist plan experiments and should therefore be intuitively understandable to non-computational bench scientists. The basic idea of the KEfED model is that scientific observations tend to have a common design: there is a significant difference between measurements of some dependent variable under conditions specified by two (or more) values of some independent variable.
Set of software modules for performing common ChIP-seq data analysis tasks across the whole genome, including positional correlation analysis, peak detection, and genome partitioning into signal-rich and signal-poor regions. The tools are designed to be simple, fast and highly modular. Each program carries out a well-defined data processing procedure that can potentially fit into a pipeline framework. ChIP-Seq is also freely available on a Web interface.
A suite of software tools for analyzing and manipulating next-generation sequencing datasets, such as FASTQ, BED and BAM format files. These tools provide a stable and modular platform for data management and analysis.
Software package for working with VCF files. Used to provide easily accessible methods for working with complex genetic variation data in the form of VCF files.Implements various utilities for processing Variant Call Format files, including validation, merging, comparing. Provides general Perl API.
Service that provides a comprehensive compilation of variant knowledge that allows you to identify pathogenic variants in human whole genome or exome sequences. It makes it easy to upload a complete genome?s worth of variations and identify the biologically relevant subset of known mutations, mutations that are novel and appear in a candidate disease genes, or mutations that are predicted to have a deleterious effect. The database includes a comprehensive collection of disease causing mutations from HGMD Professional, regulatory sites from TRANSFAC , and disease genes, drug targets and pathways from PROTEOME, as well as pharmacogenomic variants. It integrates the best public data-sets on somatic mutations, allele frequencies and clinical variants, in their most up-to-date version, for a total of more than 165 million annotations. It is possible to identify known pathogenic variants, remove harmless common variants, and obtain deleterious predictions for novel variants. With family data, it is possible to identify variants that are de novo, compound heterozygous only in the offspring. All of the results can be downloaded to Excel for further review. For core facilities and bioinformaticians, the complete underlying data is made available for download and easy integration into custom analysis pipelines. Genome Trax data is optimized to work with many other software packages, such as ANNOVARTM, CLC bio, Alamut, SimulConsult, and Cartagenia.
THIS RESOURCE IS NO LONGER IN SERVICE. Documented on August 18,2025.Software to classify the function and phylogeny of reads as short as 30 bp. It is flexible, which can utilize multiple data modules and downstream analysis scripts. It is fast, reading in signature lists of 5-500 million peptide signatures in 1-15 minutes, and subsequently processes genomic fragments at the rate of 6 Gbp/hr. It parallelizes without significant increase in memory requirements until I/O bound on multiple input files; parallelization works well on 64 processors.
Software that gives estimates of the structure and diversity of uncultured viral communities using metagenomic information.