We support boolean queries, use +,-,<,>,~,* to alter the weighting of terms
Accurate error correction in high-throughput sequencing data.
A tool for error correction of short read datasets with non-uniform coverage, such as single-cell data., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Tool to find regions of sequence similarity within selected protein databases quickly, with minimum loss of sensitivity.
A web-based tool used to search translated nucleotide databases using a translated nucleotide query.
Tool to search translated nucleotide databases using a protein query.
Software that searches for short patterns in large DNA databases, allowing for approximate matches., THIS RESOURCE IS NO LONGER IN SERVICE. Documented on September 16,2025.
Software for an accelerated version of the popular NCBI-BLAST using a general-purpose graphics processing unit (GPU). It s nearly four times faster, while producing identical results. GPU-BLAST supports: protein alignment according to blastp (it does not support psiblast), multiple CPU threads working in parallel with a single GPU, and input files with multiple protein queries.
Software package for DNA and protein sequence alignment to find regions of local or global similarity between Protein or DNA sequences, either by searching Protein or DNA databases, or by identifying local duplications within a sequence.
A multiple sequence alignment server which can align Protein, DNA and RNA sequences.
Downloadable data set designed to assess the performance of both multiple and pairwise (protein) sequence alignment algorithms, and is extremely easy to use. Currently, the database contains 2 sets, each consisting of a number of subsets with related sequences. It''s main features are: * Covers the entire known fold space (SCOP classification), with subsets provided by the ASTRAL compendium * All structures have high quality, with 100% resolved residues * Structure alignments have been derived carefully, using both SOFI and CE, and Relaxed Transitive Alignment * At most 25 sequences in each subset to avoid overrepresentation of large folds* Automated running, archiving and scoring of programs through a few Perl scripts The Twilight Zone set is divided into sequence groups that each represent a SCOP fold. All sequences within a group share a pairwise Blast e-value of at least 1, for a theoretical database size of 100 million residues. Sequence similarity is thus very low, between 0-25% identity, and a (traceable) common evolutionary origin cannot be established between most pairs even though their structures are (distantly) similar. This set therefore represents the worst case scenario for sequence alignment, which unfortunately is also the most frequent one, as most related sequences share less than 25% identity. The Superfamilies set consists of groups that each represent a SCOP superfamily, and therefore contain sequences with a (putative) common evolutionary origin. However, they share at most 50% identity, which is still challenging for any sequence alignment algorithm. Frequently, alignments are performed to establish whether or not sequences are related. To benchmark this, a second version of both the Twilight Zone and the Superfamilies set is provided, in which to each alignment problem a number of false positives, i.e. sequences not related to the original set, are added. Database specifications: * Current version: 1.65 (concurrent with PDB, SCOP and ASTRAL) * Twilight Zone set (with false positives): 209 groups, 1740 (3280) sequences, 10667 (44056) related pairs * Superfamilies set (with false positives): 425 groups, 3280 (6526) sequences, 19092 (79095) related pairs
Software for a Laboratory Information Management System (LIMS) developed to support the unpredictable workflows of Molecular biology and Protein production labs of all sizes.
Software for an open, distributed system for managing biological information that supports biological research data workflows from the source (i.e. the measurement instruments) to facilitate the process of answering biological questions by means of cross-domain queries against raw data, processed data, knowledge resources and its corresponding metadata. The openBIS software framework can be easily extended and has been customized for the following technologies: * High Content Screening * Proteomics * Deep Sequencing * Metabolomics
Software for improving multiple sequence alignment using probabilistic sampling.
Efficient protein multiple sequence alignment program, which has demonstrated a statistically significant improvement in accuracy compared to several leading alignment tools.
Multiple sequence alignment method with reduced time and space complexity.Multiple sequence alignment with high accuracy and high throughput. Data analysis service for multiple sequence comparison by log- expectation.
Software package as multiple alignment program for amino acid or nucleotide sequences. Can align up to 500 sequences or maximum file size of 1 MB. First version of MAFFT used algorithm based on progressive alignment, in which sequences were clustered with help of Fast Fourier Transform. Subsequent versions have added other algorithms and modes of operation, including options for faster alignment of large numbers of sequences, higher accuracy alignments, alignment of non-coding RNA sequences, and addition of new sequences to existing alignments.
A fast and accurate multiple sequence alignment algorithm.
Software for searching DNA sequence databases for RNA structure and sequence similarities.
Software tools for comparative genomics.Comprehensive suite of programs and databases for comparative analysis of genomic sequences. There are two ways of using VISTA - you can submit your own sequences and alignments for analysis (VISTA servers) or examine pre-computed whole-genome alignments of different species.
A generic sequence comparison tool for visualizing genome alignments both within and between species.