Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.
Spotted steed (Hemibarbus maculatus Bleeker, 1871), a small and medium-sized benthic fish, is widely distributed between the Yangtze and Heilongjiang River basins and is considered to be one of the most widely distributed freshwater fish in East Asia. It is also an economically valuable aquaculture fish and has become a commercial freshwater aquaculture species in China. Here, a high-quality chromosome-level genome of spotted steed was produced by combining PacBio single molecule sequencing technique and high-throughput chromosome conformation capture technologies. Ultimately, the genome was assembled into 1098.66 Mb with a contig N50 of 33.57 Mb and a scaffold N50 of 40.40 Mb. We constructed a chromosome-level genome assembled with 25 chromosomes, whose total lengths accounted for 97.49%, and the assembled genome represents 95.64% completeness (BUSCO). We also identified 285.16 Mb (25.96%) of repetitive genome sequences and 23233 predicted genes. A total of 23021 genes were functionally annotated, representing 99.09% of the predicted genes. These results will provide valuable genomic resources for subsequent study of the genetic, evolutionary, and biological characteristics of spotted steed.
Pubmed ID: 41888160
Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.
Quality assessment software tool for evaluating and comparing genome assemblies. It works both with and without a given reference genome. It produces many reports, summary tables and plots.
View all literature mentionsOriginal SAMTOOLS package has been split into three separate repositories including Samtools, BCFtools and HTSlib. Samtools for manipulating next generation sequencing data used for reading, writing, editing, indexing,viewing nucleotide alignments in SAM,BAM,CRAM format. BCFtools used for reading, writing BCF2,VCF, gVCF files and calling, filtering, summarising SNP and short indel sequence variants. HTSlib used for reading, writing high throughput sequencing data.
View all literature mentionsA database of protein families, each represented by multiple sequence alignments and hidden Markov models (HMMs). Users can analyze protein sequences for Pfam matches, view Pfam family annotation and alignments, see groups of related families, look at the domain organization of a protein sequence, find the domains on a PDB structure, and query Pfam by keywords. There are two components to Pfam: Pfam-A and Pfam-B. Pfam-A entries are high quality, manually curated families that may automatically generate a supplement using the ADDA database. These automatically generated entries are called Pfam-B. Although of lower quality, Pfam-B families can be useful for identifying functionally conserved regions when no Pfam-A entries are found. Pfam also generates higher-level groupings of related families, known as clans (collections of Pfam-A entries which are related by similarity of sequence, structure or profile-HMM).
View all literature mentionsSoftware package for functional analysis of sequences by classifying them into families and predicting presence of domains and sites. Scans sequences against InterPro's signatures. Characterizes nucleotide or protein function by matching it with models from several different databases. Used in large scale analysis of whole proteomes, genomes and metagenomes. Available as Web based version and standalone Perl version and SOAP Web Service.
View all literature mentionsA software package for visualizing data and information. It visualizes data in a circular layout - this makes Circos ideal for exploring relationships between objects or positions.
View all literature mentionsSoftware tool that screens DNA sequences for interspersed repeats and low complexity DNA sequences. The output of the program is a detailed annotation of the repeats that are present in the query sequence as well as a modified version of the query sequence in which all the annotated repeats have been masked (default: replaced by Ns). Currently over 56% of human genomic sequence is identified and masked by the program. Sequence comparisons in RepeatMasker are performed by one of several popular search engines including nhmmer, cross_match, ABBlast/WUBlast, RMBlast and Decypher. RepeatMasker makes use of curated libraries of repeats and currently supports Dfam ( profile HMM library ) and RepBase ( consensus sequence library ).
View all literature mentionsSoftware tool to quantitatively measure genome assembly and annotation completeness based on evolutionarily informed expectations of gene content.
View all literature mentionsGraph-based alignment of next generation sequencing reads to a population of genomes.
View all literature mentions