Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

The genome of Acorus deciphers insights into early monocot evolution.

Xing Guo | Fang Wang | Dongming Fang | Qiongqiong Lin | Sunil Kumar Sahu | Liuming Luo | Jiani Li | Yewen Chen | Shanshan Dong | Sisi Chen | Yang Liu | Shixiao Luo | Yalong Guo | Huan Liu
Nature communications | 2023

Acorales is the sister lineage to all the other extant monocot plants. Genomic resource enhancement of this genus can help to reveal early monocot genomic architecture and evolution. Here, we assemble the genome of Acorus gramineus and reveal that it has ~45% fewer genes than the majority of monocots, although they have similar genome size. Phylogenetic analyses based on both chloroplast and nuclear genes consistently support that A. gramineus is the sister to the remaining monocots. In addition, we assemble a 2.2 Mb mitochondrial genome and observe many genes exhibit higher mutation rates than that of most angiosperms, which could be the reason leading to the controversies of nuclear genes- and mitochondrial genes-based phylogenetic trees existing in the literature. Further, Acorales did not experience tau (τ) whole-genome duplication, unlike majority of monocot clades, and no large-scale gene expansion is observed. Moreover, we identify gene contractions and expansions likely linking to plant architecture, stress resistance, light harvesting, and essential oil metabolism. These findings shed light on the evolution of early monocots and genomic footprints of wetland plant adaptations.

Pubmed ID: 37339966

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


GlimmerHMM (tool)

RRID:SCR_002654

A gene finder based on a Generalized Hidden Markov Model (GHMM). Although the gene finder conforms to the overall mathematical framework of a GHMM, additionally it incorporates splice site models adapted from the GeneSplicer program and a decision tree adapted from GlimmerM. It also utilizes Interpolated Markov Models for the coding and noncoding models . Currently, GlimmerHMM's GHMM structure includes introns of each phase, intergenic regions, and four types of exons (initial, internal, final, and single).

View all literature mentions

Hmmer (tool)

RRID:SCR_005305

Tool for searching sequence databases for homologs of protein sequences, and for making protein sequence alignments. It implements methods using probabilistic models called profile hidden Markov models (profile HMMs). Compared to BLAST, FASTA, and other sequence alignment and database search tools based on older scoring methodology, HMMER aims to be significantly more accurate and more able to detect remote homologs because of the strength of its underlying mathematical models. In the past, this strength came at significant computational expense, but in the new HMMER3 project, HMMER is now essentially as fast as BLAST.

View all literature mentions

MAKER (tool)

RRID:SCR_005309

Software genome annotation pipeline. Portable and easily configurable genome annotation pipeline. Used to allow smaller eukaryotic and prokaryotic genomeprojects to independently annotate their genomes and to create genome databases. MAKER identifies repeats, aligns ESTs and proteins to genome, produces ab-initio gene predictions and automatically synthesizes these data into gene annotations having evidence based quality values.

View all literature mentions

InterProScan (tool)

RRID:SCR_005829

Software package for functional analysis of sequences by classifying them into families and predicting presence of domains and sites. Scans sequences against InterPro's signatures. Characterizes nucleotide or protein function by matching it with models from several different databases. Used in large scale analysis of whole proteomes, genomes and metagenomes. Available as Web based version and standalone Perl version and SOAP Web Service.

View all literature mentions

CAFE (tool)

RRID:SCR_005983

R software package for the detection of gross chromosomal abnormalities from gene expression microarray data.

View all literature mentions

Catalogue of Life (tool)

RRID:SCR_006701

Comprehensive and authoritative global index of species of animals, plants, fungi and micro-organisms. It consists of a single integrated species checklist and taxonomic hierarchy. The Catalogue holds essential information on the names, relationships and distributions of over 1.3 million species. This figure continues to rise as information is compiled from diverse sources around the world. There are two distinct versions of the Catalogue of Life: the Dynamic Checklist and the Annual Checklist. Choose the version most suited to your needs. If you have a taxonomic database and would like to join the Species 2000 federation of databases in the Catalogue of Life please contact the Species 2000 Secretariat: all candidate databases go through a peer review process. The Annual Checklist Exchange Format defines the format for exchanging data.

View all literature mentions

OrthoMCL DB: Ortholog Groups of Protein Sequences (tool)

RRID:SCR_007839

OrthoMCL is a genome-scale algorithm for grouping orthologous protein sequences. It provides not only groups shared by two or more species/genomes, but also groups representing species-specific gene expansion families. OrthoMCL starts with reciprocal best hits within each genome as putative in-paralog/recent paralog pairs and reciprocal best hits across any two genomes as putative ortholog pairs. Related proteins are interlinked in a similarity graph. Then MCL (Markov Clustering algorithm,Van Dongen 2000; www.micans.org/mcl) is invoked to split mega-clusters. This process is analogous to the manual review in COG construction. MCL clustering is based on weights between each pair of proteins, so to correct for differences in evolutionary distance the weights are normalized before running MCL.

View all literature mentions

SNAP (tool)

RRID:SCR_007936

A sequence analysis tool providing a simple but detailed analysis of human genes and their variations. For each gene, a gene-gene relationship network can be generated based on protein-protein interaction data, metabolic pathway connections and extended through phylogenetic relations. Snap provides tools for designing sequence primers and evaluating RNA splicing effects of single SNPs - known from the databases or defined by you. Primers can be designed for the amplification or sequencing of cDNA, genomic DNA, introns only or exons only.

View all literature mentions

Augustus (tool)

RRID:SCR_008417

Software for gene prediction in eukaryotic genomic sequences. Serves as a basis for further steps in the analysis of sequenced and assembled eukaryotic genomes.

View all literature mentions

ADJUST (tool)

RRID:SCR_009526

A completely automatic algorithm for artifact identification and removal in EEG data. ADJUST is based on Independent Component Analysis (ICA), a successful but unsupervised method for isolating artifacts from EEG recordings. ADJUST identifies artifacted ICA components by combining stereotyped artifact-specific spatial and temporal features. Features are optimised to capture blinks, eye movements and generic discontinuities. Once artifacted IC are identified, they can be simply removed from the data while leaving the activity due to neural sources almost unaffected.

View all literature mentions

MAFFT (tool)

RRID:SCR_011811

Software package as multiple alignment program for amino acid or nucleotide sequences. Can align up to 500 sequences or maximum file size of 1 MB. First version of MAFFT used algorithm based on progressive alignment, in which sequences were clustered with help of Fast Fourier Transform. Subsequent versions have added other algorithms and modes of operation, including options for faster alignment of large numbers of sequences, higher accuracy alignments, alignment of non-coding RNA sequences, and addition of new sequences to existing alignments.

View all literature mentions

MUSCLE (tool)

RRID:SCR_011812

Multiple sequence alignment method with reduced time and space complexity.Multiple sequence alignment with high accuracy and high throughput. Data analysis service for multiple sequence comparison by log- expectation.

View all literature mentions

TBLASTN (tool)

RRID:SCR_011822

Tool to search translated nucleotide databases using a protein query.

View all literature mentions

Trimmomatic (tool)

RRID:SCR_011848

Software Java pipeline for trimming tasks for Illumina paired end and single ended data. Flexible Trimmer for Illumina Sequence Data. Pair aware preprocessing tool optimized for Illumina next generation sequencing data. Includes several processing steps for read trimming and filtering. Operating systems Unix/Linux, Mac OS, Windows.

View all literature mentions

BLAT (tool)

RRID:SCR_011919

Software designed to quickly find sequences of 95% and greater similarity of length 25 bases or more.

View all literature mentions

KEGG (tool)

RRID:SCR_012773

Integrated database resource consisting of 16 main databases, broadly categorized into systems information, genomic information, and chemical information. In particular, gene catalogs in completely sequenced genomes are linked to higher-level systemic functions of cell, organism, and ecosystem. Analysis tools are also available. KEGG may be used as reference knowledge base for biological interpretation of large-scale datasets generated by sequencing and other high-throughput experimental technologies.

View all literature mentions

RepeatMasker (tool)

RRID:SCR_012954

Software tool that screens DNA sequences for interspersed repeats and low complexity DNA sequences. The output of the program is a detailed annotation of the repeats that are present in the query sequence as well as a modified version of the query sequence in which all the annotated repeats have been masked (default: replaced by Ns). Currently over 56% of human genomic sequence is identified and masked by the program. Sequence comparisons in RepeatMasker are performed by one of several popular search engines including nhmmer, cross_match, ABBlast/WUBlast, RMBlast and Decypher. RepeatMasker makes use of curated libraries of repeats and currently supports Dfam ( profile HMM library ) and RepBase ( consensus sequence library ).

View all literature mentions

Trinity (tool)

RRID:SCR_013048

Software for the efficient and robust de novo reconstruction of transcriptomes from RNA-seq data.

View all literature mentions

PhyML (tool)

RRID:SCR_014629

Web phylogeny server based on the maximum-likelihood principle.

View all literature mentions

PASA (tool)

RRID:SCR_014656

Gene structure annotation and analysis tool that uses spliced alignments of expressed transcript sequences to automatically model gene structures. It also incorporates gene structures based on transcript alignments into existing gene structure annotations. It is one component of a larger eukayotic annotation pipeline implemented at the Broad Institute.

View all literature mentions

PAML (tool)

RRID:SCR_014932

Package of programs for phylogenetic analyses of DNA or protein sequences using maximum likelihood. PAML estimates parameters and tests hypotheses to study the evolutionary process from a phylogenetic tree.

View all literature mentions

GeneWise (tool)

RRID:SCR_015054

Gene alignment tool from the EBI which predicts gene structure using similar protein sequences. See also the associated GenomeWise tool.

View all literature mentions

LTR_Finder (tool)

RRID:SCR_015247

Web software capable of scanning large-scale sequences for full-length LTR retrotranspsons.

View all literature mentions

HISAT2 (tool)

RRID:SCR_015530

Graph-based alignment of next generation sequencing reads to a population of genomes.

View all literature mentions

SignalP (tool)

RRID:SCR_015644

Web application for prediction of the presence and location of signal peptide cleavage sites in amino acid sequences from different organisms. The method incorporates a prediction of cleavage sites and a signal peptide/non-signal peptide prediction based on a combination of several artificial neural networks.

View all literature mentions

Canu (tool)

RRID:SCR_015880

Software for scalable and accurate long-read assembly via adaptive k-mer weighting and repeat separation. Canu is a fork of the Celera Assembler and is designed for high-noise single-molecule sequencing (such as the PacBio RS II/Sequel or Oxford Nanopore MinION).

View all literature mentions

GenomeScope (tool)

RRID:SCR_017014

Open source software package for fast genome analysis from unassembled short reads. Used to estimate genome heterozygosity, repeat content, and size from sequencing reads using a kmer-based statistical approach.

View all literature mentions

OrthoFinder (tool)

RRID:SCR_017118

Software Python application for comparative genomics analysis. Finds orthogroups and orthologs, infers rooted gene trees for all orthogroups and identifies all of gene duplcation events in those gene trees, infers rooted species tree for species being analysed and maps gene duplication events from gene trees to branches in species tree, improves orthogroup inference accuracy. Runs set of protein sequence files, one per species, in FASTA format.

View all literature mentions

TGS-GapCloser (tool)

RRID:SCR_017633

Software tool that uses long reads to enhance genome assembly. Fast and accurate gap closing software tool that uses low coverage of error-prone long reads generated by third generation sequence techniques (Pacbio, Oxford Nanopore, etc.) or preassembled contigs for large genomes.

View all literature mentions

TransDecoder (tool)

RRID:SCR_017647

Software tool to identify candidate coding regions within transcript sequences, such as those generated by de novo RNA-Seq transcript assembly using Trinity, or constructed based on RNA-Seq alignments to genome using Tophat and Cufflinks.Starts from FASTA or GFF file. Can scan and retain open reading frames (ORFs) for homology to known proteins by using BlastP or Pfam search and incorporate results into obtained selection. Predictions can then be visualized by using genome browser such as IGV.

View all literature mentions

TimeTree (tool)

RRID:SCR_021162

Public knowledge base for information on evolutionary timescale of life. Data from thousands of published studies are assembled into searchable tree of life scaled to time.

View all literature mentions

BWA-MEM2 (tool)

RRID:SCR_022192

Software tool for sequence mapping.The next version of BWA-MEM. Used for aligning sequencing reads against large reference genome.

View all literature mentions

Tandem Repeats Finder (tool)

RRID:SCR_022193

Software tool to locate and display tandem repeats in DNA sequences. Program to analyze DNA sequences.

View all literature mentions