Searching the Resource Information Network

Our searching services are busy right now. Please try again later

  • Register
X
Forgot Password

If you have forgotten your password you can enter your email here and get a temporary password sent to your email.

X

Leaving Community

Are you sure you want to leave this community? Leaving the community will revoke any permissions you have been granted in this community.

No
Yes

Pangenome and resequencing analyses reveal flowering evolution and genetic control in Cerasus.

Songtao Jiu | Yahui Lei | Linlin Fang | Yue Huang | Wenbo Chen | Yan Xu | Lei Wang | Zhengxin Lv | Xunju Liu | Lu Han | Boyang Liu | Chen Zhang | Xiao Dong | Taoxian Zhang | Jiawei Wang | Xu Zhang | Yuliang Cai | Wei Zheng | Fangdong Li | Congli Liu | Hongwen Li | Jijun Zhu | Fei Yu | Ming Li | Jing Wang | Long Chen | Chunjin Yin | Lei Peng | Jiyuan Wang | Zhuo Zhang | Xiaojuan An | Lixia Yu | Ruie Liu | Yan Zhang | Li Wang | Yaqin Wu | Anthony Bernard | Mitsuho Nakagomi | Shiping Wang | Dirlewanger Elisabeth | Dawei Li | Qi Zhao | Hongzhang Xue | Yang Dong | Zongyi Sun | Shengchang Duan | Caixi Zhang
Nature communications | 2026

Prunus subgenus Cerasus contains numerous species with ornamental, edible, and medicinal value. However, limited genomic resources have constrained systematic analyses of structural variation and the genetic basis of key phenological traits in this group. Here, we assemble eight genomes from diverse Cerasus species. Together with 13 published genomes, we construct a pangenome of 21 accessions representing 17 species. Phenological observations reveal substantial variation in flowering time. Integrating comparative genomics, transcriptomics, and population genetic analyses highlight candidate regulators of flowering time. We find that AGAMOUS-LIKE 9 (AGL9) is strongly associated with flowering progression. Both ectopic expression and transient overexpression of PavAGL9 can accelerate post-dormancy flowering progression. We reveal that PavBPC6 binds the PavAGL9 promoter and represses its transcription, indicating a negative regulatory role. Furthermore, PavAGL9 interacts physically with PavSEP1 and PavPMADS2, suggesting synergistic roles in floral organ development. Our pangenome resource establishes a comprehensive genomic framework for Cerasus and provides insights into the regulation of flowering progression.

Pubmed ID: 42204142

Associated grants

  • Agency: Earmarked Fund for China Agriculture Research System,
    Id: CARS-30
  • Agency: National Natural Science Foundation of China (National Science Foundation of China),
    Id: 32541114
  • Agency: National Natural Science Foundation of China (National Science Foundation of China),
    Id: 32260734

Publication data is provided by the National Library of Medicine ® and PubMed ®. Data is retrieved from PubMed ® on a weekly schedule. For terms and conditions see the National Library of Medicine Terms and Conditions.

This is a list of tools and resources that we have found mentioned in this publication.


VCFtools (tool)

RRID:SCR_001235

Software package for working with VCF files. Used to provide easily accessible methods for working with complex genetic variation data in the form of VCF files.Implements various utilities for processing Variant Call Format files, including validation, merging, comparing. Provides general Perl API.

View all literature mentions

PLINK (tool)

RRID:SCR_001757

Open source whole genome association analysis toolset, designed to perform range of basic, large scale analyses in computationally efficient manner. Used for analysis of genotype/phenotype data. Through integration with gPLINK and Haploview, there is some support for subsequent visualization, annotation and storage of results. PLINK 1.9 is improved and second generation of the software.

View all literature mentions

TAIR (tool)

RRID:SCR_004618

Database of genetic and molecular biology data for the model higher plant Arabidopsis thaliana. Data available includes the complete genome sequence along with gene structure, gene product information, metabolism, gene expression, DNA and seed stocks, genome maps, genetic and physical markers, publications, and information about the Arabidopsis research community. Gene product function data is updated every two weeks from the latest published research literature and community data submissions. Gene structures are updated 1-2 times per year using computational and manual methods as well as community submissions of new and updated genes. TAIR also provides extensive linkouts from data pages to other Arabidopsis resources. The data can be searched, viewed and analyzed. Datasets can also be downloaded. Pages on news, job postings, conference announcements, Arabidopsis lab protocols, and useful links are provided.

View all literature mentions

Pfam (tool)

RRID:SCR_004726

A database of protein families, each represented by multiple sequence alignments and hidden Markov models (HMMs). Users can analyze protein sequences for Pfam matches, view Pfam family annotation and alignments, see groups of related families, look at the domain organization of a protein sequence, find the domains on a PDB structure, and query Pfam by keywords. There are two components to Pfam: Pfam-A and Pfam-B. Pfam-A entries are high quality, manually curated families that may automatically generate a supplement using the ADDA database. These automatically generated entries are called Pfam-B. Although of lower quality, Pfam-B families can be useful for identifying functionally conserved regions when no Pfam-A entries are found. Pfam also generates higher-level groupings of related families, known as clans (collections of Pfam-A entries which are related by similarity of sequence, structure or profile-HMM).

View all literature mentions

Hmmer (tool)

RRID:SCR_005305

Tool for searching sequence databases for homologs of protein sequences, and for making protein sequence alignments. It implements methods using probabilistic models called profile hidden Markov models (profile HMMs). Compared to BLAST, FASTA, and other sequence alignment and database search tools based on older scoring methodology, HMMER aims to be significantly more accurate and more able to detect remote homologs because of the strength of its underlying mathematical models. In the past, this strength came at significant computational expense, but in the new HMMER3 project, HMMER is now essentially as fast as BLAST.

View all literature mentions

InterProScan (tool)

RRID:SCR_005829

Software package for functional analysis of sequences by classifying them into families and predicting presence of domains and sites. Scans sequences against InterPro's signatures. Characterizes nucleotide or protein function by matching it with models from several different databases. Used in large scale analysis of whole proteomes, genomes and metagenomes. Available as Web based version and standalone Perl version and SOAP Web Service.

View all literature mentions

Rfam (tool)

RRID:SCR_007891

The Rfam database is a collection of RNA families, each represented by multiple sequence alignments, consensus secondary structures and covariance models (CMs). The families in Rfam break down into three broad functional classes: Non-coding RNA genes, structured cis-regulatory elements and self-splicing RNAs. Typically these functional RNAs often have a conserved secondary structure which may be better preserved than the RNA sequence. The CMs used to describe each family are a slightly more complicated relative of the profile hidden Markov models (HMMs) used by Pfam. CMs can simultaneously model RNA sequence and the structure in an elegant and accurate fashion. Rfam is also available via FTP. You can find data in Rfam in various ways... * Analyze your RNA sequence for Rfam matches * View Rfam family annotation and alignments * View Rfam clan details * Query Rfam by keywords * Fetch families or sequences by NCBI taxonomy * Enter any type of accession or ID to jump to the page for a Rfam family, sequence or genome

View all literature mentions

Augustus (tool)

RRID:SCR_008417

Software for gene prediction in eukaryotic genomic sequences. Serves as a basis for further steps in the analysis of sequenced and assembled eukaryotic genomes.

View all literature mentions

tRNAscan-SE (tool)

RRID:SCR_008637

Web server to search for tRNA genes in genomic sequence. If you would like to run tRNAscan-SE locally, you can get the UNIX source code (gzip''d tar file).

View all literature mentions

Infernal (tool)

RRID:SCR_011809

Software for searching DNA sequence databases for RNA structure and sequence similarities.

View all literature mentions

MAFFT (tool)

RRID:SCR_011811

Software package as multiple alignment program for amino acid or nucleotide sequences. Can align up to 500 sequences or maximum file size of 1 MB. First version of MAFFT used algorithm based on progressive alignment, in which sequences were clustered with help of Fast Fourier Transform. Subsequent versions have added other algorithms and modes of operation, including options for faster alignment of large numbers of sequences, higher accuracy alignments, alignment of non-coding RNA sequences, and addition of new sequences to existing alignments.

View all literature mentions

MUSCLE (tool)

RRID:SCR_011812

Multiple sequence alignment method with reduced time and space complexity.Multiple sequence alignment with high accuracy and high throughput. Data analysis service for multiple sequence comparison by log- expectation.

View all literature mentions

KEGG (tool)

RRID:SCR_012773

Integrated database resource consisting of 16 main databases, broadly categorized into systems information, genomic information, and chemical information. In particular, gene catalogs in completely sequenced genomes are linked to higher-level systemic functions of cell, organism, and ecosystem. Analysis tools are also available. KEGG may be used as reference knowledge base for biological interpretation of large-scale datasets generated by sequencing and other high-throughput experimental technologies.

View all literature mentions

ANNOVAR (tool)

RRID:SCR_012821

An efficient software tool to utilize update-to-date information to functionally annotate genetic variants detected from diverse genomes (including human genome hg18, hg19, as well as mouse, worm, fly, yeast and many others). Given a list of variants with chromosome, start position, end position, reference nucleotide and observed nucleotides, ANNOVAR can perform: 1. gene-based annotation. 2. region-based annotation. 3. filter-based annotation. 4. other functionalities. (entry from Genetic Analysis Software)

View all literature mentions

TASSEL (tool)

RRID:SCR_012837

Software package which performs a variety of genetic analyses including association mapping, diversity estimation and calculating linkage disequilibrium. The association analysis between genotypes and phenotypes can be performed by either a general linear model or a mixed linear model. The general linear model now allows users to analyze complex field designs, environmental interactions, and epistatic interactions. The mixed model is specially designed to handle polygenic effects at multiple levels of relatedness including pedigree information. These new analyses should permit association analysis in a wide range plant and animal species. (entry from Genetic Analysis Software)

View all literature mentions

BUSCO (tool)

RRID:SCR_015008

Software tool to quantitatively measure genome assembly and annotation completeness based on evolutionarily informed expectations of gene content.

View all literature mentions

HISAT2 (tool)

RRID:SCR_015530

Graph-based alignment of next generation sequencing reads to a population of genomes.

View all literature mentions

StringTie (tool)

RRID:SCR_016323

Software application for assembling of RNA-Seq alignments into potential transcripts. It enables improved reconstruction of a transcriptome from RNA-seq reads. This transcript assembling and quantification program is implemented in C++ .

View all literature mentions

OrthoFinder (tool)

RRID:SCR_017118

Software Python application for comparative genomics analysis. Finds orthogroups and orthologs, infers rooted gene trees for all orthogroups and identifies all of gene duplcation events in those gene trees, infers rooted species tree for species being analysed and maps gene duplication events from gene trees to branches in species tree, improves orthogroup inference accuracy. Runs set of protein sequence files, one per species, in FASTA format.

View all literature mentions

LTR_retriever (tool)

RRID:SCR_017623

Software package for identification of long terminal repeat retrotransposons (LTR-RTs). Removes false positives from initial software predictions. Achieves very high specificity, accuracy, and precision without significantly sacrificing sensitivity, hence significantly outperforming existing methods. Can construct LTR libraries directly from self-corrected PacBio reads prior to genome assembly.

View all literature mentions

iTOL (tool)

RRID:SCR_018174

Web tool for display, annotation and management of phylogenetic trees. Accessible with any modern web browser.

View all literature mentions

Minimap2 (tool)

RRID:SCR_018550

Software tool as pairwise alignment for nucleotide sequences. Alignment program to map DNA or long mRNA sequences against large reference database. Versatile pairwise aligner for genomic and spliced nucleotide sequences.

View all literature mentions

RAxML Next Generation (tool)

RRID:SCR_022066

Software phylogenetic tree inference tool which uses maximum likelihood optimality criterion. Used for maximum likelihood phylogenetic inference. Offers improved accuracy, flexibility, speed, scalability, and usability compared with RAxML/ExaML.

View all literature mentions

TBtools (tool)

RRID:SCR_023018

Software integrative toolkit developed for interactive analyses of big biological data.

View all literature mentions

quarTeT (tool)

RRID:SCR_025258

Web toolkit for studies of large scale T2T genomes. Collection of tools designed for T2T genome assembly and characterization, including reference guided genome assembly, ultra long sequence based gap filling, telomere identification, and de novo centromere prediction. Includes four modules: AssemblyMapper, GapFiller, TeloExplorer, and CentroMiner. Modules can be used alone or in combination with each other for T2T genome assembly and characterization.

View all literature mentions

modeltest (tool)

RRID:SCR_026633

Software tool for selecting the best-fit model of evolution for DNA and protein alignments. Used for selection of DNA and Protein evolutionary models.

View all literature mentions