The Complete Plastid Genome of Rhododendron pulchrum and Comparative Genetic Analysis of Ericaceae Species

Jianshuang Shen; Xueqin Li; Xiangtao Zhu; Xiaoling Huang; Songheng Jin

doi:10.3390/f11020158

The Complete Plastid Genome of Rhododendron pulchrum and Comparative Genetic Analysis of Ericaceae Species

Forests ◽

10.3390/f11020158 ◽

2020 ◽

Vol 11 (2) ◽

pp. 158

Author(s):

Jianshuang Shen ◽

Xueqin Li ◽

Xiangtao Zhu ◽

Xiaoling Huang ◽

Songheng Jin

Keyword(s):

Western Europe ◽

Size Structure ◽

Gc Content ◽

Purifying Selection ◽

Rrna Genes ◽

Biological Research ◽

Genome Comparison ◽

Trna Genes ◽

Protein Coding ◽

Cp Genome

Background and Objectives: Rhododendron pulchrum Sweet (R. pulchrum) belongs to the genus Rhododendron (Ericaceae), a valuable horticultural and medicinal plant species widely used in Western Europe and the US. Despite its importance, this is the first member to have its cpGenome sequenced. Materials and Methods: In this study, the complete cp genome of R. pulchrum was sequenced with NGS Illumina HiSeq2500, analyzed, and compared to eight species in the Ericaceae family. Results: Our study reveals that the cp genome of R. pulchrum is 136,249 bp in length, with an overall GC content of 35.98% and no inverted repeat regions. The R. pulchrum chloroplast genome encodes 73 genes, including 42 protein-coding genes, 29 tRNA genes, and two rRNA genes. The synonymous (Ks) and nonsynonymous (Ka) substitution rates were estimated and the Ka/Ks ratio of R. pulchrum plastid genes were categorized; the results indicated that most of the genes have undergone purifying selection. A total of 382 forward and 259 inverted long repeats, as well as 221 simple-sequence repeat loci (SSR) were detected in the R. pulchrum cp genome. Comparison between different Ericaceae cp genomes revealed significant differences in genome size, structure, and GC content. Conclusions: The phylogenetic relationships among eight Ericaceae species suggested that R. pulchrum is closely related to Vaccinium oldhamii Miq. and Vaccinium macrocarpon Aiton. This study provides a theoretical basis for species identification and future biological research of Rhododendron resources.

Download Full-text

Complete Chloroplast Genomes from Sanguisorba: Identity and Variation Among Four Species

Molecules ◽

10.3390/molecules23092137 ◽

2018 ◽

Vol 23 (9) ◽

pp. 2137 ◽

Cited By ~ 6

Author(s):

Xiang-Xiao Meng ◽

Yan-Fang Xian ◽

Li Xiang ◽

Dong Zhang ◽

Yu-Hua Shi ◽

...

Keyword(s):

Gc Content ◽

Single Copy ◽

Rrna Genes ◽

Trna Genes ◽

Protein Coding ◽

Future Studies ◽

Chloroplast Genomes ◽

Close Relationship ◽

Cp Genome ◽

Sanguisorba Officinalis

The genus Sanguisorba, which contains about 30 species around the world and seven species in China, is the source of the medicinal plant Sanguisorba officinalis, which is commonly used as a hemostatic agent as well as to treat burns and scalds. Here we report the complete chloroplast (cp) genome sequences of four Sanguisorba species (S. officinalis, S. filiformis, S. stipulata, and S. tenuifolia var. alba). These four Sanguisorba cp genomes exhibit typical quadripartite and circular structures, and are 154,282 to 155,479 bp in length, consisting of large single-copy regions (LSC; 84,405–85,557 bp), small single-copy regions (SSC; 18,550–18,768 bp), and a pair of inverted repeats (IRs; 25,576–25,615 bp). The average GC content was ~37.24%. The four Sanguisorba cp genomes harbored 112 different genes arranged in the same order; these identical sections include 78 protein-coding genes, 30 tRNA genes, and four rRNA genes, if duplicated genes in IR regions are counted only once. A total of 39–53 long repeats and 79–91 simple sequence repeats (SSRs) were identified in the four Sanguisorba cp genomes, which provides opportunities for future studies of the population genetics of Sanguisorba medicinal plants. A phylogenetic analysis using the maximum parsimony (MP) method strongly supports a close relationship between S. officinalis and S. tenuifolia var. alba, followed by S. stipulata, and finally S. filiformis. The availability of these cp genomes provides valuable genetic information for future studies of Sanguisorba identification and provides insights into the evolution of the genus Sanguisorba.

Download Full-text

Characterization of the Complete Chloroplast Genome of Buddleja Lindleyana

Journal of AOAC International ◽

10.1093/jaoacint/qsab066 ◽

2021 ◽

Author(s):

Shanshan Liu ◽

Shiyin Feng ◽

Yuying Huang ◽

Wenli An ◽

Zerui Yang ◽

...

Keyword(s):

Gc Content ◽

Single Copy ◽

Rrna Genes ◽

Future Research ◽

Trna Genes ◽

Similar Species ◽

Protein Coding ◽

Genome Data ◽

Cp Genome ◽

Genomic Resource

Abstract Background Buddleja lindleyana Fort., which belongs to the Loganiaceae with a distribution throughout the tropics, is widely used as an ornamental plant in China. Buddleja contains several morphologically similar species, which need to be identified by molecular identification. But there is little molecular research on the genus Buddleja. Objective Using molecular biology techniques to sequence and analyze the complete chloroplast (cp) genome of B. lindleyana Methods According to next-generation sequencing to sequence the genome data, a series of bioinformatics software were used to assembly and analysis the molecular structure of cp genome of B. lindleyana. Results The complete cp genome of B. lindleyana is a circular 154,487-bp-long molecule with a GC content of 38.1%. It has a familiar quadripartite structure, including a large single-copy region (LSC; 85,489 bp), a small single-copy region (SSC; 17,898bp) and a pair of inverted repeats (IRs; 25,550 bp). A total of 133 genes were identified in the genome, including 86 protein-coding genes, 37 tRNA genes, 8 rRNA genes and 2 pseudogenes. Conclusions These results suggested that B. lindelyana cp genome could be used as a potential genomic resource to resolve the phylogenetic positions and relationships of Loganiaceae, and will offer valuable information for future research in the identification of Buddleja species and will conduce to genomic investigations of these species.

Download Full-text

Comparative Analyses of Euonymus Chloroplast Genomes: Genetic Structure, Screening for Loci With Suitable Polymorphism, Positive Selection Genes, and Phylogenetic Relationships Within Celastrineae

Frontiers in Plant Science ◽

10.3389/fpls.2020.593984 ◽

2021 ◽

Vol 11 ◽

Author(s):

Yongtan Li ◽

Yan Dong ◽

Yichao Liu ◽

Xiaoyue Yu ◽

Minsheng Yang ◽

...

Keyword(s):

Positive Selection ◽

Chloroplast Genome ◽

Gc Content ◽

Single Copy ◽

Rrna Genes ◽

Evolutionary Relationships ◽

Trna Genes ◽

Protein Coding ◽

Chloroplast Genomes ◽

Cp Genome

In this study, we assembled and annotated the chloroplast (cp) genome of the Euonymus species Euonymus fortunei, Euonymus phellomanus, and Euonymus maackii, and performed a series of analyses to investigate gene structure, GC content, sequence alignment, and nucleic acid diversity, with the objectives of identifying positive selection genes and understanding evolutionary relationships. The results indicated that the Euonymus cp genome was 156,860–157,611bp in length and exhibited a typical circular tetrad structure. Similar to the majority of angiosperm chloroplast genomes, the results yielded a large single-copy region (LSC) (85,826–86,299bp) and a small single-copy region (SSC) (18,319–18,536bp), separated by a pair of sequences (IRA and IRB; 26,341–26,700bp) with the same encoding but in opposite directions. The chloroplast genome was annotated to 130–131 genes, including 85–86 protein coding genes, 37 tRNA genes, and eight rRNA genes, with GC contents of 37.26–37.31%. The GC content was variable among regions and was highest in the inverted repeat (IR) region. The IR boundary of Euonymus happened expanding resulting that the rps19 entered into IR region and doubled completely. Such fluctuations at the border positions might be helpful in determining evolutionary relationships among Euonymus. The simple-sequence repeats (SSRs) of Euonymus species were composed primarily of single nucleotides (A)n and (T)n, and were mostly 10–12bp in length, with an obvious A/T bias. We identified several loci with suitable polymorphism with the potential use as molecular markers for inferring the phylogeny within the genus Euonymus. Signatures of positive selection were seen in rpoB protein encoding genes. Based on data from the whole chloroplast genome, common single copy genes, and the LSC, SSC, and IR regions, we constructed an evolutionary tree of Euonymus and related species, the results of which were consistent with traditional taxonomic classifications. It showed that E. fortunei sister to the Euonymus japonicus, whereby E. maackii appeared as sister to Euonymus hamiltonianus. Our study provides important genetic information to support further investigations into the phylogenetic development and adaptive evolution of Euonymus species.

Download Full-text

Comprehensive Analysis of Rhodomyrtus tomentosa Chloroplast Genome

Plants ◽

10.3390/plants8040089 ◽

2019 ◽

Vol 8 (4) ◽

pp. 89 ◽

Cited By ~ 7

Author(s):

Yuying Huang ◽

Zerui Yang ◽

Song Huang ◽

Wenli An ◽

Jing Li ◽

...

Keyword(s):

Gc Content ◽

Single Copy ◽

Rrna Genes ◽

Trna Genes ◽

Sister Relationship ◽

Protein Coding ◽

Protein Coding Genes ◽

Plastid Genomes ◽

Cp Genome ◽

Rhodomyrtus Tomentosa

In the last decade, several studies have relied on a small number of plastid genomes to deduce deep phylogenetic relationships in the species-rich Myrtaceae. Nevertheless, the plastome of Rhodomyrtus tomentosa, an important representative plant of the Rhodomyrtus (DC.) genera, has not yet been reported yet. Here, we sequenced and analyzed the complete chloroplast (CP) genome of R. tomentosa, which is a 156,129-bp-long circular molecule with 37.1% GC content. This CP genome displays a typical quadripartite structure with two inverted repeats (IRa and IRb), of 25,824 bp each, that are separated by a small single copy region (SSC, 18,183 bp) and one large single copy region (LSC, 86,298 bp). The CP genome encodes 129 genes, including 84 protein-coding genes, 37 tRNA genes, eight rRNA genes and three pseudogenes (ycf1, rps19, ndhF). A considerable number of protein-coding genes have a universal ATG start codon, except for psbL and ndhD. Premature termination codons (PTCs) were found in one protein-coding gene, namely atpE, which is rarely reported in the CP genome of plants. Phylogenetic analysis revealed that R. tomentosa has a sister relationship with Eugenia uniflora and Psidium guajava. In conclusion, this study identified unique characteristics of the R. tomentosa CP genome providing valuable information for further investigations on species identification and the phylogenetic evolution between R. tomentosa and related species.

Download Full-text

Complete Chloroplast Genome Sequence of Sonchus brachyotus Helps to Elucidate Evolutionary Relationships with Related Species of Asteraceae

BioMed Research International ◽

10.1155/2021/9410496 ◽

2021 ◽

Vol 2021 ◽

pp. 1-13

Author(s):

Caixiang Wang ◽

Juanjuan Liu ◽

Yue Su ◽

Meili Li ◽

Xiaoyu Xie ◽

...

Keyword(s):

Related Species ◽

Phylogenetic Analyses ◽

Gc Content ◽

Rrna Genes ◽

Trna Genes ◽

Evolutionary Patterns ◽

Protein Coding ◽

Variable Regions ◽

Ssr Frequency ◽

Cp Genome

Sonchus brachyotus DC. possesses both edible and medicinal properties and is widely distributed throughout China. In this study, the complete cp genome of S. brachyotus was sequenced and assembled. The total length of the complete S. brachyotus cp genome was 151,977 bp, including an LSC region of 84,553 bp, SSC region of 18,138 bp, and IR region of 24,643 bp. Sequence analyses revealed that the cp genome encoded 132 genes, including 87 protein-coding genes, 37 tRNA genes, and 8 rRNA genes. The GC content was 37.6%. One hundred mononucleotide microsatellites, 4 dinucleotide microsatellites, 67 trinucleotide microsatellites, 4 tetranucleotide microsatellites, and 1 long repeat were identified. The SSR frequency of the LSC region was significantly greater than that of the IR and SSC regions. In total, 175 SSRs and highly variable regions were recognized as potential cp markers. By analyzing the IR/LSC and IR/SSC boundaries, structural differences between S. brachyotus and 6 other species were detected. According to phylogenetic analyses, S. brachyotus was most closely related to S. arvensis and S. oleraceus. Overall, this study provides complete cp genome resources for S. brachyotus that will be beneficial for identifying potential molecular markers and evolutionary patterns of S. brachyotus and its closely related species.

Download Full-text

The complete chloroplast genome ofColobanthus apetalus(Labill.) Druce: genome organization and comparison with related species

PeerJ ◽

10.7717/peerj.4723 ◽

2018 ◽

Vol 6 ◽

pp. e4723 ◽

Cited By ~ 3

Author(s):

Piotr Androsiuk ◽

Jan Paweł Jastrzębski ◽

Łukasz Paukszto ◽

Adam Okorski ◽

Agnieszka Pszczółkowska ◽

...

Keyword(s):

Gc Content ◽

Large Family ◽

Detailed Comparison ◽

Single Copy ◽

Rrna Genes ◽

Trna Genes ◽

Protein Coding ◽

Protein Coding Genes ◽

Cp Genome ◽

Unique Genes

Colobanthus apetalusis a member of the genusColobanthus, one of the 86 genera of the large family Caryophyllaceae which groups annual and perennial herbs (rarely shrubs) that are widely distributed around the globe, mainly in the Holarctic. The genusColobanthusconsists of 25 species, includingColobanthus quitensis, an extremophile plant native to the maritime Antarctic. Complete chloroplast (cp) genomes are useful for phylogenetic studies and species identification. In this study, next-generation sequencing (NGS) was used to identify the cp genome ofC. apetalus.The complete cp genome ofC. apetalushas the length of 151,228 bp, 36.65% GC content, and a quadripartite structure with a large single copy (LSC) of 83,380 bp and a small single copy (SSC) of 17,206 bp separated by inverted repeats (IRs) of 25,321 bp. The cp genome contains 131 genes, including 112 unique genes and 19 genes which are duplicated in the IRs. The group of 112 unique genes features 73 protein-coding genes, 30 tRNA genes, four rRNA genes and five conserved chloroplast open reading frames (ORFs). A total of 12 forward repeats, 10 palindromic repeats, five reverse repeats and three complementary repeats were detected. In addition, a simple sequence repeat (SSR) analysis revealed 41 (mono-, di-, tri-, tetra-, penta- and hexanucleotide) SSRs, most of which were AT-rich. A detailed comparison ofC. apetalusandC. quitensiscp genomes revealed identical gene content and order. A phylogenetic tree was built based on the sequences of 76 protein-coding genes that are shared by the eleven sequenced representatives of Caryophyllaceae andC. apetalus,and it revealed thatC. apetalusandC. quitensisform a clade that is closely related toSilenespecies andAgrostemma githago. Moreover, the genusSileneappeared as a polymorphic taxon. The results of this study expand our knowledge about the evolution and molecular biology of Caryophyllaceae.

Download Full-text

Complete Chloroplast Genome of Argania spinosa: Structural Organization and Phylogenetic Relationships in Sapotaceae

Plants ◽

10.3390/plants9101354 ◽

2020 ◽

Vol 9 (10) ◽

pp. 1354

Author(s):

Slimane Khayi ◽

Fatima Gaboun ◽

Stacy Pirro ◽

Tatiana Tatusova ◽

Abdelhamid El Mousadik ◽

...

Keyword(s):

Chloroplast Genome ◽

Single Copy ◽

Rrna Genes ◽

Trna Genes ◽

Protein Coding ◽

Important Species ◽

Complete Chloroplast Genome ◽

Argania Spinosa ◽

Protein Coding Genes ◽

Cp Genome

Argania spinosa (Sapotaceae), an important endemic Moroccan oil tree, is a primary source of argan oil, which has numerous dietary and medicinal proprieties. The plant species occupies the mid-western part of Morocco and provides great environmental and socioeconomic benefits. The complete chloroplast (cp) genome of A. spinosa was sequenced, assembled, and analyzed in comparison with those of two Sapotaceae members. The A. spinosa cp genome is 158,848 bp long, with an average GC content of 36.8%. The cp genome exhibits a typical quadripartite and circular structure consisting of a pair of inverted regions (IR) of 25,945 bp in length separating small single-copy (SSC) and large single-copy (LSC) regions of 18,591 and 88,367 bp, respectively. The annotation of A. spinosa cp genome predicted 130 genes, including 85 protein-coding genes (CDS), 8 ribosomal RNA (rRNA) genes, and 37 transfer RNA (tRNA) genes. A total of 44 long repeats and 88 simple sequence repeats (SSR) divided into mononucleotides (76), dinucleotides (7), trinucleotides (3), tetranucleotides (1), and hexanucleotides (1) were identified in the A. spinosa cp genome. Phylogenetic analyses using the maximum likelihood (ML) method were performed based on 69 protein-coding genes from 11 species of Ericales. The results confirmed the close position of A. spinosa to the Sideroxylon genus, supporting the revisiting of its taxonomic status. The complete chloroplast genome sequence will be valuable for further studies on the conservation and breeding of this medicinally and culinary important species and also contribute to clarifying the phylogenetic position of the species within Sapotaceae.

Download Full-text

The complete mitochondrial genome sequence of Scolopendra mutilans L. Koch, 1878 (Scolopendromorpha, Scolopendridae), with a comparative analysis of other centipede genomes

ZooKeys ◽

10.3897/zookeys.925.47820 ◽

2020 ◽

Vol 925 ◽

pp. 73-88

Author(s):

Chaoyi Hu ◽

Shuaibin Wang ◽

Bisheng Huang ◽

Hegang Liu ◽

Lei Xu ◽

...

Keyword(s):

Mitochondrial Genome ◽

Complete Mitochondrial Genome ◽

Gc Content ◽

Sister Group ◽

Rrna Genes ◽

Sister Group Relationship ◽

Trna Genes ◽

Protein Coding ◽

Generation Sequencing ◽

Simple Sequence

Scolopendra mutilans L. Koch, 1878 is an important Chinese animal with thousands of years of medicinal history. However, the genomic information of this species is limited, which hinders its further application. Here, the complete mitochondrial genome (mitogenome) of S. mutilans was sequenced and assembled by next-generation sequencing. The genome is 15,011 bp in length, consisting of 13 protein-coding genes (PCGs), 14 tRNA genes, and two rRNA genes. Most PCGs start with the ATN initiation codon, and all PCGs have the conventional stop codons TAA and TAG. The S. mutilans mitogenome revealed nine simple sequence repeats (SSRs), and an obviously lower GC content compared with other seven centipede mitogenomes previously sequenced. After analysis of homologous regions between the eight centipede mitogenomes, the S. mutilans mitogenome further showed clear genomic rearrangements. The phylogenetic analysis of eight centipedes using 13 conserved PCG genes was finally performed. The phylogenetic reconstructions showed Scutigeromorpha as a separate group, and Scolopendromorpha in a sister-group relationship with Lithobiomorpha and Geophilomorpha. Collectively, the S. mutilans mitogenome provided new genomic resources, which will improve its medicinal research and applications in the future.

Download Full-text

Characterization of the complete chloroplast genome sequence and phylogenetic analysis of B. oleracea var. italica

10.21203/rs.2.20976/v1 ◽

2020 ◽

Author(s):

Zhenchao Zhang ◽

Zhongliang Dai ◽

Yuemei Yao ◽

Yongfei Pan ◽

Guosheng Sun ◽

...

Keyword(s):

Chloroplast Genome ◽

Genome Sequence ◽

Genomic Structure ◽

Gc Content ◽

Single Copy ◽

Biological Research ◽

Protein Coding ◽

Protein Coding Genes ◽

Cp Genome ◽

Functional Components

Abstract Backgrounds: Broccoli (Brassica. oleracea var. italica L.) is known as one of the most nutritionally rich vegetables, as well as rich in functional components that benefit to health. The main purposes of this research were sequencing, assembling and annotation of chloroplast genome of broccoli based on Illumina HiSeq2500 sequencing platform. Results: The size of the broccoli cp genome is 153,364 bp, including two inverted repeat (IR) regions of 26,197 bp each, separated by a small single copy (SSC) region of 17,834 bp and a large single copy (LSC) region of 83,136 bp. The GC content of the complete genome is 36.36%, while those of SSC, LSC, and IR are 29.1%, 34.15% and 42.35%, respectively. It harbors 134 functional genes, including 87 protein-coding genes, 39 tRNAs and 8 rRNAs, with 31 duplicates in the IRs. The most abundant amino acid in the protein-coding genes is leucine, while the least is cysteine. Codon usage frequency showed bias for A/T-ending codons in the cp genome. In the repeat structure analysis, a total of 34 repeat sequences and 291 simple sequence repeat (SSRs) were detected in the work. Although cp genomic structure and size are highly conserved, the SC-IR boundary regions are variable between the 7 cp genomes. The phylogenetic relationships based on complete cp genome from 9 species suggest that B. oleracea var. italica is closely related to Brassica juncea. Conclusions: The complete cp genome sequence was obtained and annotated for broccoli for the first time. The information acquired from this research will be useful for further species identification, population genetics and biological research of broccoli.

Download Full-text

Draft Genome Sequence of Bacillus pacificus KVCMST-8A-12, Isolated from a Marine Sediment Sample from the Kanyakumari Coast, India

Microbiology Resource Announcements ◽

10.1128/mra.01011-21 ◽

2021 ◽

Vol 10 (50) ◽

Author(s):

Asha Santhi ◽

Venkatesh Subramanian ◽

Krishnaveni Muthan

Keyword(s):

Genome Sequence ◽

Draft Genome ◽

Gc Content ◽

Draft Genome Sequence ◽

Whole Genome Sequence ◽

Rrna Genes ◽

Trna Genes ◽

Protein Coding ◽

Content Type ◽

Marine Sediment Sample

A DNase-producing Bacillus pacificus strain was isolated, and the whole-genome sequence is reported in this paper. The draft genome sequence of Bacillus pacificus KVCMST-8A-12 constitutes 2.4 Gbp of raw reads, with a GC content of 35.24%. In total, 5,661 protein-coding genes, 64 tRNA genes, and 4 rRNA genes were predicted.

Download Full-text