National Center for Biotechnology Information Archives - Page 2 of 10

September 22, 2019 |

Full-length transcriptome survey and expression analysis of Cassia obtusifolia to discover putative genes related to aurantio-obtusin biosynthesis, seed formation and development, and stress response.

The seed is the pharmaceutical and breeding organ of Cassia obtusifolia, a well-known medical herb containing aurantio-obtusin (a kind of anthraquinone), food, and landscape. In order to understand the molecular mechanism of the biosynthesis of aurantio-obtusin, seed formation and development, and stress response of C. obtusifolia, it is necessary to understand the genomics information. Although previous seed transcriptome of C. obtusifolia has been carried out by short-read next-generation sequencing (NGS) technology, the vast majority of the resulting unigenes did not represent full-length cDNA sequences and supply enough gene expression profile information of the various organs or tissues. In this study, fifteen cDNA libraries, which were constructed from the seed, root, stem, leaf, and flower (three repetitions with each organ) of C. obtusifolia, were sequenced using hybrid approach combining single-molecule real-time (SMRT) and NGS platform. More than 4,315,774 long reads with 9.66 Gb sequencing data and 361,427,021 short reads with 108.13 Gb sequencing data were generated by SMRT and NGS platform, respectively. 67,222 consensus isoforms were clustered from the reads and 81.73% (61,016) of which were longer than 1000 bp. Furthermore, the 67,222 consensus isoforms represented 58,106 nonredundant transcripts, 98.25% (57,092) of which were annotated and 25,573 of which were assigned to specific metabolic pathways by KEGG. CoDXS and CoDXR genes were directly used for functional characterization to validate the accuracy of sequences obtained from transcriptome. A total of 658 seed-specific transcripts indicated their special roles in physiological processes in seed. Analysis of transcripts which were involved in the early stage of anthraquinone biosynthesis suggested that the aurantio-obtusin in C. obtusifolia was mainly generated from isochorismate and Mevalonate/methylerythritol phosphate (MVA/MEP) pathway, and three reactions catalyzed by Menaquinone-specific isochorismate synthase (ICS), 1-deoxy-d-xylulose-5-phosphate synthase (DXS) and isopentenyl diphosphate (IPPS) might be the limited steps. Several seed-specific CYPs, SAM-dependent methyltransferase, and UDP-glycosyltransferase (UDPG) supplied promising candidate genes in the late stage of anthraquinone biosynthesis. In addition, four seed-specific transcriptional factors including three MYB Transcription Factor (MYB) and one MADS-box Transcription Factor (MADS) transcriptional factors) and alternative splicing might be involved with seed formation and development. Meanwhile, most members of Hsp20 genes showed high expression level in seed and flower; seven of which might have chaperon activities under various abiotic stresses. Finally, the expressional patterns of genes with particular interests showed similar trends in both transcriptome assay and qRT-PCR. In conclusion, this is the first full-length transcriptome sequencing reported in Caesalpiniaceae family, and thus providing a more complete insight into aurantio-obtusin biosynthesis, seed formation and development, and stress response as well in C. obtusifolia.

September 22, 2019 |

Isoform sequencing and state-of-art applications for unravelling complexity of plant transcriptomes

Single-molecule real-time (SMRT) sequencing developed by PacBio, also called third-generation sequencing (TGS), offers longer reads than the second-generation sequencing (SGS). Given its ability to obtain full-length transcripts without assembly, isoform sequencing (Iso-Seq) of transcriptomes by PacBio is advantageous for genome annotation, identification of novel genes and isoforms, as well as the discovery of long non-coding RNA (lncRNA). In addition, Iso-Seq gives access to the direct detection of alternative splicing, alternative polyadenylation (APA), gene fusion, and DNA modifications. Such applications of Iso-Seq facilitate the understanding of gene structure, post-transcriptional regulatory networks, and subsequently proteomic diversity. In this review, we summarize its applications in plant transcriptome study, specifically pointing out challenges associated with each step in the experimental design and highlight the development of bioinformatic pipelines. We aim to provide the community with an integrative overview and a comprehensive guidance to Iso-Seq, and thus to promote its applications in plant research.

September 22, 2019 |

Metagenomic SMRT sequencing-based exploration of novel lignocellulose-degrading capability in wood detritus from Torreya nucifera in Bija forest on Jeju Island.

Lignocellulose, mostly composed of cellulose, hemicellulose and lignin generated through secondary growth of woody plant, is considered as promising resources for bio-fuel. In order to use lignocellulose as a biofuel, the biodegradation besides high-cost chemical treatments were applied, but its knowledge on decomposition of lignocellulose occurring in a natural environment were insufficient. We analyzed 16S rRNA gene and metagenome to understand how the lignocellulose are decomposed naturally in decayed Torreya nucifera (L) of Bija forest (Bijarim) in Gotjawal, an ecologically distinct environment. A total of 464,360 reads were obtained from 16S rRNA gene sequencing, representing diverse phyla; Proteobacteria (51%), Bacteroidetes (11%) and Actinobacteria (10%). The metagenome analysis using Single Molecules Real-Time Sequencing revealed that the assembled contigs determined by originated from Proteobacteria (58%) and Actinobacteria (10.3%). Carbohydrate Active enZYmes (CAZy) and Protein families (Pfam) based analysis showed that Proteobacteria was involved in degrading whole lignocellulose and Actinobacteria played a role only in a part of hemicellulose degradation. Combining these results, it suggested that Proteobacteria and Actinobacteria had selective biodegradation potential for different lignocellulose substrate. Thus, it is considered that understanding of the systemic microbial degradation pathways may be a useful strategy for recycle of lignocellulosic biomass and the microbial enzymes in Bija forest can be useful natural resources in industrial processes.

September 22, 2019 |

High-resolution comparative analysis of great ape genomes.

Genetic studies of human evolution require high-quality contiguous ape genome assemblies that are not guided by the human reference. We coupled long-read sequence assembly and full-length complementary DNA sequencing with a multiplatform scaffolding approach to produce ab initio chimpanzee and orangutan genome assemblies. By comparing these with two long-read de novo human genome assemblies and a gorilla genome assembly, we characterized lineage-specific and shared great ape genetic variation ranging from single- to mega-base pair-sized variants. We identified ~17,000 fixed human-specific structural variants identifying genic and putative regulatory changes that have emerged in humans since divergence from nonhuman apes. Interestingly, these variants are enriched near genes that are down-regulated in human compared to chimpanzee cerebral organoids, particularly in cells analogous to radial glial neural progenitors. Copyright © 2018 The Authors, some rights reserved; exclusive licensee American Association for the Advancement of Science. No claim to original U.S. Government Works.

September 22, 2019 |

Transcriptome-referenced association study of clove shape traits in garlic.

Genome-wide association studies are a powerful approach for identifying genes related to complex traits in organisms, but are limited by the requirement for a reference genome sequence of the species under study. To circumvent this problem, we propose a transcriptome-referenced association study (TRAS) that utilizes a transcriptome generated by single-molecule long-read sequencing as a reference sequence to score population variation at both transcript sequence and expression levels. Candidate transcripts are identified when both scores are associated with a trait and their potential interactions are ascertained by expression quantitative trait loci analysis. Applying this method to characterize garlic clove shape traits in 102 landraces, we identified 22 candidate transcripts, most of which showed extensive interactions. Eight transcripts were long non-coding RNAs (lncRNAs), and the others were proteins involved mainly in carbohydrate metabolism, protein degradation, etc. TRAS, as an efficient tool for association study independent of a reference genome, extends the applicability of association studies to a broad range of species.

September 22, 2019 |

Laboratory colonization stabilizes the naturally dynamic microbiome composition of field collected Dermacentor andersoni ticks.

Nearly a quarter of emerging infectious diseases identified in the last century are arthropod-borne. Although ticks and insects can carry pathogenic microorganisms, non-pathogenic microbes make up the majority of their microbial communities. The majority of tick microbiome research has had a focus on discovery and description; very few studies have analyzed the ecological context and functional responses of the bacterial microbiome of ticks. The goal of this analysis was to characterize the stability of the bacterial microbiome of Dermacentor andersoni ticks between generations and two populations within a species.The bacterial microbiome of D. andersoni midguts and salivary glands was analyzed from populations collected at two different ecologically distinct sites by comparing field (F1) and lab-reared populations (F1-F3) over three generations. The microbiome composition of pooled and individual samples was analyzed by sequencing nearly full-length 16S rRNA gene amplicons using a Pacific Biosciences CCS platform that allows identification of bacteria to the species level.In this study, we found that the D. andersoni microbiome was distinct in different geographic populations and was tissue specific, differing between the midgut and the salivary gland, over multiple generations. Additionally, our study showed that the microbiomes of laboratory-reared populations were not necessarily representative of their respective field populations. Furthermore, we demonstrated that the microbiome of a few individual ticks does not represent the microbiome composition at the population level.We demonstrated that the bacterial microbiome of D. andersoni was complex over three generations and specific to tick tissue (midgut vs. salivary glands) as well as geographic location (Burns, Oregon vs. Lake Como, Montana vs. laboratory setting). These results provide evidence that habitat of the tick population is a vital component of the complexity of the bacterial microbiome of ticks, and that the microbiome of lab colonies may not allow for comparative analyses with field populations. A broader understanding of microbiome variation will be required if we are to employ manipulation of the microbiome as a method for interfering with acquisition and transmission of tick-borne pathogens.

September 22, 2019 |

Hybrid sequencing of full-length cDNA transcripts of stems and leaves in Dendrobium officinale.

Dendrobium officinale is an extremely valuable orchid used in traditional Chinese medicine, so sought after that it has a higher market value than gold. Although the expression profiles of some genes involved in the polysaccharide synthesis have previously been investigated, little research has been carried out on their alternatively spliced isoforms in D. officinale. In addition, information regarding the translocation of sugars from leaves to stems in D. officinale also remains limited. We analyzed the polysaccharide content of D. officinale leaves and stems, and completed in-depth transcriptome sequencing of these two diverse tissue types using second-generation sequencing (SGS) and single-molecule real-time (SMRT) sequencing technology. The results of this study yielded a digital inventory of gene and mRNA isoform expressions. A comparative analysis of both transcriptomes uncovered a total of 1414 differentially expressed genes, including 844 that were up-regulated and 570 that were down-regulated in stems. Of these genes, one sugars will eventually be exported transporter (SWEET) and one sucrose transporter (SUT) are expressed to a greater extent in D. officinale stems than in leaves. Two glycosyltransferase (GT) and four cellulose synthase (Ces) genes undergo a distinct degree of alternative splicing. In the stems, the content of polysaccharides is twice as much as that in the leaves. The differentially expressed GT and transcription factor (TF) genes will be the focus of further study. The genes DoSWEET4 and DoSUT1 are significantly expressed in the stem, and are likely to be involved in sugar loading in the phloem.

September 22, 2019 |

Full-length transcriptome sequencing and modular organization analysis of naringin/neoeriocitrin related gene expression pattern in Drynaria roosii.

Drynaria roosii (Nakaike) is a traditional Chinese medicinal fern, known as ‘GuSuiBu’. The effective components, naringin and neoeriocitrin, share a highly similar chemical structure and medicinal function. Our HPLC-tandem mass spectrometry (MS/MS) results showed that the accumulation of naringin/neoeriocitrin depended on specific tissues or ages. However, little was known about the expression patterns of naringin/neoeriocitrin-related genes involved in their regulatory pathways. Due to a lack of basic genetic information, we applied a combination of single molecule real-time (SMRT) sequencing and second-generation sequencing (SGS) to generate the complete and full-length transcriptome of D. roosii. According to the SGS data, the differentially expressed gene (DEG)-based heat map analysis revealed that naringin/neoeriocitrin-related gene expression exhibited obvious tissue- and time-specific transcriptomic differences. Using the systems biology method of modular organization analysis, we clustered 16,472 DEGs into 17 gene modules and studied the relationships between modules and tissue/time point samples, as well as modules and naringin/neoeriocitrin contents. We found that naringin/neoeriocitrin-related DEGs distributed in nine distinct modules, and DEGs in these modules showed significantly different patterns of transcript abundance to be linked to specific tissues or ages. Moreover, weighted gene co-expression network analysis (WGCNA) results further identified that PAL, 4CL and C4H, and C3H and HCT acted as the major hub genes involved in naringin and neoeriocitrin synthesis, respectively, and exhibited high co-expression with MYB- and basic helix-leucine-helix (bHLH)-regulated genes. In this work, modular organization and co-expression networks elucidated the tissue and time specificity of the gene expression pattern, as well as hub genes associated with naringin/neoeriocitrin synthesis in D. roosii. Simultaneously, the comprehensive transcriptome data set provided important genetic information for further research on D. roosii.

September 22, 2019 |

Single-cell (meta-)genomics of a dimorphic Candidatus Thiomargarita nelsonii reveals genomic plasticity.

The genus Thiomargarita includes the world’s largest bacteria. But as uncultured organisms, their physiology, metabolism, and basis for their gigantism are not well understood. Thus, a genomics approach, applied to a single Candidatus Thiomargarita nelsonii cell was employed to explore the genetic potential of one of these enigmatic giant bacteria. The Thiomargarita cell was obtained from an assemblage of budding Ca. T. nelsonii attached to a provannid gastropod shell from Hydrate Ridge, a methane seep offshore of Oregon, USA. Here we present a manually curated genome of Bud S10 resulting from a hybrid assembly of long Pacific Biosciences and short Illumina sequencing reads. With respect to inorganic carbon fixation and sulfur oxidation pathways, the Ca. T. nelsonii Hydrate Ridge Bud S10 genome was similar to marine sister taxa within the family Beggiatoaceae. However, the Bud S10 genome contains genes suggestive of the genetic potential for lithotrophic growth on arsenite and perhaps hydrogen. The genome also revealed that Bud S10 likely respires nitrate via two pathways: a complete denitrification pathway and a dissimilatory nitrate reduction to ammonia pathway. Both pathways have been predicted, but not previously fully elucidated, in the genomes of other large, vacuolated, sulfur-oxidizing bacteria. Surprisingly, the genome also had a high number of unusual features for a bacterium to include the largest number of metacaspases and introns ever reported in a bacterium. Also present, are a large number of other mobile genetic elements, such as insertion sequence (IS) transposable elements and miniature inverted-repeat transposable elements (MITEs). In some cases, mobile genetic elements disrupted key genes in metabolic pathways. For example, a MITE interrupts hupL, which encodes the large subunit of the hydrogenase in hydrogen oxidation. Moreover, we detected a group I intron in one of the most critical genes in the sulfur oxidation pathway, dsrA. The dsrA group I intron also carried a MITE sequence that, like the hupL MITE family, occurs broadly across the genome. The presence of a high degree of mobile elements in genes central to Thiomargarita’s core metabolism has not been previously reported in free-living bacteria and suggests a highly mutable genome.

September 22, 2019 |

Long-term changes of bacterial and viral compositions in the intestine of a recovered Clostridium difficile patient after fecal microbiota transplantation

Fecal microbiota transplantation (FMT) is an effective treatment for recurrent Clostridium difficile infections (RCDIs). However, long-term effects on the patients’ gut microbiota and the role of viruses remain to be elucidated. Here, we characterized bacterial and viral microbiota in the feces of a cured RCDI patient at various time points until 4.5 yr post-FMT compared with the stool donor. Feces were subjected to DNA sequencing to characterize bacteria and double-stranded DNA (dsDNA) viruses including phages. The patient’s microbial communities varied over time and showed little overall similarity to the donor until 7 mo post-FMT, indicating ongoing gut microbiota adaption in this time period. After 4.5 yr, the patient’s bacteria attained donor-like compositions at phylum, class, and order levels with similar bacterial diversity. Differences in the bacterial communities between donor and patient after 4.5 yr were seen at lower taxonomic levels. C. difficile remained undetectable throughout the entire timespan. This demonstrated sustainable donor feces engraftment and verified long-term therapeutic success of FMT on the molecular level. Full engraftment apparently required longer than previously acknowledged, suggesting the implementation of year-long patient follow-up periods into clinical practice. The identified dsDNA viruses were mainly Caudovirales phages. Unexpectedly, sequences related to giant algae–infecting Chlorella viruses were also detected. Our findings indicate that intestinal viruses may be implicated in the establishment of gut microbiota. Therefore, virome analyses should be included in gut microbiota studies to determine the roles of phages and other viruses—such as Chlorella viruses—in human health and disease, particularly during RCDI.

September 22, 2019 |

Comparative genome and transcriptome analysis reveals distinctive surface characteristics and unique physiological potentials of Pseudomonas aeruginosa ATCC 27853.

Pseudomonas aeruginosa ATCC 27853 was isolated from a hospital blood specimen in 1971 and has been widely used as a model strain to survey antibiotics susceptibilities, biofilm development, and metabolic activities of Pseudomonas spp.. Although four draft genomes of P. aeruginosa ATCC 27853 have been sequenced, the complete genome of this strain is still lacking, hindering a comprehensive understanding of its physiology and functional genome.Here we sequenced and assembled the complete genome of P. aeruginosa ATCC 27853 using the Pacific Biosciences SMRT (PacBio) technology and Illumina sequencing platform. We found that accessory genes of ATCC 27853 including prophages and genomic islands (GIs) mainly contribute to the difference between P. aeruginosa ATCC 27853 and other P. aeruginosa strains. Seven prophages were identified within the genome of P. aeruginosa ATCC 27853. Of the predicted 25 GIs, three contain genes that encode monoxoygenases, dioxygenases and hydrolases that could be involved in the metabolism of aromatic compounds. Surveying virulence-related genes revealed that a series of genes that encode the B-band O-antigen of LPS are lacking in ATCC 27853. Distinctive SNPs in genes of cellular adhesion proteins such as type IV pili and flagella biosynthesis were also observed in this strain. Colony morphology analysis confirmed an enhanced biofilm formation capability of ATCC 27853 on solid agar surface compared to Pseudomonas aeruginosa PAO1. We then performed transcriptome analysis of ATCC 27853 and PAO1 using RNA-seq and compared the expression of orthologous genes to understand the functional genome and the genomic details underlying the distinctive colony morphogenesis. These analyses revealed an increased expression of genes involved in cellular adhesion and biofilm maturation such as type IV pili, exopolysaccharide and electron transport chain components in ATCC 27853 compared with PAO1. In addition, distinctive expression profiles of the virulence genes lecA, lasB, quorum sensing regulators LasI/R, and the type I, III and VI secretion systems were observed in the two strains.The complete genome sequence of P. aeruginosa ATCC 27853 reveals the comprehensive genetic background of the strain, and provides genetic basis for several interesting findings about the functions of surface associated proteins, prophages, and genomic islands. Comparative transcriptome analysis of P. aeruginosa ATCC 27853 and PAO1 revealed several classes of differentially expressed genes in the two strains, underlying the genetic and molecular details of several known and yet to be explored morphological and physiological potentials of P. aeruginosa ATCC 27853.

September 22, 2019 |

The state of play in higher eukaryote gene annotation.

A genome sequence is worthless if it cannot be deciphered; therefore, efforts to describe – or ‘annotate’ – genes began as soon as DNA sequences became available. Whereas early work focused on individual protein-coding genes, the modern genomic ocean is a complex maelstrom of alternative splicing, non-coding transcription and pseudogenes. Scientists – from clinicians to evolutionary biologists – need to navigate these waters, and this has led to the design of high-throughput, computationally driven annotation projects. The catalogues that are being produced are key resources for genome exploration, especially as they become integrated with expression, epigenomic and variation data sets. Their creation, however, remains challenging.

September 22, 2019 |

Complete genome sequence of Enterococcus durans Oregon-R-modENCODE strain BDGP3, a lactic acid bacterium found in the Drosophila melanogaster gut

Enterococcus durans Oregon-R-modENCODE strain BDGP3 was isolated from the Drosophila melanogaster gut for functional host-microbe interaction studies. The complete genome is composed of a single circular genome of 2,983,334 bp, with a G+C content of 38%, and a single plasmid of 5,594 bp. Copyright © 2017 Wan et al.

September 22, 2019 |

Assessment of the physicochemical properties and bacterial composition of Lactobacillus plantarum and Enterococcus faecium-fermented Astragalus membranaceus using single molecule, real-time sequencing technology.

We investigated if fermentation with probiotic cultures could improve the production of health-promoting biological compounds in Astragalus membranaceus. We tested the probiotics Enterococcus faecium, Lactobacillus plantarum and Enterococcus faecium?+?Lactobacillus plantarum and applied PacBio single molecule, real-time sequencing technology (SMRT) to evaluate the quality of Astragalus fermentation. We found that the production rates of acetic acid, methylacetic acid, aethyl acetic acid and lactic acid using E. faecium?+?L. plantarum were 1866.24?mg/kg on day 15, 203.80?mg/kg on day 30, 996.04?mg/kg on day 15, and 3081.99?mg/kg on day 20, respectively. Other production rates were: polysaccharides, 9.43%, 8.51%, and 7.59% on day 10; saponins, 19.6912?mg/g, 21.6630?mg/g and 20.2084?mg/g on day 15; and flavonoids, 1.9032?mg/g, 2.0835?mg/g, and 1.7086?mg/g on day 20 using E. faecium, L. plantarum and E. faecium?+?L. plantarum, respectively. SMRT was used to analyze microbial composition, and we found that E. faecium and L. plantarum were the most prevalent species after fermentation for 3 days. E. faecium?+?L. plantarum gave more positive effects than single strains in the Astragalus solid state fermentation process. Our data demonstrated that the SMRT sequencing platform is applicable to quality assessment of Astragalus fermentation.

September 22, 2019 |

The first whole transcriptomic exploration of pre-oviposited early chicken embryos using single and bulked embryonic RNA-sequencing.

The chicken is a valuable model organism, especially in evolutionary and embryology research because its embryonic development occurs in the egg. However, despite its scientific importance, no transcriptome data have been generated for deciphering the early developmental stages of the chicken because of practical and technical constraints in accessing pre-oviposited embryos.Here, we determine the entire transcriptome of pre-oviposited avian embryos, including oocyte, zygote, and intrauterine embryos from Eyal-giladi and Kochav stage I (EGK.I) to EGK.X collected using a noninvasive approach for the first time. We also compare RNA-sequencing data obtained using a bulked embryo sequencing and single embryo/cell sequencing technique. The raw sequencing data were preprocessed with two genome builds, Galgal4 and Galgal5, and the expression of 17,108 and 26,102 genes was quantified in the respective builds. There were some differences between the two techniques, as well as between the two genome builds, and these were affected by the emergence of long intergenic noncoding RNA annotations.The first transcriptome datasets of pre-oviposited early chicken embryos based on bulked and single embryo sequencing techniques will serve as a valuable resource for investigating early avian embryogenesis, for comparative studies among vertebrates, and for novel gene annotation in the chicken genome.

Auto Tag: National Center for Biotechnology Information

Full-length transcriptome survey and expression analysis of Cassia obtusifolia to discover putative genes related to aurantio-obtusin biosynthesis, seed formation and development, and stress response.

Isoform sequencing and state-of-art applications for unravelling complexity of plant transcriptomes

Metagenomic SMRT sequencing-based exploration of novel lignocellulose-degrading capability in wood detritus from Torreya nucifera in Bija forest on Jeju Island.

High-resolution comparative analysis of great ape genomes.

Transcriptome-referenced association study of clove shape traits in garlic.

Laboratory colonization stabilizes the naturally dynamic microbiome composition of field collected Dermacentor andersoni ticks.

Hybrid sequencing of full-length cDNA transcripts of stems and leaves in Dendrobium officinale.

Full-length transcriptome sequencing and modular organization analysis of naringin/neoeriocitrin related gene expression pattern in Drynaria roosii.

Single-cell (meta-)genomics of a dimorphic Candidatus Thiomargarita nelsonii reveals genomic plasticity.

Long-term changes of bacterial and viral compositions in the intestine of a recovered Clostridium difficile patient after fecal microbiota transplantation

Comparative genome and transcriptome analysis reveals distinctive surface characteristics and unique physiological potentials of Pseudomonas aeruginosa ATCC 27853.

The state of play in higher eukaryote gene annotation.

Complete genome sequence of Enterococcus durans Oregon-R-modENCODE strain BDGP3, a lactic acid bacterium found in the Drosophila melanogaster gut

Assessment of the physicochemical properties and bacterial composition of Lactobacillus plantarum and Enterococcus faecium-fermented Astragalus membranaceus using single molecule, real-time sequencing technology.

The first whole transcriptomic exploration of pre-oviposited early chicken embryos using single and bulked embryonic RNA-sequencing.

Subscribe for blog updates:

Filter by topic

Talk with an expert

ALS case study

Subscribe for blog updates:

Filter by topic

Talk with an expert