Read length Archives - Page 8 of 29

September 22, 2019

Single-molecule long-read transcriptome dataset of halophyte Halogeton glomeratus.

Soil salinization has become a major challenge for sustainable development of global agriculture. As a result, cultivation of salt-tolerant crop varieties has become a focus of plant breeding. However, development of effective breeding strategies would be significantly enhanced by improving our understanding of salt tolerance mechanisms in plants and identifying genes required for adaptation.

September 22, 2019

Complete genome sequence of Petrimonas sp. strain IBARAKI, assembled from the metagenome data of a culture containing Dehalococcoides spp.

The complete genome sequence of Petrimonas sp. strain IBARAKI in a Dehalococcoides-containing culture was determined using the PacBio RS II platform. The genome is a single circular chromosome of 3,693,233 nucleotides (nt), with a GC content of 44%. This is the first genome sequence of a Petrimonas species. Copyright © 2018 Ikegami et al.

September 22, 2019

HapIso: An accurate method for the haplotype-specific isoforms reconstruction from long single-molecule reads

Sequencing of RNA provides the possibility to study an individual’s transcriptome landscape and determine allelic expression ratios. Single-molecule protocols generate multi-kilobase reads longer than most transcripts allowing sequencing of complete haplotype isoforms. This allows partitioning the reads into two parental haplotypes. While the read length of the single-molecule protocols is long, the relatively high error rate limits the ability to accurately detect the genetic variants and assemble them into the haplotype-specific isoforms. In this paper, we present HapIso (Haplotype-specific Isoform Reconstruction), a method able to tolerate the relatively high error-rate of the single-molecule platform and partition the isoform reads into the parental alleles. Phasing the reads according to the allele of origin allows our method to efficiently distinguish between the read errors and the true biological mutations. HapIso uses a k-means clustering algorithm aiming to group the reads into two meaningful clusters maximizing the similarity of the reads within cluster and minimizing the similarity of the reads from different clusters. Each cluster corresponds to a parental haplotype. We use family pedigree information to evaluate our approach. Experimental validation suggests that HapIso is able to tolerate the relatively high error-rate and accurately partition the reads into the parental alleles of the isoform transcripts. Furthermore, our method is the first method able to reconstruct the haplotype-specific isoforms from long single-molecule reads. The open source Python implementation of HapIso is freely available for download at https://?github.?com/?smangul1/?HapIso/?.

September 22, 2019

Profiling of metabolome and bacterial community dynamics in ensiled Medicago sativa inoculated without or with Lactobacillus plantarum or Lactobacillus buchneri.

Using gas chromatography mass spectrometry and the PacBio single molecule with real-time sequencing technology (SMRT), we analyzed the detailed metabolomic profiles and microbial community dynamics involved in ensiled Medicago sativa (alfalfa) inoculated without or with the homofermenter Lactobacillus plantarum or heterofermenter Lactobacillus buchneri. Our results revealed that 280 substances and 102 different metabolites were present in ensiled alfalfa. Inoculation of L. buchneri led to remarkable up-accumulation in concentrations of 4-aminobutyric acid, some free amino acids, and polyols in ensiled alfalfa, whereas considerable down-accumulation in cadaverine and succinic acid were observed in L. plantarum-inoculated silages. Completely different microbial flora and their successions during ensiling were observed in the control and two types of inoculant-treated silages. Inoculation of the L. plantarum or L. buchneri alters the microbial composition dynamics of the ensiled forage in very different manners. Our study demonstrates that metabolomic profiling analysis provides a deep insight in metabolites in silage. Moreover, the PacBio SMRT method revealed the microbial composition and its succession during the ensiling process at the species level. This provides information regarding the microbial processes underlying silage formation and may contribute to target-based regulation methods to achieve high-quality silage production.

September 22, 2019

Assessing the gene content of the megagenome: sugar pine (Pinus lambertiana).

Sugar pine (Pinus lambertiana Douglas) is within the subgenus Strobus with an estimated genome size of 31 Gbp. Transcriptomic resources are of particular interest in conifers due to the challenges presented in their megagenomes for gene identification. In this study, we present the first comprehensive survey of the P. lambertiana transcriptome through deep sequencing of a variety of tissue types to generate more than 2.5 billion short reads. Third generation, long reads generated through PacBio Iso-Seq has been included for the first time in conifers to combat the challenges associated with de novo transcriptome assembly. A technology comparison is provided here contribute to the otherwise scarce comparisons of 2nd and 3rd generation transcriptome sequencing approaches in plant species. In addition, the transcriptome reference was essential for gene model identification and quality assessment in the parallel project responsible for sequencing and assembly of the entire genome. In this study, the transcriptomic data was also used to address some of the questions surrounding lineage-specific Dicer-like proteins in conifers. These proteins play a role in the control of transposable element proliferation and the related genome expansion in conifers. Copyright © 2016 Author et al.

September 22, 2019

Lactobacillus fermentum FTDC 8312 combats hypercholesterolemia via alteration of gut microbiota.

In this study, hypercholesterolemic mice fed with Lactobacillus fermentum FTDC 8312 after a seven-week feeding trial showed a reduction in serum total cholesterol (TC) levels, accompanied by a decrease in serum low-density lipoprotein cholesterol (LDL-C) levels, an increase in serum high-density lipoprotein cholesterol (HDL-C) levels, and a decreased ratio of apoB100:apoA1 when compared to those fed with control or a type strain, L. fermentum JCM 1173. These have contributed to a decrease in atherogenic indices (TC/HDL-C) of mice on the FTDC 8312 diet. Serum triglyceride (TG) levels of mice fed with FTDC 8312 and JCM 1173 were comparable to those of the controls. A decreased ratio of cholesterol and phospholipids (C/P) was also observed for mice fed with FTDC 8312, leading to a decreased number of spur red blood cells (RBC) formation in mice. Additionally, there was an increase in fecal TC, TG, and total bile acid levels in mice on FTDC 8312 diet compared to those with JCM 1173 and controls. The administration of FTDC 8312 also altered the gut microbiota population such as an increase in the members of genera Akkermansia and Oscillospira, affecting lipid metabolism and fecal bile excretion in the mice. Overall, we demonstrated that FTDC 8312 exerted a cholesterol lowering effect that may be attributed to gut microbiota modulation. Copyright © 2017 Elsevier B.V. All rights reserved.

September 22, 2019

Long-read isoform sequencing reveals a hidden complexity of the transcriptional landscape of Herpes Simplex Virus Type 1.

In this study, we used the amplified isoform sequencing technique from Pacific Biosciences to characterize the poly(A)(+) fraction of the lytic transcriptome of the herpes simplex virus type 1 (HSV-1). Our analysis detected 34 formerly unidentified protein-coding genes, 10 non-coding RNAs, as well as 17 polycistronic and complex transcripts. This work also led us to identify many transcript isoforms, including 13 splice and 68 transcript end variants, as well as several transcript overlaps. Additionally, we determined previously unascertained transcriptional start and polyadenylation sites. We analyzed the transcriptional activity from the complementary DNA strand in five convergent HSV gene pairs with quantitative RT-PCR and detected antisense RNAs in each gene. This part of the study revealed an inverse correlation between the expressions of convergent partners. Our work adds new insights for understanding the complexity of the pervasive transcriptional overlaps by suggesting that there is a crosstalk between adjacent and distal genes through interaction between their transcription apparatuses. We also identified transcripts overlapping the HSV replication origins, which may indicate an interplay between the transcription and replication machineries. The relative abundance of HSV-1 transcripts has also been established by using a novel method based on the calculation of sequencing reads for the analysis.

September 22, 2019

Iso-Seq analysis of Nepenthes ampullaria, Nepenthes rafflesiana and Nepenthes × hookeriana for hybridisation study in pitcher plants.

Tropical pitcher plants in the species-rich Nepenthaceae family of carnivorous plants possess unique pitcher organs. Hybridisation, natural or artificial, in this family is extensive resulting in pitchers with diverse features. The pitcher functions as a passive insect trap with digestive fluid for nutrient acquisition in nitrogen-poor habitats. This organ shows specialisation according to the dietary habit of different Nepenthes species. In this study, we performed the first single-molecule real-time isoform sequencing (Iso-Seq) analysis of full-length cDNA from Nepenthes ampullaria which can feed on leaf litter, compared to carnivorous Nepenthes rafflesiana, and their carnivorous hybrid Nepenthes × hookeriana. This allows the comparison of pitcher transcriptomes from the parents and the hybrid to understand how hybridisation could shape the evolution of dietary habit in Nepenthes. Raw reads have been deposited to SRA database with the accession numbers SRX2692198 (N. ampullaria), SRX2692197 (N. rafflesiana), and SRX2692196 (N. × hookeriana).

September 22, 2019

Comprehensive profiling of rhizome-associated alternative splicing and alternative polyadenylation in moso bamboo (Phyllostachys edulis).

Moso bamboo (Phyllostachys edulis) represents one of the fastest-spreading plants in the world, due in part to its well-developed rhizome system. However, the post-transcriptional mechanism for the development of the rhizome system in bamboo has not been comprehensively studied. We therefore used a combination of single-molecule long-read sequencing technology and polyadenylation site sequencing (PAS-seq) to re-annotate the bamboo genome, and identify genome-wide alternative splicing (AS) and alternative polyadenylation (APA) in the rhizome system. In total, 145 522 mapped full-length non-chimeric (FLNC) reads were analyzed, resulting in the correction of 2241 mis-annotated genes and the identification of 8091 previously unannotated loci. Notably, more than 42 280 distinct splicing isoforms were derived from 128 667 intron-containing full-length FLNC reads, including a large number of AS events associated with rhizome systems. In addition, we characterized 25 069 polyadenylation sites from 11 450 genes, 6311 of which have APA sites. Further analysis of intronic polyadenylation revealed that LTR/Gypsy and LTR/Copia were two major transposable elements within the intronic polyadenylation region. Furthermore, this study provided a quantitative atlas of poly(A) usage. Several hundred differential poly(A) sites in the rhizome-root system were identified. Taken together, these results suggest that post-transcriptional regulation may potentially have a vital role in the underground rhizome-root system.© 2017 The Authors The Plant Journal © 2017 John Wiley & Sons Ltd.

September 22, 2019

Full-length transcriptome sequencing and modular organization analysis of naringin/neoeriocitrin related gene expression pattern in Drynaria roosii.

Drynaria roosii (Nakaike) is a traditional Chinese medicinal fern, known as ‘GuSuiBu’. The effective components, naringin and neoeriocitrin, share a highly similar chemical structure and medicinal function. Our HPLC-tandem mass spectrometry (MS/MS) results showed that the accumulation of naringin/neoeriocitrin depended on specific tissues or ages. However, little was known about the expression patterns of naringin/neoeriocitrin-related genes involved in their regulatory pathways. Due to a lack of basic genetic information, we applied a combination of single molecule real-time (SMRT) sequencing and second-generation sequencing (SGS) to generate the complete and full-length transcriptome of D. roosii. According to the SGS data, the differentially expressed gene (DEG)-based heat map analysis revealed that naringin/neoeriocitrin-related gene expression exhibited obvious tissue- and time-specific transcriptomic differences. Using the systems biology method of modular organization analysis, we clustered 16,472 DEGs into 17 gene modules and studied the relationships between modules and tissue/time point samples, as well as modules and naringin/neoeriocitrin contents. We found that naringin/neoeriocitrin-related DEGs distributed in nine distinct modules, and DEGs in these modules showed significantly different patterns of transcript abundance to be linked to specific tissues or ages. Moreover, weighted gene co-expression network analysis (WGCNA) results further identified that PAL, 4CL and C4H, and C3H and HCT acted as the major hub genes involved in naringin and neoeriocitrin synthesis, respectively, and exhibited high co-expression with MYB- and basic helix-leucine-helix (bHLH)-regulated genes. In this work, modular organization and co-expression networks elucidated the tissue and time specificity of the gene expression pattern, as well as hub genes associated with naringin/neoeriocitrin synthesis in D. roosii. Simultaneously, the comprehensive transcriptome data set provided important genetic information for further research on D. roosii.

September 22, 2019

Use of a draft genome of coffee (Coffea arabica) to identify SNPs associated with caffeine content.

Arabica coffee (Coffea arabica) has a small gene pool limiting genetic improvement. Selection for caffeine content within this gene pool would be assisted by identification of the genes controlling this important trait. Sequencing of DNA bulks from 18 genotypes with extreme high- or low-caffeine content from a population of 232 genotypes was used to identify linked polymorphisms. To obtain a reference genome, a whole genome assembly of arabica coffee (variety K7) was achieved by sequencing using short read (Illumina) and long-read (PacBio) technology. Assembly was performed using a range of assembly tools resulting in 76 409 scaffolds with a scaffold N50 of 54 544 bp and a total scaffold length of 1448 Mb. Validation of the genome assembly using different tools showed high completeness of the genome. More than 99% of transcriptome sequences mapped to the C. arabica draft genome, and 89% of BUSCOs were present. The assembled genome annotated using AUGUSTUS yielded 99 829 gene models. Using the draft arabica genome as reference in mapping and variant calling allowed the detection of 1444 nonsynonymous single nucleotide polymorphisms (SNPs) associated with caffeine content. Based on Kyoto Encyclopaedia of Genes and Genomes pathway-based analysis, 65 caffeine-associated SNPs were discovered, among which 11 SNPs were associated with genes encoding enzymes involved in the conversion of substrates, which participate in the caffeine biosynthesis pathways. This analysis demonstrated the complex genetic control of this key trait in coffee.© 2018 The Authors. Plant Biotechnology Journal published by Society for Experimental Biology and The Association of Applied Biologists and John Wiley & Sons Ltd.

September 22, 2019

Species groups distributed across elevational gradients reveal convergent and continuous genetic adaptation to high elevations.

Although many cases of genetic adaptations to high elevations have been reported, the processes driving these modifications and the pace of their evolution remain unclear. Many high-elevation adaptations (HEAs) are thought to have arisen in situ as populations rose with growing mountains. In contrast, most high-elevation lineages of the Qinghai-Tibetan Plateau appear to have colonized from low-elevation areas. These lineages provide an opportunity for studying recent HEAs and comparing them with ancestral low-elevation alternatives. Herein, we compare four frogs (three species of Nanorana and a close lowland relative) and four lizards (Phrynocephalus) that inhabit a range of elevations on or along the slopes of the Qinghai-Tibetan Plateau. The sequential cladogenesis of these species across an elevational gradient allows us to examine the gradual accumulation of HEA at increasing elevations. Many adaptations to high elevations appear to arise gradually and evolve continuously with increasing elevational distributions. Numerous related functions, especially DNA repair and energy metabolism pathways, exhibit rapid change and continuous positive selection with increasing elevations. Although the two studied genera are distantly related, they exhibit numerous convergent evolutionary changes, especially at the functional level. This functional convergence appears to be more extensive than convergence at the individual gene level, although we found 32 homologous genes undergoing positive selection for change in both high-elevation groups. We argue that species groups distributed along a broad elevational gradient provide a more powerful system for testing adaptations to high-elevation environments compared with studies that compare only pairs of high-elevation versus low-elevation species.

September 22, 2019

Single-molecule long-read 16S sequencing to characterize the lung microbiome from mechanically ventilated patients with suspected pneumonia.

In critically ill patients, the development of pneumonia results in significant morbidity and mortality and additional health care costs. The accurate and rapid identification of the microbial pathogens in patients with pulmonary infections might lead to targeted antimicrobial therapy with potentially fewer adverse effects and lower costs. Major advances in next-generation sequencing (NGS) allow culture-independent identification of pathogens. The present study used NGS of essentially full-length PCR-amplified 16S ribosomal DNA from the bronchial aspirates of intubated patients with suspected pneumonia. The results from 61 patients demonstrated that sufficient DNA was obtained from 72% of samples, 44% of which (27 samples) yielded PCR amplimers suitable for NGS. Out of the 27 sequenced samples, only 20 had bacterial culture growth, while the microbiological and NGS identification of bacteria coincided in 17 (85%) of these samples. Despite the lack of bacterial growth in 7 samples that yielded amplimers and were sequenced, the NGS identified a number of bacterial species in these samples. Overall, a significant diversity of bacterial species was identified from the same genus as the predominant cultured pathogens. The numbers of NGS-identifiable bacterial genera were consistently higher than identified by standard microbiological methods. As technical advances reduce the processing and sequencing times, NGS-based methods will ultimately be able to provide clinicians with rapid, precise, culture-independent identification of bacterial, fungal, and viral pathogens and their antimicrobial sensitivity profiles. Copyright © 2014, American Society for Microbiology. All Rights Reserved.

September 22, 2019

Multi-platform sequencing approach reveals a novel transcriptome profile in pseudorabies virus.

Third-generation sequencing is an emerging technology that is capable of solving several problems that earlier approaches were not able to, including the identification of transcripts isoforms and overlapping transcripts. In this study, we used long-read sequencing for the analysis of pseudorabies virus (PRV) transcriptome, including Oxford Nanopore Technologies MinION, PacBio RS-II, and Illumina HiScanSQ platforms. We also used data from our previous short-read and long-read sequencing studies for the comparison of the results and in order to confirm the obtained data. Our investigations identified 19 formerly unknown putative protein-coding genes, all of which are 5′ truncated forms of earlier annotated longer PRV genes. Additionally, we detected 19 non-coding RNAs, including 5′ and 3′ truncated transcripts without in-frame ORFs, antisense RNAs, as well as RNA molecules encoded by those parts of the viral genome where no transcription had been detected before. This study has also led to the identification of three complex transcripts and 50 distinct length isoforms, including transcription start and end variants. We also detected 121 novel transcript overlaps, and two transcripts that overlap the replication origins of PRV. Furthermore,in silicoanalysis revealed 145 upstream ORFs, many of which are located on the longer 5′ isoforms of the transcripts.

September 22, 2019

Single-cell (meta-)genomics of a dimorphic Candidatus Thiomargarita nelsonii reveals genomic plasticity.

The genus Thiomargarita includes the world’s largest bacteria. But as uncultured organisms, their physiology, metabolism, and basis for their gigantism are not well understood. Thus, a genomics approach, applied to a single Candidatus Thiomargarita nelsonii cell was employed to explore the genetic potential of one of these enigmatic giant bacteria. The Thiomargarita cell was obtained from an assemblage of budding Ca. T. nelsonii attached to a provannid gastropod shell from Hydrate Ridge, a methane seep offshore of Oregon, USA. Here we present a manually curated genome of Bud S10 resulting from a hybrid assembly of long Pacific Biosciences and short Illumina sequencing reads. With respect to inorganic carbon fixation and sulfur oxidation pathways, the Ca. T. nelsonii Hydrate Ridge Bud S10 genome was similar to marine sister taxa within the family Beggiatoaceae. However, the Bud S10 genome contains genes suggestive of the genetic potential for lithotrophic growth on arsenite and perhaps hydrogen. The genome also revealed that Bud S10 likely respires nitrate via two pathways: a complete denitrification pathway and a dissimilatory nitrate reduction to ammonia pathway. Both pathways have been predicted, but not previously fully elucidated, in the genomes of other large, vacuolated, sulfur-oxidizing bacteria. Surprisingly, the genome also had a high number of unusual features for a bacterium to include the largest number of metacaspases and introns ever reported in a bacterium. Also present, are a large number of other mobile genetic elements, such as insertion sequence (IS) transposable elements and miniature inverted-repeat transposable elements (MITEs). In some cases, mobile genetic elements disrupted key genes in metabolic pathways. For example, a MITE interrupts hupL, which encodes the large subunit of the hydrogenase in hydrogen oxidation. Moreover, we detected a group I intron in one of the most critical genes in the sulfur oxidation pathway, dsrA. The dsrA group I intron also carried a MITE sequence that, like the hupL MITE family, occurs broadly across the genome. The presence of a high degree of mobile elements in genes central to Thiomargarita’s core metabolism has not been previously reported in free-living bacteria and suggests a highly mutable genome.

Auto Tag: Read length

Single-molecule long-read transcriptome dataset of halophyte Halogeton glomeratus.

Complete genome sequence of Petrimonas sp. strain IBARAKI, assembled from the metagenome data of a culture containing Dehalococcoides spp.

HapIso: An accurate method for the haplotype-specific isoforms reconstruction from long single-molecule reads

Profiling of metabolome and bacterial community dynamics in ensiled Medicago sativa inoculated without or with Lactobacillus plantarum or Lactobacillus buchneri.

Assessing the gene content of the megagenome: sugar pine (Pinus lambertiana).

Lactobacillus fermentum FTDC 8312 combats hypercholesterolemia via alteration of gut microbiota.

Long-read isoform sequencing reveals a hidden complexity of the transcriptional landscape of Herpes Simplex Virus Type 1.

Iso-Seq analysis of Nepenthes ampullaria, Nepenthes rafflesiana and Nepenthes × hookeriana for hybridisation study in pitcher plants.

Comprehensive profiling of rhizome-associated alternative splicing and alternative polyadenylation in moso bamboo (Phyllostachys edulis).

Full-length transcriptome sequencing and modular organization analysis of naringin/neoeriocitrin related gene expression pattern in Drynaria roosii.

Use of a draft genome of coffee (Coffea arabica) to identify SNPs associated with caffeine content.

Species groups distributed across elevational gradients reveal convergent and continuous genetic adaptation to high elevations.

Single-molecule long-read 16S sequencing to characterize the lung microbiome from mechanically ventilated patients with suspected pneumonia.

Multi-platform sequencing approach reveals a novel transcriptome profile in pseudorabies virus.

Single-cell (meta-)genomics of a dimorphic Candidatus Thiomargarita nelsonii reveals genomic plasticity.

Subscribe for blog updates:

Filter by topic

Talk with an expert

Antimicrobial resistance research

Subscribe for blog updates:

Filter by topic

Talk with an expert