Bioinformatics Archives - Page 48 of 267

September 22, 2019

Influenza virus infection causes global RNAPII termination defects.

Viral infection perturbs host cells and can be used to uncover regulatory mechanisms controlling cellular responses and susceptibility to infections. Using cell biological, biochemical, and genetic tools, we reveal that influenza A virus (IAV) infection induces global transcriptional defects at the 3′ ends of active host genes and RNA polymerase II (RNAPII) run-through into extragenic regions. Deregulated RNAPII leads to expression of aberrant RNAs (3′ extensions and host-gene fusions) that ultimately cause global transcriptional downregulation of physiological transcripts, an effect influencing antiviral response and virulence. This phenomenon occurs with multiple strains of IAV, is dependent on influenza NS1 protein, and can be modulated by SUMOylation of an intrinsically disordered region (IDR) of NS1 expressed by the 1918 pandemic IAV strain. Our data identify a strategy used by IAV to suppress host gene expression and indicate that polymorphisms in IDRs of viral proteins can affect the outcome of an infection.

September 22, 2019

A comprehensive quality evaluation system for complex herbal medicine using PacBio sequencing, PCR-denaturing gradient gel electrophoresis, and several chemical approaches.

Herbal medicine is a major component of complementary and alternative medicine, contributing significantly to the health of many people and communities. Quality control of herbal medicine is crucial to ensure that it is safe and sound for use. Here, we investigated a comprehensive quality evaluation system for a classic herbal medicine, Danggui Buxue Formula, by applying genetic-based and analytical chemistry approaches to authenticate and evaluate the quality of its samples. For authenticity, we successfully applied two novel technologies, third-generation sequencing and PCR-DGGE (denaturing gradient gel electrophoresis), to analyze the ingredient composition of the tested samples. For quality evaluation, we used high performance liquid chromatography assays to determine the content of chemical markers to help estimate the dosage relationship between its two raw materials, plant roots of Huangqi and Danggui. A series of surveys were then conducted against several exogenous contaminations, aiming to further access the efficacy and safety of the samples. In conclusion, the quality evaluation system demonstrated here can potentially address the authenticity, quality, and safety of herbal medicines, thus providing novel insight for enhancing their overall quality control. Highlight: We established a comprehensive quality evaluation system for herbal medicine, by combining two genetic-based approaches third-generation sequencing and DGGE (denaturing gradient gel electrophoresis) with analytical chemistry approaches to achieve the authentication and quality connotation of the samples.

September 22, 2019

Antagonism between Staphylococcus epidermidis and Propionibacterium acnes and its genomic basis.

Propionibacterium acnes and Staphylococcus epidermidis live in close proximity on human skin, and both bacterial species can be isolated from normal and acne vulgaris-affected skin sites. The antagonistic interactions between the two species are poorly understood, as well as the potential significance of bacterial interferences for the skin microbiota. Here, we performed simultaneous antagonism assays to detect inhibitory activities between multiple isolates of the two species. Selected strains were sequenced to identify the genomic basis of their antimicrobial phenotypes.First, we screened 77 P. acnes strains isolated from healthy and acne-affected skin, and representing all known phylogenetic clades (I, II, and III), for their antimicrobial activities against 12?S. epidermidis isolates. One particular phylogroup (I-2) exhibited a higher antimicrobial activity than other P. acnes phylogroups. All genomes of type I-2 strains carry an island encoding the biosynthesis of a thiopeptide with possible antimicrobial activity against S. epidermidis. Second, 20?S. epidermidis isolates were examined for inhibitory activity against 25 P. acnes strains. The majority of S. epidermidis strains were able to inhibit P. acnes. Genomes of S. epidermidis strains with strong, medium and no inhibitory activities against P. acnes were sequenced. Genome comparison underlined the diversity of S. epidermidis and detected multiple clade- or strain-specific mobile genetic elements encoding a variety of functions important in antibiotic and stress resistance, biofilm formation and interbacterial competition, including bacteriocins such as epidermin. One isolate with an extraordinary antimicrobial activity against P. acnes harbors a functional ESAT-6 secretion system that might be involved in the antimicrobial activity against P. acnes via the secretion of polymorphic toxins.Taken together, our study suggests that interspecies interactions could potentially jeopardize balances in the skin microbiota. In particular, S. epidermidis strains possess an arsenal of different mechanisms to inhibit P. acnes. However, if such interactions are relevant in skin disorders such as acne vulgaris remains questionable, since no difference in the antimicrobial activity against, or the sensitivity towards S. epidermidis could be detected between health- and acne-associated strains of P. acnes.

September 22, 2019

Contrasting distribution patterns between aquatic and terrestrial Phytophthora species along a climatic gradient are linked to functional traits.

Diversity of microbial organisms is linked to global climatic gradients. The genus Phytophthora includes both aquatic and terrestrial plant pathogenic species that display a large variation of functional traits. The extent to which the physical environment (water or soil) modulates the interaction of microorganisms with climate is unknown. Here, we explored the main environmental drivers of diversity and functional trait composition of Phytophthora communities. Communities were obtained by a novel metabarcoding setup based on PacBio sequencing of river filtrates in 96 river sites along a geographical gradient. Species were classified as terrestrial or aquatic based on their phylogenetic clade. Overall, terrestrial and aquatic species showed contrasting patterns of diversity. For terrestrial species, precipitation was a stronger driver than temperature, and diversity and functional diversity decreased with decreasing temperature and precipitation. In cold and dry areas, the dominant species formed resistant structures and had a low optimum temperature. By contrast, for aquatic species, temperature and water chemistry were the strongest drivers, and diversity increased with decreasing temperature and precipitation. Within the same area, environmental filtering affected terrestrial species more strongly than aquatic species (20% versus 3% of the studied communities, respectively). Our results highlight the importance of functional traits and the physical environment in which microorganisms develop their life cycle when predicting their distribution under changing climatic conditions. Temperature and rainfall may be buffered differently by water and soil, and thus pose contrasting constrains to microbial assemblies.

September 22, 2019

cDNA library enrichment of full length transcripts for SMRT long read sequencing.

The utility of genome assemblies does not only rely on the quality of the assembled genome sequence, but also on the quality of the gene annotations. The Pacific Biosciences Iso-Seq technology is a powerful support for accurate eukaryotic gene model annotation as it allows for direct readout of full-length cDNA sequences without the need for noisy short read-based transcript assembly. We propose the implementation of the TeloPrime Full Length cDNA Amplification kit to the Pacific Biosciences Iso-Seq technology in order to enrich for genuine full-length transcripts in the cDNA libraries. We provide evidence that TeloPrime outperforms the commonly used SMARTer PCR cDNA Synthesis Kit in identifying transcription start and end sites in Arabidopsis thaliana. Furthermore, we show that TeloPrime-based Pacific Biosciences Iso-Seq can be successfully applied to the polyploid genome of bread wheat (Triticum aestivum) not only to efficiently annotate gene models, but also to identify novel transcription sites, gene homeologs, splicing isoforms and previously unidentified gene loci.

September 22, 2019

Gene activity in primary T cells infected with HIV89.6: intron retention and induction of genomic repeats.

HIV infection has been reported to alter cellular gene activity, but published studies have commonly assayed transformed cell lines and lab-adapted HIV strains, yielding inconsistent results. Here we carried out a deep RNA-Seq analysis of primary human T cells infected with the low passage HIV isolate HIV89.6.Seventeen percent of cellular genes showed altered activity 48 h after infection. In a meta-analysis including four other studies, our data differed from studies of HIV infection in cell lines but showed more parallels with infections of primary cells. We found a global trend toward retention of introns after infection, suggestive of a novel cellular response to infection. HIV89.6 infection was also associated with activation of several human endogenous retroviruses (HERVs) and retrotransposons, of interest as possible novel antigens that could serve as vaccine targets. The most highly activated group of HERVs was a subset of the ERV-9. Analysis showed that activation was associated with a particular variant of ERV-9 long terminal repeats that contains an indel near the U3-R border. These data also allowed quantification of >70 splice forms of the HIV89.6 RNA and specified the main types of chimeric HIV89.6-host RNAs. Comparison to over 100,000 integration site sequences from the same infected cell populations allowed quantification of authentic versus artifactual chimeric reads, showing that 5′ read-in, splicing out of HIV89.6 from the D4 donor and 3′ read-through were the most common HIV89.6-host cell chimeric RNA forms.Analysis of RNA abundance after infection of primary T cells with the low passage HIV89.6 isolate disclosed multiple novel features of HIV-host interactions, notably intron retention and induction of transcription of retrotransposons and endogenous retroviruses.

September 22, 2019

Chromosome-level reference genome and alternative splicing atlas of moso bamboo (Phyllostachys edulis).

Bamboo is one of the most important nontimber forestry products worldwide. However, a chromosome-level reference genome is lacking, and an evolutionary view of alternative splicing (AS) in bamboo remains unclear despite emerging omics data and improved technologies.Here, we provide a chromosome-level de novo genome assembly of moso bamboo (Phyllostachys edulis) using additional abundance sequencing data and a Hi-C scaffolding strategy. The significantly improved genome is a scaffold N50 of 79.90 Mb, approximately 243 times longer than the previous version. A total of 51,074 high-quality protein-coding loci with intact structures were identified using single-molecule real-time sequencing and manual verification. Moreover, we provide a comprehensive AS profile based on the identification of 266,711 unique AS events in 25,225 AS genes by large-scale transcriptomic sequencing of 26 representative bamboo tissues using both the Illumina and Pacific Biosciences sequencing platforms. Through comparisons with orthologous genes in related plant species, we observed that the AS genes are concentrated among more conserved genes that tend to accumulate higher transcript levels and share less tissue specificity. Furthermore, gene family expansion, abundant AS, and positive selection were identified in crucial genes involved in the lignin biosynthetic pathway of moso bamboo.These fundamental studies provide useful information for future in-depth analyses of comparative genome and AS features. Additionally, our results highlight a global perspective of AS during evolution and diversification in bamboo.

September 22, 2019

Using PacBio long-read high-throughput microbial gene amplicon sequencing to evaluate infant formula safety.

Infant formula (IF) requires a strict microbiological standard because of the high vulnerability of infants to foodborne diseases. The current study used the PacBio single molecule real-time (SMRT) sequencing platform to generate full-length 16S rRNA-based bacterial microbiota profiles of thirty Chinese domestic and imported IF samples. A total of 600 species were identified, dominated by Streptococcus thermophilus, Lactococcus lactis and Lactococcus piscium. Distinctive bacterial profiles were observed between the two sample groups, as confirmed with both principal coordinate analysis and multivariate analysis of variance. Moreover, the product whey protein nitrogen index (WPNI), representing the degree of preheating, negatively correlated with the relative abundances of the Bacillus genus. Our study has demonstrated the application of the PacBio SMRT sequencing platform in assessing the bacterial contamination of IF products, which is of interest to the dairy industry for effective monitoring of microbial quality and safety during production.

September 22, 2019

Gaining comprehensive biological insight into the transcriptome by performing a broad-spectrum RNA-seq analysis.

RNA-sequencing (RNA-seq) is an essential technique for transcriptome studies, hundreds of analysis tools have been developed since it was debuted. Although recent efforts have attempted to assess the latest available tools, they have not evaluated the analysis workflows comprehensively to unleash the power within RNA-seq. Here we conduct an extensive study analysing a broad spectrum of RNA-seq workflows. Surpassing the expression analysis scope, our work also includes assessment of RNA variant-calling, RNA editing and RNA fusion detection techniques. Specifically, we examine both short- and long-read RNA-seq technologies, 39 analysis tools resulting in ~120 combinations, and ~490 analyses involving 15 samples with a variety of germline, cancer and stem cell data sets. We report the performance and propose a comprehensive RNA-seq analysis protocol, named RNACocktail, along with a computational pipeline achieving high accuracy. Validation on different samples reveals that our proposed protocol could help researchers extract more biologically relevant predictions by broad analysis of the transcriptome.RNA-seq is widely used for transcriptome analysis. Here, the authors analyse a wide spectrum of RNA-seq workflows and present a comprehensive analysis protocol named RNACocktail as well as a computational pipeline leveraging the widely used tools for accurate RNA-seq analysis.

September 22, 2019

Distinguishing highly similar gene isoforms with a clustering-based bioinformatics analysis of PacBio single-molecule long reads.

Gene isoforms are commonly found in both prokaryotes and eukaryotes. Since each isoform may perform a specific function in response to changing environmental conditions, studying the dynamics of gene isoforms is important in understanding biological processes and disease conditions. However, genome-wide identification of gene isoforms is technically challenging due to the high degree of sequence identity among isoforms. Traditional targeted sequencing approach, involving Sanger sequencing of plasmid-cloned PCR products, has low throughput and is very tedious and time-consuming. Next-generation sequencing technologies such as Illumina and 454 achieve high throughput but their short read lengths are a critical barrier to accurate assembly of highly similar gene isoforms, and may result in ambiguities and false joining during sequence assembly. More recently, the third generation sequencer represented by the PacBio platform offers sufficient throughput and long reads covering the full length of typical genes, thus providing a potential to reliably profile gene isoforms. However, the PacBio long reads are error-prone and cannot be effectively analyzed by traditional assembly programs.We present a clustering-based analysis pipeline integrated with PacBio sequencing data for profiling highly similar gene isoforms. This approach was first evaluated in comparison to de novo assembly of 454 reads using a benchmark admixture containing 10 known, cloned msg genes encoding the major surface glycoprotein of Pneumocystis jirovecii. All 10 msg isoforms were successfully reconstructed with the expected length (~1.5 kb) and correct sequence by the new approach, while 454 reads could not be correctly assembled using various assembly programs. When using an additional benchmark admixture containing 22 known P. jirovecii msg isoforms, this approach accurately reconstructed all but 4 these isoforms in their full-length (~3 kb); these 4 isoforms were present in low concentrations in the admixture. Finally, when applied to the original clinical sample from which the 22 known msg isoforms were cloned, this approach successfully identified not only all known isoforms accurately (~3 kb each) but also 48 novel isoforms.PacBio sequencing integrated with the clustering-based analysis pipeline achieves high-throughput and high-resolution discrimination of highly similar sequences, and can serve as a new approach for genome-wide characterization of gene isoforms and other highly repetitive sequences.

September 22, 2019

Genome-wide identification and analysis of the ALTERNATIVE OXIDASE gene family in diploid and hexaploid wheat.

A comprehensive understanding of wheat responses to environmental stress will contribute to the long-term goal of feeding the planet. ALERNATIVE OXIDASE (AOX) genes encode proteins involved in a bypass of the electron transport chain and are also known to be involved in stress tolerance in multiple species. Here, we report the identification and characterization of the AOX gene family in diploid and hexaploid wheat. Four genes each were found in the diploid ancestors Triticum urartu, and Aegilops tauschii, and three in Aegilops speltoides. In hexaploid wheat (Triticum aestivum), 20 genes were identified, some with multiple splice variants, corresponding to a total of 24 proteins for those with observed transcription and translation. These proteins were classified as AOX1a, AOX1c, AOX1e or AOX1d via phylogenetic analysis. Proteins lacking most or all signature AOX motifs were assigned to putative regulatory roles. Analysis of protein-targeting sequences suggests mixed localization to the mitochondria and other organelles. In comparison to the most studied AOX from Trypanosoma brucei, there were amino acid substitutions at critical functional domains indicating possible role divergence in wheat or grasses in general. In hexaploid wheat, AOX genes were expressed at specific developmental stages as well as in response to both biotic and abiotic stresses such as fungal pathogens, heat and drought. These AOX expression patterns suggest a highly regulated and diverse transcription and expression system. The insights gained provide a framework for the continued and expanded study of AOX genes in wheat for stress tolerance through breeding new varieties, as well as resistance to AOX-targeted herbicides, all of which can ultimately be used synergistically to improve crop yield.

September 22, 2019

Analysis of microbial community structure of pit mud for Chinese strong-flavor liquor fermentation using next generation DNA sequencing of full-length 16S rRNA

The pit is the necessary bioreactor for brewing process of Chinese strong-flavor liquor. Pit mud in pits contains a large number of microorganisms and is a complex ecosystem. The analysis of bacterial flora in pit mud is of great significance to understand liquor fermentation mechanisms. To overcome taxonomic limitations of short reads in 16S rRNA variable region sequencing, we used high-throughput DNA sequencing of near full-length 16S rRNA gene to analyze microbial compositions of different types of pit mud that produce different qualities of strong-flavor liquor. The results showed that the main species in pit mud were Pseudomonas extremaustralis 14-3, Pseudomonas veronii, Serratia marcescens WW4, and Clostridium leptum in Ruminiclostridium. The microbial diversity of pit mud with different quality was significantly different. From poor to good quality of pit mud (thus the quality of liquor), the relative abundances of Ruminiclostridium and Syntrophomonas in Firmicutes was increased, and the relative abundance of Olsenella in Actinobacteria also increased, but the relative abundances of Pseudomonas and Serratia in Proteobacteria were decreased. The surprising findings of this study include that the diversity of intermediate level quality of N pit mud was the lowest, and the diversity levels of high quality pit mud G and poor quality pit mud B were similar. Correlation analysis showed that there were high positive correlations (r > 0.8) among different microbial groups in the flora. Based on the analysis of the microbial structures of pit mud in different quality, the good quality pit mud has a higher microbial diversity, but how this higher diversity and differential microbial compositions contribute to better quality of liquor fermentation remains obscure.

September 22, 2019

PacBio sequencing of gene families – a case study with wheat gluten genes.

Amino acids in wheat (Triticum aestivum) seeds mainly accumulate in storage proteins called gliadins and glutenins. Gliadins contain a/ß-, ?- and ?-types whereas glutenins contain HMW- and LMW-types. Known gliadin and glutenin sequences were largely determined through cloning and sequencing by capillary electrophoresis. This time-consuming process prevents us to intensively study the variation of each orthologous gene copy among cultivars. The throughput and sequencing length of Pacific Bioscience RS (PacBio) single molecule sequencing platform make it feasible to construct contiguous and non-chimeric RNA sequences. We assembled 424 wheat storage protein transcripts from ten wheat cultivars by using just one single-molecule-real-time cell. The protein genes from wheat cultivar Chinese Spring are comparable to known sequences from NCBI. We demonstrated real-time sequencing of gene families with high-throughput and low-cost. This method can be applied to studies of gene amplification and copy number variation among species and cultivars. © 2013 Elsevier B.V. All rights reserved.

September 22, 2019

A high-quality annotated transcriptome of swine peripheral blood.

High throughput gene expression profiling assays of peripheral blood are widely used in biomedicine, as well as in animal genetics and physiology research. Accurate, comprehensive, and precise interpretation of such high throughput assays relies on well-characterized reference genomes and/or transcriptomes. However, neither the reference genome nor the peripheral blood transcriptome of the pig have been sufficiently assembled and annotated to support such profiling assays in this emerging biomedical model organism. We aimed to assemble published and novel RNA-seq data to provide a comprehensive, well-annotated blood transcriptome for pigs by integrating a de novo assembly with a genome-guided assembly.A de novo and a genome-guided transcriptome of porcine whole peripheral blood was assembled with ~162 million pairs of paired-end and ~183 million single-end, trimmed and normalized Illumina RNA-seq reads (~6 billion initial reads from 146 RNA-seq libraries) from five independent studies by using the Trinity and Cufflinks software, respectively. We then removed putative transcripts (PTs) of low confidence from both assemblies and merged the remaining PTs into an integrated transcriptome consisting of 132,928 PTs, with 126,225 (~95%) PTs from the de novo assembly and more than 91% of PTs spliced. In the integrated transcriptome, ~90% and 63% of PTs had significant sequence similarity to sequences in the NCBI NT and NR databases, respectively; 68,754 (~52%) PTs were annotated with 15,965 unique gene ontology (GO) terms; and 7618 PTs annotated with Enzyme Commission codes were assigned to 134 pathways curated by the Kyoto Encyclopedia of Genes and Genomes (KEGG). Full exon-intron junctions of 17,528 PTs were validated by PacBio IsoSeq full-length cDNA reads from 3 other porcine tissues, NCBI pig RefSeq mRNAs and transcripts from Ensembl Sscrofa10.2 annotation. Completeness of the 5′ termini of 37,569 PTs was validated by public cap analysis of gene expression (CAGE) data. By comparison to the Ensembl transcripts, we found that (1) the deduced precursors of 54,402 PTs shared at least one intron or exon with those of 18,437 Ensembl transcripts; (2) 12,262 PTs had both longer 5′ and 3′ termini than their maximally overlapping Ensembl transcripts; and (3) 41,838 spliced PTs were totally missing from the Sscrofa10.2 annotation. Similar results were obtained when the PTs were compared to the pig NCBI RefSeq mRNA collection.We built, validated and annotated a comprehensive porcine blood transcriptome with significant improvement over the annotation of Ensembl Sscrofa10.2 and the pig NCBI RefSeq mRNAs, and laid a foundation for blood-based high throughput transcriptomic assays in pigs and for advancing annotation of the pig genome.

September 22, 2019

Isoform sequencing provides a more comprehensive view of the Panax ginseng transcriptome.

Korean ginseng (Panax ginseng C.A. Meyer) has been widely used for medicinal purposes and contains potent plant secondary metabolites, including ginsenosides. To obtain transcriptomic data that offers a more comprehensive view of functional genomics in P. ginseng, we generated genome-wide transcriptome data from four different P. ginseng tissues using PacBio isoform sequencing (Iso-Seq) technology. A total of 135,317 assembled transcripts were generated with an average length of 3.2 kb and high assembly completeness. Of those unigenes, 67.5% were predicted to be complete full-length (FL) open reading frames (ORFs) and exhibited a high gene annotation rate. Furthermore, we successfully identified unique full-length genes involved in triterpenoid saponin synthesis and plant hormonal signaling pathways, including auxin and cytokinin. Studies on the functional genomics of P. ginseng seedlings have confirmed the rapid upregulation of negative feed-back loops by auxin and cytokinin signaling cues. The conserved evolutionary mechanisms in the auxin and cytokinin canonical signaling pathways of P. ginseng are more complex than those in Arabidopsis thaliana. Our analysis also revealed a more detailed view of transcriptome-wide alternative isoforms for 88 genes. Finally, transposable elements (TEs) were also identified, suggesting transcriptional activity of TEs in P. ginseng. In conclusion, our results suggest that long-read, full-length or partial-unigene data with high-quality assemblies are invaluable resources as transcriptomic references in P. ginseng and can be used for comparative analyses in closely related medicinal plants.

Auto Tag: Bioinformatics

Influenza virus infection causes global RNAPII termination defects.

A comprehensive quality evaluation system for complex herbal medicine using PacBio sequencing, PCR-denaturing gradient gel electrophoresis, and several chemical approaches.

Antagonism between Staphylococcus epidermidis and Propionibacterium acnes and its genomic basis.

Contrasting distribution patterns between aquatic and terrestrial Phytophthora species along a climatic gradient are linked to functional traits.

cDNA library enrichment of full length transcripts for SMRT long read sequencing.

Gene activity in primary T cells infected with HIV89.6: intron retention and induction of genomic repeats.

Chromosome-level reference genome and alternative splicing atlas of moso bamboo (Phyllostachys edulis).

Using PacBio long-read high-throughput microbial gene amplicon sequencing to evaluate infant formula safety.

Gaining comprehensive biological insight into the transcriptome by performing a broad-spectrum RNA-seq analysis.

Distinguishing highly similar gene isoforms with a clustering-based bioinformatics analysis of PacBio single-molecule long reads.

Genome-wide identification and analysis of the ALTERNATIVE OXIDASE gene family in diploid and hexaploid wheat.

Analysis of microbial community structure of pit mud for Chinese strong-flavor liquor fermentation using next generation DNA sequencing of full-length 16S rRNA

PacBio sequencing of gene families – a case study with wheat gluten genes.

A high-quality annotated transcriptome of swine peripheral blood.

Isoform sequencing provides a more comprehensive view of the Panax ginseng transcriptome.

Subscribe for blog updates:

Filter by topic

Talk with an expert

Antimicrobial resistance research

Subscribe for blog updates:

Filter by topic

Talk with an expert