de novo assembly Archives - Page 87 of 359

September 22, 2019

A high-quality annotated transcriptome of swine peripheral blood.

High throughput gene expression profiling assays of peripheral blood are widely used in biomedicine, as well as in animal genetics and physiology research. Accurate, comprehensive, and precise interpretation of such high throughput assays relies on well-characterized reference genomes and/or transcriptomes. However, neither the reference genome nor the peripheral blood transcriptome of the pig have been sufficiently assembled and annotated to support such profiling assays in this emerging biomedical model organism. We aimed to assemble published and novel RNA-seq data to provide a comprehensive, well-annotated blood transcriptome for pigs by integrating a de novo assembly with a genome-guided assembly.A de novo and a genome-guided transcriptome of porcine whole peripheral blood was assembled with ~162 million pairs of paired-end and ~183 million single-end, trimmed and normalized Illumina RNA-seq reads (~6 billion initial reads from 146 RNA-seq libraries) from five independent studies by using the Trinity and Cufflinks software, respectively. We then removed putative transcripts (PTs) of low confidence from both assemblies and merged the remaining PTs into an integrated transcriptome consisting of 132,928 PTs, with 126,225 (~95%) PTs from the de novo assembly and more than 91% of PTs spliced. In the integrated transcriptome, ~90% and 63% of PTs had significant sequence similarity to sequences in the NCBI NT and NR databases, respectively; 68,754 (~52%) PTs were annotated with 15,965 unique gene ontology (GO) terms; and 7618 PTs annotated with Enzyme Commission codes were assigned to 134 pathways curated by the Kyoto Encyclopedia of Genes and Genomes (KEGG). Full exon-intron junctions of 17,528 PTs were validated by PacBio IsoSeq full-length cDNA reads from 3 other porcine tissues, NCBI pig RefSeq mRNAs and transcripts from Ensembl Sscrofa10.2 annotation. Completeness of the 5′ termini of 37,569 PTs was validated by public cap analysis of gene expression (CAGE) data. By comparison to the Ensembl transcripts, we found that (1) the deduced precursors of 54,402 PTs shared at least one intron or exon with those of 18,437 Ensembl transcripts; (2) 12,262 PTs had both longer 5′ and 3′ termini than their maximally overlapping Ensembl transcripts; and (3) 41,838 spliced PTs were totally missing from the Sscrofa10.2 annotation. Similar results were obtained when the PTs were compared to the pig NCBI RefSeq mRNA collection.We built, validated and annotated a comprehensive porcine blood transcriptome with significant improvement over the annotation of Ensembl Sscrofa10.2 and the pig NCBI RefSeq mRNAs, and laid a foundation for blood-based high throughput transcriptomic assays in pigs and for advancing annotation of the pig genome.

September 22, 2019

A workflow for studying specialized metabolism in nonmodel eukaryotic organisms

Eukaryotes contain a diverse tapestry of specialized metabolites, many of which are of significant pharmaceutical and industrial importance to humans. Nevertheless, exploration of specialized metabolic pathways underlying specific chemical traits in nonmodel eukaryotic organisms has been technically challenging and historically lagged behind that of the bacterial systems. Recent advances in genomics, metabolomics, phylogenomics, and synthetic biology now enable a new workflow for interrogating unknown specialized metabolic systems in nonmodel eukaryotic hosts with greater efficiency and mechanistic depth. This chapter delineates such workflow by providing a collection of state-of-the-art approaches and tools, ranging from multiomics-guided candidate gene identification to in vitro and in vivo functional and structural characterization of specialized metabolic enzymes. As already demonstrated by several recent studies, this new workflow opens up a gateway into the largely untapped world of natural product biochemistry in eukaryotes. © 2016 Elsevier Inc. All rights reserved.

September 22, 2019

P_RNA_scaffolder: a fast and accurate genome scaffolder using paired-end RNA-sequencing reads.

Obtaining complete gene structures is one major goal of genome assembly. Some gene regions are fragmented in low quality and high-quality assemblies. Therefore, new approaches are needed to recover gene regions. Genomes are widely transcribed, generating messenger and non-coding RNAs. These widespread transcripts can be used to scaffold genomes and complete transcribed regions.We present P_RNA_scaffolder, a fast and accurate tool using paired-end RNA-sequencing reads to scaffold genomes. This tool aims to improve the completeness of both protein-coding and non-coding genes. After this tool was applied to scaffolding human contigs, the structures of both protein-coding genes and circular RNAs were almost completely recovered and equivalent to those in a complete genome, especially for long proteins and long circular RNAs. Tested in various species, P_RNA_scaffolder exhibited higher speed and efficiency than the existing state-of-the-art scaffolders. This tool also improved the contiguity of genome assemblies generated by current mate-pair scaffolding and third-generation single-molecule sequencing assembly.The P_RNA_scaffolder can improve the contiguity of genome assembly and benefit gene prediction. This tool is available at http://www.fishbrowser.org/software/P_RNA_scaffolder .

September 22, 2019

Biodegradation of nonylphenol during aerobic composting of sewage sludge under two intermittent aeration treatments in a full-scale plant.

The urbanization and industrialization of cities around the coastal region of the Bohai Sea have produced large amounts of sewage sludge from sewage treatment plants. Research on the biodegradation of nonylphenol (NP) and the influencing factors of such biodegradation during sewage sludge composting is important to control pollution caused by land application of sewage sludge. The present study investigated the effect of aeration on NP biodegradation and the microbe community during aerobic composting under two intermittent aeration treatments in a full-scale plant of sewage sludge, sawdust, and returned compost at a ratio of 6:3:1. The results showed that 65% of NP was biodegraded and that Bacillus was the dominant bacterial species in the mesophilic phase. The amount of NP biodegraded in the mesophilic phase was 68.3%, which accounted for 64.6% of the total amount of biodegraded NP. The amount of NP biodegraded under high-volume aeration was 19.6% higher than that under low-volume aeration. Bacillus was dominant for 60.9% of the composting period under high-volume aeration, compared to 22.7% dominance under low-volume aeration. In the thermophilic phase, high-volume aeration promoted the biodegradation of NP and Bacillus remained the dominant bacterial species. In the cooling and stable phases, the contents of NP underwent insignificant change while different dominant bacteria were observed in the two treatments. NP was mostly biodegraded by Bacillus, and the rate of biodegradation was significantly correlated with the abundance of Bacillus (r?=?0.63, p?

September 22, 2019

Genetic determinants of in vivo fitness and diet responsiveness in multiple human gut Bacteroides.

Libraries of tens of thousands of transposon mutants generated from each of four human gut Bacteroides strains, two representing the same species, were introduced simultaneously into gnotobiotic mice together with 11 other wild-type strains to generate a 15-member artificial human gut microbiota. Mice received one of two distinct diets monotonously, or both in different ordered sequences. Quantifying the abundance of mutants in different diet contexts allowed gene-level characterization of fitness determinants, niche, stability, and resilience and yielded a prebiotic (arabinoxylan) that allowed targeted manipulation of the community. The approach described is generalizable and should be useful for defining mechanisms critical for sustaining and/or approaches for deliberately reconfiguring the highly adaptive and durable relationship between the human gut microbiota and host in ways that promote wellness. Copyright © 2015, American Association for the Advancement of Science.

September 22, 2019

Lentinula edodes genome survey and postharvest transcriptome analysis.

Lentinula edodes is a popular, cultivated edible and medicinal mushroom. Lentinula edodes is susceptible to postharvest problems, such as gill browning, fruiting body softening, and lentinan degradation. We constructed a de novo assembly draft genome sequence and performed gene prediction for Lentinula edodesDe novo assembly was carried out using short reads from paired-end and mate-paired libraries and by using long reads by PacBio, resulting in a contig number of 1,951 and an N50 of 1 Mb. Furthermore, we predicted genes by Augustus using transcriptome sequencing (RNA-seq) data from the whole life cycle of Lentinula edodes, resulting in 12,959 predicted genes. This analysis revealed that Lentinula edodes lacks lignin peroxidase. To reveal genes involved in the loss of quality of Lentinula edodes postharvest fruiting bodies, transcriptome analysis was carried out using serial analysis of gene expression (SuperSAGE). This analysis revealed that many cell wall-related enzymes are upregulated after harvest, such as ß-1,3-1,6-glucan-degrading enzymes in glycoside hydrolase (GH) families GH5, GH16, GH30, GH55, and GH128, and thaumatin-like proteins. In addition, we found that several chitin-related genes are upregulated, such as putative chitinases in GH family 18, exochitinases in GH20, and a putative chitosanase in GH family 75. The results suggest that cell wall-degrading enzymes synergistically cooperate for rapid fruiting body autolysis. Many putative transcription factor genes were upregulated postharvest, such as genes containing high-mobility-group (HMG) domains and zinc finger domains. Several cell death-related proteins were also upregulated postharvest.IMPORTANCE Our data collectively suggest that there is a rapid fruiting body autolysis system in Lentinula edodes The genes for the loss of postharvest quality newly found in this research will be targets for the future breeding of strains that keep fresh longer than present strains. De novoLentinula edodes genome assembly data will be used for the construction of a complete Lentinula edodes chromosome map for future breeding. Copyright © 2017 American Society for Microbiology.

September 22, 2019

The habu genome reveals accelerated evolution of venom protein genes.

Evolution of novel traits is a challenging subject in biological research. Several snake lineages developed elaborate venom systems to deliver complex protein mixtures for prey capture. To understand mechanisms involved in snake venom evolution, we decoded here the ~1.4-Gb genome of a habu, Protobothrops flavoviridis. We identified 60 snake venom protein genes (SV) and 224 non-venom paralogs (NV), belonging to 18 gene families. Molecular phylogeny reveals early divergence of SV and NV genes, suggesting that one of the four copies generated through two rounds of whole-genome duplication was modified for use as a toxin. Among them, both SV and NV genes in four major components were extensively duplicated after their diversification, but accelerated evolution is evident exclusively in the SV genes. Both venom-related SV and NV genes are significantly enriched in microchromosomes. The present study thus provides a genetic background for evolution of snake venom composition.

September 22, 2019

PLEK: a tool for predicting long non-coding RNAs and messenger RNAs based on an improved k-mer scheme.

High-throughput transcriptome sequencing (RNA-seq) technology promises to discover novel protein-coding and non-coding transcripts, particularly the identification of long non-coding RNAs (lncRNAs) from de novo sequencing data. This requires tools that are not restricted by prior gene annotations, genomic sequences and high-quality sequencing.We present an alignment-free tool called PLEK (predictor of long non-coding RNAs and messenger RNAs based on an improved k-mer scheme), which uses a computational pipeline based on an improved k-mer scheme and a support vector machine (SVM) algorithm to distinguish lncRNAs from messenger RNAs (mRNAs), in the absence of genomic sequences or annotations. The performance of PLEK was evaluated on well-annotated mRNA and lncRNA transcripts. 10-fold cross-validation tests on human RefSeq mRNAs and GENCODE lncRNAs indicated that our tool could achieve accuracy of up to 95.6%. We demonstrated the utility of PLEK on transcripts from other vertebrates using the model built from human datasets. PLEK attained >90% accuracy on most of these datasets. PLEK also performed well using a simulated dataset and two real de novo assembled transcriptome datasets (sequenced by PacBio and 454 platforms) with relatively high indel sequencing errors. In addition, PLEK is approximately eightfold faster than a newly developed alignment-free tool, named Coding-Non-Coding Index (CNCI), and 244 times faster than the most popular alignment-based tool, Coding Potential Calculator (CPC), in a single-threading running manner.PLEK is an efficient alignment-free computational tool to distinguish lncRNAs from mRNAs in RNA-seq transcriptomes of species lacking reference genomes. PLEK is especially suitable for PacBio or 454 sequencing data and large-scale transcriptome data. Its open-source software can be freely downloaded from https://sourceforge.net/projects/plek/files/.

September 22, 2019

Long-read, Single Molecule, Real-Time (SMRT) DNA Sequencing for metagenomic applications

In this chapter, we describe applications of single molecule, real-time (SMRT) DNA sequencing toward metagenomic research. The long sequence reads, combined with a lack of bias with respect to DNA sequence context or GC content, facilitate a more comprehensive analysis of the genomic constitution of microbial communities. Full-length 16S RNA gene sequencing at high (>99%) accuracy allows for species-level characterization of community members concomitant with the determination of community structure. The application of SMRT sequencing to whole-community shotgun microbial metagenomics has also been discussed.

September 22, 2019

Single Molecule Sequencing: new outlooks for solving genome assembly and transcripts identification challenges

In this review, we introduce a novel sequencing technology, named Single Molecule Real Time sequencing. Also called Single Molecule Sequencing, as it do not requires any amplification, this new technology is able to pro- duce much longer reads than previous NGS technologies such as Illumina. This read size improvements, which can reach 150 fold, will solve many challenges caused by the actual NGS technologies. Short NGS reads, reach- ing a maximum size of 300 bp, make it hard to reconstitute a whole genome and are always leading to fragmented genome assembly. It is also difficult to correctly infer transcript quantification and identification when there is a high isoforms diversity. Despite their higher error rate, long reads have shown very promising result concerning these actual issues. We show that longer reads can produce less fragmented assembly, with a better quality, but also sequence from start to end mRNA, making it much more easier to infer correct transcript quantification, and even allow new intron structure and so new isoforms discovery.

September 22, 2019

ALE: a generic assembly likelihood evaluation framework for assessing the accuracy of genome and metagenome assemblies.

Researchers need general purpose methods for objectively evaluating the accuracy of single and metagenome assemblies and for automatically detecting any errors they may contain. Current methods do not fully meet this need because they require a reference, only consider one of the many aspects of assembly quality or lack statistical justification, and none are designed to evaluate metagenome assemblies.In this article, we present an Assembly Likelihood Evaluation (ALE) framework that overcomes these limitations, systematically evaluating the accuracy of an assembly in a reference-independent manner using rigorous statistical methods. This framework is comprehensive, and integrates read quality, mate pair orientation and insert length (for paired-end reads), sequencing coverage, read alignment and k-mer frequency. ALE pinpoints synthetic errors in both single and metagenomic assemblies, including single-base errors, insertions/deletions, genome rearrangements and chimeric assemblies presented in metagenomes. At the genome level with real-world data, ALE identifies three large misassemblies from the Spirochaeta smaragdinae finished genome, which were all independently validated by Pacific Biosciences sequencing. At the single-base level with Illumina data, ALE recovers 215 of 222 (97%) single nucleotide variants in a training set from a GC-rich Rhodobacter sphaeroides genome. Using real Pacific Biosciences data, ALE identifies 12 of 12 synthetic errors in a Lambda Phage genome, surpassing even Pacific Biosciences’ own variant caller, EviCons. In summary, the ALE framework provides a comprehensive, reference-independent and statistically rigorous measure of single genome and metagenome assembly accuracy, which can be used to identify misassemblies or to optimize the assembly process.ALE is released as open source software under the UoI/NCSA license at http://www.alescore.org. It is implemented in C and Python.

September 22, 2019

Metagenomic SMRT sequencing-based exploration of novel lignocellulose-degrading capability in wood detritus from Torreya nucifera in Bija forest on Jeju Island.

Lignocellulose, mostly composed of cellulose, hemicellulose and lignin generated through secondary growth of woody plant, is considered as promising resources for bio-fuel. In order to use lignocellulose as a biofuel, the biodegradation besides high-cost chemical treatments were applied, but its knowledge on decomposition of lignocellulose occurring in a natural environment were insufficient. We analyzed 16S rRNA gene and metagenome to understand how the lignocellulose are decomposed naturally in decayed Torreya nucifera (L) of Bija forest (Bijarim) in Gotjawal, an ecologically distinct environment. A total of 464,360 reads were obtained from 16S rRNA gene sequencing, representing diverse phyla; Proteobacteria (51%), Bacteroidetes (11%) and Actinobacteria (10%). The metagenome analysis using Single Molecules Real-Time Sequencing revealed that the assembled contigs determined by originated from Proteobacteria (58%) and Actinobacteria (10.3%). Carbohydrate Active enZYmes (CAZy) and Protein families (Pfam) based analysis showed that Proteobacteria was involved in degrading whole lignocellulose and Actinobacteria played a role only in a part of hemicellulose degradation. Combining these results, it suggested that Proteobacteria and Actinobacteria had selective biodegradation potential for different lignocellulose substrate. Thus, it is considered that understanding of the systemic microbial degradation pathways may be a useful strategy for recycle of lignocellulosic biomass and the microbial enzymes in Bija forest can be useful natural resources in industrial processes.

September 22, 2019

The industrial melanism mutation in British peppered moths is a transposable element.

Discovering the mutational events that fuel adaptation to environmental change remains an important challenge for evolutionary biology. The classroom example of a visible evolutionary response is industrial melanism in the peppered moth (Biston betularia): the replacement, during the Industrial Revolution, of the common pale typica form by a previously unknown black (carbonaria) form, driven by the interaction between bird predation and coal pollution. The carbonaria locus has been coarsely localized to a 200-kilobase region, but the specific identity and nature of the sequence difference controlling the carbonaria-typica polymorphism, and the gene it influences, are unknown. Here we show that the mutation event giving rise to industrial melanism in Britain was the insertion of a large, tandemly repeated, transposable element into the first intron of the gene cortex. Statistical inference based on the distribution of recombined carbonaria haplotypes indicates that this transposition event occurred around 1819, consistent with the historical record. We have begun to dissect the mode of action of the carbonaria transposable element by showing that it increases the abundance of a cortex transcript, the protein product of which plays an important role in cell-cycle regulation, during early wing disc development. Our findings fill a substantial knowledge gap in the iconic example of microevolutionary change, adding a further layer of insight into the mechanism of adaptation in response to natural selection. The discovery that the mutation itself is a transposable element will stimulate further debate about the importance of ‘jumping genes’ as a source of major phenotypic novelty.

September 22, 2019

Complex rearrangements and oncogene amplifications revealed by long-read DNA and RNA sequencing of a breast cancer cell line.

The SK-BR-3 cell line is one of the most important models for HER2+ breast cancers, which affect one in five breast cancer patients. SK-BR-3 is known to be highly rearranged, although much of the variation is in complex and repetitive regions that may be underreported. Addressing this, we sequenced SK-BR-3 using long-read single molecule sequencing from Pacific Biosciences and develop one of the most detailed maps of structural variations (SVs) in a cancer genome available, with nearly 20,000 variants present, most of which were missed by short-read sequencing. Surrounding the important ERBB2 oncogene (also known as HER2), we discover a complex sequence of nested duplications and translocations, suggesting a punctuated progression. Full-length transcriptome sequencing further revealed several novel gene fusions within the nested genomic variants. Combining long-read genome and transcriptome sequencing enables an in-depth analysis of how SVs disrupt the genome and sheds new light on the complex mechanisms involved in cancer genome evolution.© 2018 Nattestad et al.; Published by Cold Spring Harbor Laboratory Press.

September 22, 2019

Next-generation sequencing for pathogen detection and identification

Over the past decade, the field of genomics has seen such drastic improvements in sequencing chemistries that high-throughput sequencing, or next-generation sequencing (NGS), is being applied to generate data across many disciplines. NGS instruments are becoming less expensive, faster, and smaller, and therefore are being adopted in an increasing number of laboratories, including clinical laboratories. Thus far, clinical use of NGS has been mostly focused on the human genome, for purposes such as characterizing the molecular basis of cancer or for diagnosing and understanding the basis of rare genetic disorders. There are, however, an increasing number of examples whereby NGS is employed to discover novel pathogens, and these cases provide precedent for the use of NGS in microbial diagnostics. NGS has many advantages over traditional microbial diagnostic methods, such as unbiased rather than pathogen-specific protocols, ability to detect fastidious or non-culturable organisms, and ability to detect co-infections. One of the most impressive advantages of NGS is that it requires little or no prior knowledge of the pathogen, unlike many other diagnostic assays; therefore for pathogen discovery, NGS is very valuable. However, despite these advantages, there are challenges involved in implementing NGS for routine clinical microbiological diagnosis. We discuss these advantages and challenges in the context of recently described research studies.

Asset Tag: de novo assembly

A high-quality annotated transcriptome of swine peripheral blood.

A workflow for studying specialized metabolism in nonmodel eukaryotic organisms

P_RNA_scaffolder: a fast and accurate genome scaffolder using paired-end RNA-sequencing reads.

Biodegradation of nonylphenol during aerobic composting of sewage sludge under two intermittent aeration treatments in a full-scale plant.

Genetic determinants of in vivo fitness and diet responsiveness in multiple human gut Bacteroides.

Lentinula edodes genome survey and postharvest transcriptome analysis.

The habu genome reveals accelerated evolution of venom protein genes.

PLEK: a tool for predicting long non-coding RNAs and messenger RNAs based on an improved k-mer scheme.

Long-read, Single Molecule, Real-Time (SMRT) DNA Sequencing for metagenomic applications

Single Molecule Sequencing: new outlooks for solving genome assembly and transcripts identification challenges

ALE: a generic assembly likelihood evaluation framework for assessing the accuracy of genome and metagenome assemblies.

Metagenomic SMRT sequencing-based exploration of novel lignocellulose-degrading capability in wood detritus from Torreya nucifera in Bija forest on Jeju Island.

The industrial melanism mutation in British peppered moths is a transposable element.

Complex rearrangements and oncogene amplifications revealed by long-read DNA and RNA sequencing of a breast cancer cell line.

Next-generation sequencing for pathogen detection and identification

Subscribe for blog updates:

Filter by topic

Talk with an expert

Antimicrobial resistance research

Subscribe for blog updates:

Filter by topic

Talk with an expert