Genome assembly Archives - Page 114 of 196

July 7, 2019

High resolution assembly and characterization of genomes of Canadian isolates of Salmonella Enteritidis.

There is a need to characterize genomes of the foodborne pathogen, Salmonella enterica serovar Enteritidis (SE) and identify genetic information that could be ultimately deployed for differentiating strains of the organism, a need that is yet to be addressed mainly because of the high degree of clonality of the organism. In an effort to achieve the first characterization of the genomes of SE of Canadian origin, we carried out massively parallel sequencing of the nucleotide sequence of 11 SE isolates obtained from poultry production environments (n?=?9), a clam and a chicken, assembled finished genomes and investigated diversity of the SE genome.The median genome size was 4,678,683 bp. A total of 4,833 chromosomal genes defined the pan genome of our field SE isolates consisting of 4,600 genes present in all the genomes, i.e., core genome, and 233 genes absent in at least one genome (accessory genome). Genome diversity was demonstrable by the presence of 1,360 loci showing single nucleotide polymorphism (SNP) in the core genome which was used to portray the genetic distances by means of a phylogenetic tree for the SE isolates. The accessory genome consisted mostly of previously identified SE prophage sequences as well as two, apparently full-sized, novel prophages namely a 28 kb sequence provisionally designated as SE-OLF-10058 (3) prophage and a 43 kb sequence provisionally designated as SE-OLF-10012 prophage.The number of SNPs identified in the relatively large core genome of SE is a reflection of substantial diversity that could be exploited for strain differentiation as shown by the development of an informative phylogenetic tree. Prophage sequences can also be exploited for SE strain differentiation and lineage tracking. This work has laid the ground work for further studies to develop a readily adoptable laboratory test for the subtyping of SE.

July 7, 2019

Fully assembled genome sequence for Salmonella enterica subsp. enterica serovar Javiana CFSAN001992.

We report a closed genome of Salmonella enterica subsp. enterica serovar Javiana (S. Javiana). This serotype is a common food-borne pathogen and is often associated with fresh-cut produce. Complete (finished) genome assemblies will support pilot studies testing the utility of next-generation sequencing (NGS) technologies in public health laboratories.

July 7, 2019

The first 50 plant genomes

Fifty-five plant genomes have been published to date representing 49 different species (Table 1 includes PubMed IDs for complete reference). What have we learned from the first wave of plant genomes? It has been said that plant genome papers (and genome papers in general) are dry and lack “biology” and that the days of high impact plant genome papers are drawing to a close unless they explore significant biology. However, with each new genome, earlier observations are refined and plant genome papers continue to reveal novel aspects of genome biology. For example, the tomato and banana genome papers refined current thinking on the whole genome duplications (WGD) that shaped dicot and monocot genome evolution (D’Hont et al., 2012; Tomato Genome Consortium, 2012). These observations were enabled not only by high quality genome assemblies but also by a greater number of genomes available for com- parisons. In addition, the initial round of plant genomes enabled the first generation of functional genomics that helped to define the roles of hundreds of genes, provided unprecedented access to sequence-based markers for breeding, and provided glimpses into plant evolutionary history. More genomes, representing the diverse array of species in Viridiplantae are still required to gain a full understanding of plant genome structure, evolution, and complexity.

July 7, 2019

De novo assembly of the Streptomyces sp. strain Mg1 genome using PacBio single-molecule sequencing.

We report a draft genome assembly of Streptomyces sp. strain Mg1, a competitive soil isolate with multiple secondary metabolite gene clusters.

July 7, 2019

Whole-exome targeted sequencing of the uncharacterized pine genome.

The large genome size of many species hinders the development and application of genomic tools to study them. For instance, loblolly pine (Pinus taeda L.), an ecologically and economically important conifer, has a large and yet uncharacterized genome of 21.7 Gbp. To characterize the pine genome, we performed exome capture and sequencing of 14 729 genes derived from an assembly of expressed sequence tags. Efficiency of sequence capture was evaluated and shown to be similar across samples with increasing levels of complexity, including haploid cDNA, haploid genomic DNA and diploid genomic DNA. However, this efficiency was severely reduced for probes that overlapped multiple exons, presumably because intron sequences hindered probe:exon hybridizations. Such regions could not be entirely avoided during probe design, because of the lack of a reference sequence. To improve the throughput and reduce the cost of sequence capture, a method to multiplex the analysis of up to eight samples was developed. Sequence data showed that multiplexed capture was reproducible among 24 haploid samples, and can be applied for high-throughput analysis of targeted genes in large populations. Captured sequences were de novo assembled, resulting in 11 396 expanded and annotated gene models, significantly improving the knowledge about the pine gene space. Interspecific capture was also evaluated with over 98% of all probes designed from P. taeda that were efficient in sequence capture, were also suitable for analysis of the related species Pinus elliottii Engelm.© 2013 The Authors The Plant Journal © 2013 John Wiley & Sons Ltd.

July 7, 2019

The genome sequence of Streptomyces lividans 66 reveals a novel tRNA-dependent peptide biosynthetic system within a metal-related genomic island.

The complete genome sequence of the original isolate of the model actinomycete Streptomyces lividans 66, also referred to as 1326, was deciphered after a combination of next-generation sequencing platforms and a hybrid assembly pipeline. Comparative analysis of the genomes of S. lividans 66 and closely related strains, including S. coelicolor M145 and S. lividans TK24, was used to identify strain-specific genes. The genetic diversity identified included a large genomic island with a mosaic structure, present in S. lividans 66 but not in the strain TK24. Sequence analyses showed that this genomic island has an anomalous (G + C) content, suggesting recent acquisition and that it is rich in metal-related genes. Sequences previously linked to a mobile conjugative element, termed plasmid SLP3 and defined here as a 94 kb region, could also be identified within this locus. Transcriptional analysis of the response of S. lividans 66 to copper was used to corroborate a role of this large genomic island, including two SLP3-borne “cryptic” peptide biosynthetic gene clusters, in metal homeostasis. Notably, one of these predicted biosynthetic systems includes an unprecedented nonribosomal peptide synthetase–tRNA-dependent transferase biosynthetic hybrid organization. This observation implies the recruitment of members of the leucyl/phenylalanyl-tRNA-protein transferase family to catalyze peptide bond formation within the biosynthesis of natural products. Thus, the genome sequence of S. lividans 66 not only explains long-standing genetic and phenotypic differences but also opens the door for further in-depth comparative genomic analyses of model Streptomyces strains, as well as for the discovery of novel natural products following genome-mining approaches.

July 7, 2019

In transition: primate genomics at a time of rapid change.

The field of nonhuman primate genomics is undergoing rapid change and making impressive progress. Exploiting new technologies for DNA sequencing, researchers have generated new whole-genome sequence assemblies for multiple primate species over the past 6 years. In addition, investigations of within-species genetic variation, gene expression and RNA sequences, conservation of non-protein-coding regions of the genome, and other aspects of comparative genomics are moving at an accelerating speed. This progress is opening a wide array of new research opportunities in the analysis of comparative primate genome content and evolution. It also creates new possibilities for the use of nonhuman primates as model organisms in biomedical research. This transition, based on both new technology and the new information being generated in regard to human genetics, provides an important justification for reevaluating the research goals, strategies, and study designs used in primate genetics and genomics.

July 7, 2019

Haplotype assembly in polyploid genomes and identical by descent shared tracts.

Genome-wide haplotype reconstruction from sequence data, or haplotype assembly, is at the center of major challenges in molecular biology and life sciences. For complex eukaryotic organisms like humans, the genome is vast and the population samples are growing so rapidly that algorithms processing high-throughput sequencing data must scale favorably in terms of both accuracy and computational efficiency. Furthermore, current models and methodologies for haplotype assembly (i) do not consider individuals sharing haplotypes jointly, which reduces the size and accuracy of assembled haplotypes, and (ii) are unable to model genomes having more than two sets of homologous chromosomes (polyploidy). Polyploid organisms are increasingly becoming the target of many research groups interested in the genomics of disease, phylogenetics, botany and evolution but there is an absence of theory and methods for polyploid haplotype reconstruction.In this work, we present a number of results, extensions and generalizations of compass graphs and our HapCompass framework. We prove the theoretical complexity of two haplotype assembly optimizations, thereby motivating the use of heuristics. Furthermore, we present graph theory-based algorithms for the problem of haplotype assembly using our previously developed HapCompass framework for (i) novel implementations of haplotype assembly optimizations (minimum error correction), (ii) assembly of a pair of individuals sharing a haplotype tract identical by descent and (iii) assembly of polyploid genomes. We evaluate our methods on 1000 Genomes Project, Pacific Biosciences and simulated sequence data.HapCompass is available for download at http://www.brown.edu/Research/Istrail_Lab/.Supplementary data are available at Bioinformatics online.

July 7, 2019

Combining de novo and reference-guided assembly with scaffold_builder.

Genome sequencing has become routine, however genome assembly still remains a challenge despite the computational advances in the last decade. In particular, the abundance of repeat elements in genomes makes it difficult to assemble them into a single complete sequence. Identical repeats shorter than the average read length can generally be assembled without issue. However, longer repeats such as ribosomal RNA operons cannot be accurately assembled using existing tools. The application Scaffold_builder was designed to generate scaffolds – super contigs of sequences joined by N-bases – based on the similarity to a closely related reference sequence. This is independent of mate-pair information and can be used complementarily for genome assembly, e.g. when mate-pairs are not available or have already been exploited. Scaffold_builder was evaluated using simulated pyrosequencing reads of the bacterial genomes Escherichia coli 042, Lactobacillus salivarius UCC118 and Salmonella enterica subsp. enterica serovar Typhi str. P-stx-12. Moreover, we sequenced two genomes from Salmonella enterica serovar Typhimurium LT2 G455 and Salmonella enterica serovar Typhimurium SDT1291 and show that Scaffold_builder decreases the number of contig sequences by 53% while more than doubling their average length. Scaffold_builder is written in Python and is available at http://edwards.sdsu.edu/scaffold_builder. A web-based implementation is additionally provided to allow users to submit a reference genome and a set of contigs to be scaffolded.

July 7, 2019

Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species.

The process of generating raw genome sequence data continues to become cheaper, faster, and more accurate. However, assembly of such data into high-quality, finished genome sequences remains challenging. Many genome assembly tools are available, but they differ greatly in terms of their performance (speed, scalability, hardware requirements, acceptance of newer read technologies) and in their final output (composition of assembled sequence). More importantly, it remains largely unclear how to best assess the quality of assembled genome sequences. The Assemblathon competitions are intended to assess current state-of-the-art methods in genome assembly.In Assemblathon 2, we provided a variety of sequence data to be assembled for three vertebrate species (a bird, a fish, and snake). This resulted in a total of 43 submitted assemblies from 21 participating teams. We evaluated these assemblies using a combination of optical map data, Fosmid sequences, and several statistical methods. From over 100 different metrics, we chose ten key measures by which to assess the overall quality of the assemblies.Many current genome assemblers produced useful assemblies, containing a significant representation of their genes and overall genome structure. However, the high degree of variability between the entries suggests that there is still much room for improvement in the field of genome assembly and that approaches which work well in assembling the genome of one species may not necessarily work well for another.

July 7, 2019

Complete genome sequence of a multidrug-resistant Salmonella enterica serovar Typhimurium var. 5- strain isolated from chicken breast.

Salmonella enterica subsp. enterica serovar Typhimurium is a leading cause of salmonellosis. Here, we report a closed genome sequence, including sequences of 3 plasmids, of Salmonella serovar Typhimurium var. 5- CFSAN001921 (National Antimicrobial Resistance Monitoring System [NARMS] strain ID N30688), which was isolated from chicken breast meat and shows resistance to 10 different antimicrobials. Whole-genome and plasmid sequence analyses of this isolate will help enhance our understanding of this pathogenic multidrug-resistant serovar.

July 7, 2019

Genome of an arbuscular mycorrhizal fungus provides insight into the oldest plant symbiosis.

The mutualistic symbiosis involving Glomeromycota, a distinctive phylum of early diverging Fungi, is widely hypothesized to have promoted the evolution of land plants during the middle Paleozoic. These arbuscular mycorrhizal fungi (AMF) perform vital functions in the phosphorus cycle that are fundamental to sustainable crop plant productivity. The unusual biological features of AMF have long fascinated evolutionary biologists. The coenocytic hyphae host a community of hundreds of nuclei and reproduce clonally through large multinucleated spores. It has been suggested that the AMF maintain a stable assemblage of several different genomes during the life cycle, but this genomic organization has been questioned. Here we introduce the 153-Mb haploid genome of Rhizophagus irregularis and its repertoire of 28,232 genes. The observed low level of genome polymorphism (0.43 SNP per kb) is not consistent with the occurrence of multiple, highly diverged genomes. The expansion of mating-related genes suggests the existence of cryptic sex-related processes. A comparison of gene categories confirms that R. irregularis is close to the Mucoromycotina. The AMF obligate biotrophy is not explained by genome erosion or any related loss of metabolic complexity in central metabolism, but is marked by a lack of genes encoding plant cell wall-degrading enzymes and of genes involved in toxin and thiamine synthesis. A battery of mycorrhiza-induced secreted proteins is expressed in symbiotic tissues. The present comprehensive repertoire of R. irregularis genes provides a basis for future research on symbiosis-related mechanisms in Glomeromycota.

July 7, 2019

PBSIM: PacBio reads simulator–toward accurate genome assembly.

PacBio sequencers produce two types of characteristic reads (continuous long reads: long and high error rate and circular consensus sequencing: short and low error rate), both of which could be useful for de novo assembly of genomes. Currently, there is no available simulator that targets the specific generation of PacBio libraries.Our analysis of 13 PacBio datasets showed characteristic features of PacBio reads (e.g. the read length of PacBio reads follows a log-normal distribution). We have developed a read simulator, PBSIM, that captures these features using either a model-based or sampling-based method. Using PBSIM, we conducted several hybrid error correction and assembly tests for PacBio reads, suggesting that a continuous long reads coverage depth of at least 15 in combination with a circular consensus sequencing coverage depth of at least 30 achieved extensive assembly results.PBSIM is freely available from the web under the GNU GPL v2 license (http://code.google.com/p/pbsim/).

July 7, 2019

Complete genome sequence of Leifsonia xyli subsp. cynodontis strain DSM46306, a gram-positive bacterial pathogen of grasses.

We announce the complete genome sequence of Leifsonia xyli subsp. cynodontis, a vascular pathogen of Bermuda grass. The species also comprises Leifsonia xyli subsp. xyli, a sugarcane pathogen. Since these two subspecies have genome sequences available, a comparative analysis will contribute to our understanding of the differences in their biology and host specificity.

July 7, 2019

Technology feature: The genome jigsaw.

Advances in high-throughput sequencing are accelerating genomics research, but crucial gaps in data remain.

Auto Tag: Genome assembly

High resolution assembly and characterization of genomes of Canadian isolates of Salmonella Enteritidis.

Fully assembled genome sequence for Salmonella enterica subsp. enterica serovar Javiana CFSAN001992.

The first 50 plant genomes

De novo assembly of the Streptomyces sp. strain Mg1 genome using PacBio single-molecule sequencing.

Whole-exome targeted sequencing of the uncharacterized pine genome.

The genome sequence of Streptomyces lividans 66 reveals a novel tRNA-dependent peptide biosynthetic system within a metal-related genomic island.

In transition: primate genomics at a time of rapid change.

Haplotype assembly in polyploid genomes and identical by descent shared tracts.

Combining de novo and reference-guided assembly with scaffold_builder.

Assemblathon 2: evaluating de novo methods of genome assembly in three vertebrate species.

Complete genome sequence of a multidrug-resistant Salmonella enterica serovar Typhimurium var. 5- strain isolated from chicken breast.

Genome of an arbuscular mycorrhizal fungus provides insight into the oldest plant symbiosis.

PBSIM: PacBio reads simulator–toward accurate genome assembly.

Complete genome sequence of Leifsonia xyli subsp. cynodontis strain DSM46306, a gram-positive bacterial pathogen of grasses.

Technology feature: The genome jigsaw.

Subscribe for blog updates:

Filter by topic

Talk with an expert

Antimicrobial resistance research

Subscribe for blog updates:

Filter by topic

Talk with an expert