PacBio assembly Archives - Page 3 of 11

July 19, 2019 |

The genome of Chenopodium quinoa.

Chenopodium quinoa (quinoa) is a highly nutritious grain identified as an important crop to improve world food security. Unfortunately, few resources are available to facilitate its genetic improvement. Here we report the assembly of a high-quality, chromosome-scale reference genome sequence for quinoa, which was produced using single-molecule real-time sequencing in combination with optical, chromosome-contact and genetic maps. We also report the sequencing of two diploids from the ancestral gene pools of quinoa, which enables the identification of sub-genomes in quinoa, and reduced-coverage genome sequences for 22 other samples of the allotetraploid goosefoot complex. The genome sequence facilitated the identification of the transcription factor likely to control the production of anti-nutritional triterpenoid saponins found in quinoa seeds, including a mutation that appears to cause alternative splicing and a premature stop codon in sweet quinoa strains. These genomic resources are an important first step towards the genetic improvement of quinoa.

July 19, 2019 |

The impact of third generation genomic technologies on plant genome assembly.

Since the introduction of next generation sequencing, plant genome assembly projects do not need to rely on dedicated research facilities or community-wide consortia anymore, even individual research groups can sequence and assemble the genomes they are interested in. However, such assemblies are typically not based on the entire breadth of genomic technologies including genetic and physical maps and their contiguities tend to be low compared to the full-length gold standard reference sequences. Recently emerging third generation genomic technologies like long-read sequencing or optical mapping promise to bridge this quality gap and enable simple and cost-effective solutions for chromosomal-level assemblies.

July 19, 2019 |

Chromosomal integration of the Klebsiella pneumoniae carbapenemase gene, blaKPC, in Klebsiella species is elusive but not rare.

Carbapenemase genes in Enterobacteriaceae are mostly described as being plasmid associated. However, the genetic context of carbapenemase genes is not always confirmed in epidemiological surveys, and the frequency of their chromosomal integration therefore is unknown. A previously sequenced collection of blaKPC-positive Enterobacteriaceae from a single U.S. institution (2007 to 2012; n = 281 isolates from 182 patients) was analyzed to identify chromosomal insertions of Tn4401, the transposon most frequently harboring blaKPC Using a combination of short- and long-read sequencing, we confirmed five independent chromosomal integration events from 6/182 (3%) patients, corresponding to 15/281 (5%) isolates. Three patients had isolates identified by perirectal screening, and three had infections which were all successfully treated. When a single copy of blaKPC was in the chromosome, one or both of the phenotypic carbapenemase tests were negative. All chromosomally integrated blaKPC genes were from Klebsiella spp., predominantly K. pneumoniae clonal group 258 (CG258), even though these represented only a small proportion of the isolates. Integration occurred via IS15-?I-mediated transposition of a larger, composite region encompassing Tn4401 at one locus of chromosomal integration, seen in the same strain (K. pneumoniae ST340) in two patients. In summary, we identified five independent chromosomal integrations of blaKPC in a large outbreak, demonstrating that this is not a rare event. blaKPC was more frequently integrated into the chromosome of epidemic CG258 K. pneumoniae lineages (ST11, ST258, and ST340) and was more difficult to detect by routine phenotypic methods in this context. The presence of chromosomally integrated blaKPC within successful, globally disseminated K. pneumoniae strains therefore is likely underestimated. Copyright © 2017 Mathers et al.

July 19, 2019 |

Single-molecule sequencing and chromatin conformation capture enable de novo reference assembly of the domestic goat genome.

The decrease in sequencing cost and increased sophistication of assembly algorithms for short-read platforms has resulted in a sharp increase in the number of species with genome assemblies. However, these assemblies are highly fragmented, with many gaps, ambiguities, and errors, impeding downstream applications. We demonstrate current state of the art for de novo assembly using the domestic goat (Capra hircus) based on long reads for contig formation, short reads for consensus validation, and scaffolding by optical and chromatin interaction mapping. These combined technologies produced what is, to our knowledge, the most continuous de novo mammalian assembly to date, with chromosome-length scaffolds and only 649 gaps. Our assembly represents a ~400-fold improvement in continuity due to properly assembled gaps, compared to the previously published C. hircus assembly, and better resolves repetitive structures longer than 1 kb, representing the largest repeat family and immune gene complex yet produced for an individual of a ruminant species.

July 19, 2019 |

Genomic confirmation of vancomycin-resistant Enterococcus transmission from deceased donor to liver transplant recipient.

In a liver transplant recipient with vancomycin-resistant Enterococcus (VRE) surgical site and bloodstream infection, a combination of pulsed-field gel electrophoresis, multilocus sequence typing, and whole genome sequencing identified that donor and recipient VRE isolates were highly similar when compared to time-matched hospital isolates. Comparison of de novo assembled isolate genomes was highly suggestive of transplant transmission rather than hospital-acquired transmission and also identified subtle internal rearrangements between donor and recipient missed by other genomic approaches. Given the improved resolution, whole-genome assembly of pathogen genomes is likely to become an essential tool for investigation of potential organ transplant transmissions.

July 19, 2019 |

Genomic structure of the horse major histocompatibility complex class II region resolved using PacBio long-read sequencing technology.

The mammalian Major Histocompatibility Complex (MHC) region contains several gene families characterized by highly polymorphic loci with extensive nucleotide diversity, copy number variation of paralogous genes, and long repetitive sequences. This structural complexity has made it difficult to construct a reliable reference sequence of the horse MHC region. In this study, we used long-read single molecule, real-time (SMRT) sequencing technology from Pacific Biosciences (PacBio) to sequence eight Bacterial Artificial Chromosome (BAC) clones spanning the horse MHC class II region. The final assembly resulted in a 1,165,328?bp continuous gap free sequence with 35 manually curated genomic loci of which 23 were considered to be functional and 12 to be pseudogenes. In comparison to the MHC class II region in other mammals, the corresponding region in horse shows extraordinary copy number variation and different relative location and directionality of the Eqca-DRB, -DQA, -DQB and -DOB loci. This is the first long-read sequence assembly of the horse MHC class II region with rigorous manual gene annotation, and it will serve as an important resource for association studies of immune-mediated equine diseases and for evolutionary analysis of genetic diversity in this region.

July 19, 2019 |

Single-molecule sequencing resolves the detailed structure of complex satellite DNA loci in Drosophila melanogaster.

Highly repetitive satellite DNA (satDNA) repeats are found in most eukaryotic genomes. SatDNAs are rapidly evolving and have roles in genome stability and chromosome segregation. Their repetitive nature poses a challenge for genome assembly and makes progress on the detailed study of satDNA structure difficult. Here, we use single-molecule sequencing long reads from Pacific Biosciences (PacBio) to determine the detailed structure of all major autosomal complex satDNA loci in Drosophila melanogaster, with a particular focus on the 260-bp and Responder satellites. We determine the optimal de novo assembly methods and parameter combinations required to produce a high-quality assembly of these previously unassembled satDNA loci and validate this assembly using molecular and computational approaches. We determined that the computationally intensive PBcR-BLASR assembly pipeline yielded better assemblies than the faster and more efficient pipelines based on the MHAP hashing algorithm, and it is essential to validate assemblies of repetitive loci. The assemblies reveal that satDNA repeats are organized into large arrays interrupted by transposable elements. The repeats in the center of the array tend to be homogenized in sequence, suggesting that gene conversion and unequal crossovers lead to repeat homogenization through concerted evolution, although the degree of unequal crossing over may differ among complex satellite loci. We find evidence for higher-order structure within satDNA arrays that suggest recent structural rearrangements. These assemblies provide a platform for the evolutionary and functional genomics of satDNAs in pericentric heterochromatin. © 2017 Khost et al.; Published by Cold Spring Harbor Laboratory Press.

July 19, 2019 |

Complete genome sequences of isolates of Enterococcus faecium sequence type 117, a globally disseminated multidrug-resistant clone.

The emergence of nosocomial infections by multidrug-resistant sequence type 117 (ST117) Enterococcus faecium has been reported in several European countries. ST117 has been detected in Spanish hospitals as one of the main causes of bloodstream infections. We analyzed genome variations of ST117 strains isolated in Madrid and describe the first ST117 closed genome sequences. Copyright © 2017 Tedim et al.

July 19, 2019 |

Contrasting evolutionary genome dynamics between domesticated and wild yeasts.

Structural rearrangements have long been recognized as an important source of genetic variation, with implications in phenotypic diversity and disease, yet their detailed evolutionary dynamics remain elusive. Here we use long-read sequencing to generate end-to-end genome assemblies for 12 strains representing major subpopulations of the partially domesticated yeast Saccharomyces cerevisiae and its wild relative Saccharomyces paradoxus. These population-level high-quality genomes with comprehensive annotation enable precise definition of chromosomal boundaries between cores and subtelomeres and a high-resolution view of evolutionary genome dynamics. In chromosomal cores, S. paradoxus shows faster accumulation of balanced rearrangements (inversions, reciprocal translocations and transpositions), whereas S. cerevisiae accumulates unbalanced rearrangements (novel insertions, deletions and duplications) more rapidly. In subtelomeres, both species show extensive interchromosomal reshuffling, with a higher tempo in S. cerevisiae. Such striking contrasts between wild and domesticated yeasts are likely to reflect the influence of human activities on structural genome evolution.

July 19, 2019 |

Multiple independent changes in mitochondrial genome conformation in chlamydomonadalean algae

Chlamydomonadalean green algae are no stranger to linear mitochondrial genomes, particularly members of the Reinhardtinia clade. At least nine different Reinhardtinia species are known to have linear mitochondrial DNAs (mtDNAs), including the model species Chlamydomonas reinhardtii. Thus, it is no surprise that some have suggested that the most recent common ancestor of the Reinhardtinia clade had a linear mtDNA. But the recent uncovering of circular-mapping mtDNAs in a range of Reinhardtinia algae, such as Volvox carteri and Tetrabaena socialis, has shed doubt on this hypothesis. Here, we explore mtDNA sequence and structure within the colonial Reinhardtinia algae Yamagishiella unicocca and Eudorina sp. NIES-3984, which occupy phylogenetically intermediate positions between species with opposing mtDNA mapping structures. Sequencing and gel electrophoresis data indicate that Y. unicocca has a linear monomeric mitochondrial genome with long (3?kb) palindromic telomeres. Conversely, the mtDNA of Eudorina sp., despite having an identical gene order to that of Y. unicocca, assembled as a circular-mapping molecule. Restriction digests of Eudorina sp. mtDNA supported its circular map, but also revealed a linear monomeric form with a matching architecture and gene order to the Y. unicocca mtDNA. Based on these data, we suggest that there have been at least three separate shifts in mtDNA conformation in the Reinhardtinia, and that the common ancestor of this clade had a linear monomeric mitochondrial genome with palindromic telomeres.

July 19, 2019 |

Evolutionary restoration of fertility in an interspecies hybrid yeast, by whole-genome duplication after a failed mating-type switch.

Many interspecies hybrids have been discovered in yeasts, but most of these hybrids are asexual and can replicate only mitotically. Whole-genome duplication has been proposed as a mechanism by which interspecies hybrids can regain fertility, restoring their ability to perform meiosis and sporulate. Here, we show that this process occurred naturally during the evolution of Zygosaccharomyces parabailii, an interspecies hybrid that was formed by mating between 2 parents that differed by 7% in genome sequence and by many interchromosomal rearrangements. Surprisingly, Z. parabailii has a full sexual cycle and is genetically haploid. It goes through mating-type switching and autodiploidization, followed by immediate sporulation. We identified the key evolutionary event that enabled Z. parabailii to regain fertility, which was breakage of 1 of the 2 homeologous copies of the mating-type (MAT) locus in the hybrid, resulting in a chromosomal rearrangement and irreparable damage to 1 MAT locus. This rearrangement was caused by HO endonuclease, which normally functions in mating-type switching. With 1 copy of MAT inactivated, the interspecies hybrid now behaves as a haploid. Our results provide the first demonstration that MAT locus damage is a naturally occurring evolutionary mechanism for whole-genome duplication and restoration of fertility to interspecies hybrids. The events that occurred in Z. parabailii strongly resemble those postulated to have caused ancient whole-genome duplication in an ancestor of Saccharomyces cerevisiae.

July 19, 2019 |

A case study into microbial genome assembly gap sequences and finishing strategies.

This study characterized regions of DNA which remained unassembled by either PacBio and Illumina sequencing technologies for seven bacterial genomes. Two genomes were manually finished using bioinformatics and PCR/Sanger sequencing approaches and regions not assembled by automated software were analyzed. Gaps present within Illumina assemblies mostly correspond to repetitive DNA regions such as multiple rRNA operon sequences. PacBio gap sequences were evaluated for several properties such as GC content, read coverage, gap length, ability to form strong secondary structures, and corresponding annotations. Our hypothesis that strong secondary DNA structures blocked DNA polymerases and contributed to gap sequences was not accepted. PacBio assemblies had few limitations overall and gaps were explained as cumulative effect of lower than average sequence coverage and repetitive sequences at contig termini. An important aspect of the present study is the compilation of biological features that interfered with assembly and included active transposons, multiple plasmid sequences, phage DNA integration, and large sequence duplication. Our targeted genome finishing approach and systematic evaluation of the unassembled DNA will be useful for others looking to close, finish, and polish microbial genome sequences.

July 19, 2019 |

PacBio but not Illumina technology can achieve fast, accurate and complete closure of the high GC, complex Burkholderia pseudomallei two-chromosome genome

Although PacBio third-generation sequencers have improved the read lengths of genome sequencing which facilitates the assembly of complete genomes, no study has reported success in using PacBio data alone to completely sequence a two-chromosome bacterial genome from a single library in a single run. Previous studies using earlier versions of sequencing chemistries have at most been able to finish bacterial genomes containing only one chromosome with de novo assembly. In this study, we compared the robustness of PacBio RS II, using one SMRT cell and the latest P6-C4 chemistry, with Illumina HiSeq 1500 in sequencing the genome of Burkholderia pseudomallei, a bacterium which contains two large circular chromosomes, very high G+C content of 68–69%, highly repetitive regions and substantial genomic diversity, and represents one of the largest and most complex bacterial genomes sequenced, using a reference genome generated by hybrid assembly using PacBio and Illumina datasets with subsequent manual validation. Results showed that PacBio data with de novo assembly, but not Illumina, was able to completely sequence the B. pseudomallei genome without any gaps or mis-assemblies. The two large contigs of the PacBio assembly aligned unambiguously to the reference genome, sharing >99.9% nucleotide identities. Conversely, Illumina data assembled using three different assemblers resulted in fragmented assemblies (201–366 contigs), sharing only 92.2–100% and 92.0–100% nucleotide identities to chromosomes I and II reference sequences, respectively, with no indication that the B. pseudomallei genome consisted of two chromosomes with four copies of ribosomal operons. Among all assemblies, the PacBio assembly recovered the highest number of core and virulence proteins, and housekeeping genes based on whole-genome multilocus sequence typing (wgMLST). Most notably, assembly solely based on PacBio outperformed even hybrid assembly using both PacBio and Illumina datasets. Hybrid approach generated only 74 contigs, while the PacBio data alone with de novo assembly achieved complete closure of the two-chromosome B. pseudomallei genome without additional costly bench work and further sequencing. PacBio RS II using P6-C4 chemistry is highly robust and cost-effective and should be the platform of choice in sequencing bacterial genomes, particularly for those that are well-known to be difficult-to-sequence.

July 19, 2019 |

PacBio sequencing reveals transposable element as a key contributor to genomic plasticity and virulence variation in Magnaporthe oryzae.

The sustainable cultivation of rice, which serves as staple food crop for more than half of the world’s population, is under serious threat due to the huge yield losses inflicted by rice blast disease caused by the globally destructive fungus Magnaporthe oryzae (Pyricularia oryzae) (Dean et al., 2012, Nalley et al., 2016, Deng et al., 2017). This filamentous ascomycete fungus is also capable of causing blast infection on other economically important cereal crops, including wheat, millet, and barley, making it the world’s most important plant pathogenic fungus (Zhong et al., 2016). The advent of whole-genome sequencing technology and the subsequent deployment of next-generation sequencing (NGS) strategies have successfully generated genome assemblies for over 50 isolates of M. oryzae, which have played an instrumental role in enhancing our understanding of how rice blast fungus undertakes host adaptation, host specificity, and host range expansion to overcome host resistance (Dean et al., 2005, Xue et al., 2012, Wu et al., 2015, Zhang et al., 2016). However, research findings obtained from comparative genomic studies conducted using the NGS-assembled genome do not present an in-depth account of the genomic features that contribute to the prevailing genomic variations among M. oryzae species, because NGS assemblies are highly fragmented and lack most of the lineage-specific (LS) regions, which are more plastic than the core genome and enriched with repeats and effector proteins (Raffaele and Kamoun, 2012, Faino et al., 2016).

July 19, 2019 |

Gapless genome assembly of Colletotrichum higginsianum reveals chromosome structure and association of transposable elements with secondary metabolite gene clusters.

The ascomycete fungus Colletotrichum higginsianum causes anthracnose disease of brassica crops and the model plant Arabidopsis thaliana. Previous versions of the genome sequence were highly fragmented, causing errors in the prediction of protein-coding genes and preventing the analysis of repetitive sequences and genome architecture. Here, we re-sequenced the genome using single-molecule real-time (SMRT) sequencing technology and, in combination with optical map data, this provided a gapless assembly of all twelve chromosomes except for the ribosomal DNA repeat cluster on chromosome 7. The more accurate gene annotation made possible by this new assembly revealed a large repertoire of secondary metabolism (SM) key genes (89) and putative biosynthetic pathways (77 SM gene clusters). The two mini-chromosomes differed from the ten core chromosomes in being repeat- and AT-rich and gene-poor but were significantly enriched with genes encoding putative secreted effector proteins. Transposable elements (TEs) were found to occupy 7% of the genome by length. Certain TE families showed a statistically significant association with effector genes and SM cluster genes and were transcriptionally active at particular stages of fungal development. All 24 subtelomeres were found to contain one of three highly-conserved repeat elements which, by providing sites for homologous recombination, were probably instrumental in four segmental duplications.The gapless genome of C. higginsianum provides access to repeat-rich regions that were previously poorly assembled, notably the mini-chromosomes and subtelomeres, and allowed prediction of the complete SM gene repertoire. It also provides insights into the potential role of TEs in gene and genome evolution and host adaptation in this asexual pathogen.

Auto Tag: PacBio assembly

The genome of Chenopodium quinoa.

The impact of third generation genomic technologies on plant genome assembly.

Chromosomal integration of the Klebsiella pneumoniae carbapenemase gene, blaKPC, in Klebsiella species is elusive but not rare.

Single-molecule sequencing and chromatin conformation capture enable de novo reference assembly of the domestic goat genome.

Genomic confirmation of vancomycin-resistant Enterococcus transmission from deceased donor to liver transplant recipient.

Genomic structure of the horse major histocompatibility complex class II region resolved using PacBio long-read sequencing technology.

Single-molecule sequencing resolves the detailed structure of complex satellite DNA loci in Drosophila melanogaster.

Complete genome sequences of isolates of Enterococcus faecium sequence type 117, a globally disseminated multidrug-resistant clone.

Contrasting evolutionary genome dynamics between domesticated and wild yeasts.

Multiple independent changes in mitochondrial genome conformation in chlamydomonadalean algae

Evolutionary restoration of fertility in an interspecies hybrid yeast, by whole-genome duplication after a failed mating-type switch.

A case study into microbial genome assembly gap sequences and finishing strategies.

PacBio but not Illumina technology can achieve fast, accurate and complete closure of the high GC, complex Burkholderia pseudomallei two-chromosome genome

PacBio sequencing reveals transposable element as a key contributor to genomic plasticity and virulence variation in Magnaporthe oryzae.

Gapless genome assembly of Colletotrichum higginsianum reveals chromosome structure and association of transposable elements with secondary metabolite gene clusters.

Subscribe for blog updates:

Filter by topic

Talk with an expert

ALS case study

Subscribe for blog updates:

Filter by topic

Talk with an expert