Data management Archives

June 1, 2021 |

The “Art” of shotgun sequencing

2015 SMRT Informatics Developers Conference Presentation Slides: Jason Chin of PacBio highlighted some of the challenges for shotgun assembly while suggesting some potential solutions to obtain diploid assemblies, including the FALCON method.

June 1, 2021 |

Multiplexed complete microbial genomes on the Sequel System

Microbes play an important role in nearly every part of our world, as they affect human health, our environment, agriculture, and aid in waste management. Complete closed genome sequences, which have become the gold standard with PacBio long-read sequencing, can be key to understanding microbial functional characteristics. However, input requirements, consumables costs, and the labor required to prepare and sequence a microbial genome have in the past put PacBio sequencing out of reach for some larger projects. We have developed a multiplexed library prep approach that is simple, fast, and cost-effective, and can produce 4 to 16 closed bacterial genomes from one Sequel SMRT Cell. Additionally, we are introducing a streamlined analysis pipeline for processing multiplexed genome sequence data through de novo HGAP assembly, making the entire process easy for lab personnel to perform. Here we present the entire workflow from shearing through assembly, with times for each step. We show HGAP assembly results with single or very few contigs from bacteria from different size genomes, sequenced without or with size selection. These data illustrate the benefits and potential of the PacBio multiplexed library prep and the Sequel System for sequencing large numbers of microbial genomes.

June 1, 2021 |

Single chromosomal genome assemblies on the Sequel System with Circulomics high molecular weight DNA extraction for microbes

Background: The Nanobind technology from Circulomics provides an elegant HMW DNA extraction solution for genome sequencing of Gram-positive and -negative microbes. Nanobind is a nanostructured magnetic disk that can be used for rapid extraction of high molecular weight (HMW) DNA from diverse sample types including cultured cells, blood, plant nuclei, and bacteria. Processing can be completed in <1 hour for most sample types and can be performed manually or automated with common instruments. Methods:We have validated several critical steps for generating high-quality microbial genome assemblies in a streamlined microbial multiplexing workflow. This new workflow enables high-volume, cost-effective sequencing of up to 16 microbes totaling 30 Mb in genome size on a single SMRT Cell 1M using a target shear size of 10 kb. We also evaluated this method on a pool of four “class 3” microbes that contain >7 kb repeats. Fragment size was increased to ~14 kb, with some fragments >30 kb. Results: Here we present a demonstration of these capabilities using isolates relevant to high-throughput sequencing applications, including common foodborne pathogens (Shigella, Listeria, Salmonella), and species often seen in hospital settings (Klebsiella, Staphylococcus). For nearly all microbes, including difficult-to-assemble class III microbes, we achieved complete de novo microbial assemblies of =5 chromosomal contigs with minimum quality scores of 40 (99.99% accuracy) using data from multiplexed SMRTbell libraries. Each library was sequenced on a single SMRT Cell 1M with the PacBio Sequel System and analyzed with streamlined SMRT Analysis assembly methods. Conclusions: We achieved high-quality, closed microbial genomes using a combination of Circulomics Nanobind extraction and PacBio SMRT Sequencing, along with a newly streamlined workflow that includes automated demultiplexing and push-button assembly.

February 5, 2021 |

Tutorial: SMRT Link overview

This tutorial provides a high-level overview of the features contained within the SMRT Link software. SMRT Link is the web-based end-to-end software workflow manager for run design and set-up on…

February 5, 2021 |

Tutorial: SMRT Link v5.0 overview

In this video Roberto Lleras shares new module-based features included in SMRT Link v5.0. He summarizes updates to data management, new applications for minor variant analysis and structural variant analysis…

February 5, 2021 |

Webinar: Beginner’s guide to PacBio SMRT Sequencing data analysis

PacBio SMRT Sequencing is fast changing the genomics space with its long reads and high consensus sequence accuracy, providing the most comprehensive view of the genome and transcriptome. In this…

February 5, 2021 |

PAG Conference: The impact of highly accurate PacBio sequence data on the assembly of a tetraploid rose

In this presentation at PAG 2020, Bart Nijland of Genetwister Technologies explains how his team set out to make a haplotype-aware assembly of the highly complex tetraploid Rosa x hybrida…

April 21, 2020 |

The comparative genomics and complex population history of Papio baboons.

Recent studies suggest that closely related species can accumulate substantial genetic and phenotypic differences despite ongoing gene flow, thus challenging traditional ideas regarding the genetics of speciation. Baboons (genus Papio) are Old World monkeys consisting of six readily distinguishable species. Baboon species hybridize in the wild, and prior data imply a complex history of differentiation and introgression. We produced a reference genome assembly for the olive baboon (Papio anubis) and whole-genome sequence data for all six extant species. We document multiple episodes of admixture and introgression during the radiation of Papio baboons, thus demonstrating their value as a model of complex evolutionary divergence, hybridization, and reticulation. These results help inform our understanding of similar cases, including modern humans, Neanderthals, Denisovans, and other ancient hominins.

April 21, 2020 |

Characterizing the major structural variant alleles of the human genome.

In order to provide a comprehensive resource for human structural variants (SVs), we generated long-read sequence data and analyzed SVs for fifteen human genomes. We sequence resolved 99,604 insertions, deletions, and inversions including 2,238 (1.6 Mbp) that are shared among all discovery genomes with an additional 13,053 (6.9 Mbp) present in the majority, indicating minor alleles or errors in the reference. Genotyping in 440 additional genomes confirms the most common SVs in unique euchromatin are now sequence resolved. We report a ninefold SV bias toward the last 5 Mbp of human chromosomes with nearly 55% of all VNTRs (variable number of tandem repeats) mapping to this portion of the genome. We identify SVs affecting coding and noncoding regulatory loci improving annotation and interpretation of functional variation. These data provide the framework to construct a canonical human reference and a resource for developing advanced representations capable of capturing allelic diversity. Copyright © 2018 Elsevier Inc. All rights reserved.

April 21, 2020 |

Computational aspects underlying genome to phenome analysis in plants.

Recent advances in genomics technologies have greatly accelerated the progress in both fundamental plant science and applied breeding research. Concurrently, high-throughput plant phenotyping is becoming widely adopted in the plant community, promising to alleviate the phenotypic bottleneck. While these technological breakthroughs are significantly accelerating quantitative trait locus (QTL) and causal gene identification, challenges to enable even more sophisticated analyses remain. In particular, care needs to be taken to standardize, describe and conduct experiments robustly while relying on plant physiology expertise. In this article, we review the state of the art regarding genome assembly and the future potential of pangenomics in plant research. We also describe the necessity of standardizing and describing phenotypic studies using the Minimum Information About a Plant Phenotyping Experiment (MIAPPE) standard to enable the reuse and integration of phenotypic data. In addition, we show how deep phenotypic data might yield novel trait-trait correlations and review how to link phenotypic data to genomic data. Finally, we provide perspectives on the golden future of machine learning and their potential in linking phenotypes to genomic features. © 2018 The Authors The Plant Journal published by John Wiley & Sons Ltd and Society for Experimental Biology.

April 21, 2020 |

Rapid and Focused Maturation of a VRC01-Class HIV Broadly Neutralizing Antibody Lineage Involves Both Binding and Accommodation of the N276-Glycan.

The VH1-2 restricted VRC01-class of antibodies targeting the HIV envelope CD4 binding site are a major focus of HIV vaccine strategies. However, a detailed analysis of VRC01-class antibody development has been limited by the rare nature of these responses during natural infection and the lack of longitudinal sampling of such responses. To inform vaccine strategies, we mapped the development of a VRC01-class antibody lineage (PCIN63) in the subtype C infected IAVI Protocol C neutralizer PC063. PCIN63 monoclonal antibodies had the hallmark VRC01-class features and demonstrated neutralization breadth similar to the prototype VRC01 antibody, but were 2- to 3-fold less mutated. Maturation occurred rapidly within ~24 months of emergence of the lineage and somatic hypermutations accumulated at key contact residues. This longitudinal study of broadly neutralizing VRC01-class antibody lineage reveals early binding to the N276-glycan during affinity maturation, which may have implications for vaccine design.Copyright © 2019 The Authors. Published by Elsevier Inc. All rights reserved.

April 21, 2020 |

Genomic Characterization of a Newly Isolated Rhizobacteria Sphingomonas panacis Reveals Plant Growth Promoting Effect to Rice

This article reports the full genome sequence of Sphingomonas panacis DCY99T (=KCTC 42347T =JCM30806T), which is a Gram-negative rod-shaped, non-spore forming, motile bacterium isolated from rusty ginseng root in South Korea. A draft genome of S. panacis DCY99T and a single circular plasmid were generated using the PacBio platform. Antagonistic activity experiment showed S. panacis DCY99T has the plant growth promoting effect. Thus, the genome sequence of S. panacis DCY99T may contribute to biotechnological application of the genus Sphingomonas in agriculture.

April 21, 2020 |

Copy-number variants in clinical genome sequencing: deployment and interpretation for rare and undiagnosed disease.

Current diagnostic testing for genetic disorders involves serial use of specialized assays spanning multiple technologies. In principle, genome sequencing (GS) can detect all genomic pathogenic variant types on a single platform. Here we evaluate copy-number variant (CNV) calling as part of a clinically accredited GS test.We performed analytical validation of CNV calling on 17 reference samples, compared the sensitivity of GS-based variants with those from a clinical microarray, and set a bound on precision using orthogonal technologies. We developed a protocol for family-based analysis of GS-based CNV calls, and deployed this across a clinical cohort of 79 rare and undiagnosed cases.We found that CNV calls from GS are at least as sensitive as those from microarrays, while only creating a modest increase in the number of variants interpreted (~10 CNVs per case). We identified clinically significant CNVs in 15% of the first 79 cases analyzed, all of which were confirmed by an orthogonal approach. The pipeline also enabled discovery of a uniparental disomy (UPD) and a 50% mosaic trisomy 14. Directed analysis of select CNVs enabled breakpoint level resolution of genomic rearrangements and phasing of de novo CNVs.Robust identification of CNVs by GS is possible within a clinical testing environment.

April 21, 2020 |

Complete genome sequence analysis of the thermoacidophilic verrucomicrobial methanotroph “Candidatus Methylacidiphilum kamchatkense” strain Kam1 and comparison with its closest relatives.

The candidate genus “Methylacidiphilum” comprises thermoacidophilic aerobic methane oxidizers belonging to the Verrucomicrobia phylum. These are the first described non-proteobacterial aerobic methane oxidizers. The genes pmoCAB, encoding the particulate methane monooxygenase do not originate from horizontal gene transfer from proteobacteria. Instead, the “Ca. Methylacidiphilum” and the sister genus “Ca. Methylacidimicrobium” represent a novel and hitherto understudied evolutionary lineage of aerobic methane oxidizers. Obtaining and comparing the full genome sequences is an important step towards understanding the evolution and physiology of this novel group of organisms.Here we present the closed genome of “Ca. Methylacidiphilum kamchatkense” strain Kam1 and a comparison with the genomes of its two closest relatives “Ca. Methylacidiphilum fumariolicum” strain SolV and “Ca. Methylacidiphilum infernorum” strain V4. The genome consists of a single 2,2 Mbp chromosome with 2119 predicted protein coding sequences. Genome analysis showed that the majority of the genes connected with metabolic traits described for one member of “Ca. Methylacidiphilum” is conserved between all three genomes. All three strains encode class I CRISPR-cas systems. The average nucleotide identity between “Ca. M. kamchatkense” strain Kam1 and strains SolV and V4 is =95% showing that they should be regarded as separate species. Whole genome comparison revealed a high degree of synteny between the genomes of strains Kam1 and SolV. In contrast, comparison of the genomes of strains Kam1 and V4 revealed a number of rearrangements. There are large differences in the numbers of transposable elements found in the genomes of the three strains with 12, 37 and 80 transposable elements in the genomes of strains Kam1, V4 and SolV respectively. Genomic rearrangements and the activity of transposable elements explain much of the genomic differences between strains. For example, a type 1h uptake hydrogenase is conserved between strains Kam1 and SolV but seems to have been lost from strain V4 due to genomic rearrangements.Comparing three closed genomes of “Ca. Methylacidiphilum” spp. has given new insights into the evolution of these organisms and revealed large differences in numbers of transposable elements between strains, the activity of these explains much of the genomic differences between strains.

April 21, 2020 |

High-coverage, long-read sequencing of Han Chinese trio reference samples.

Single-molecule long-read sequencing datasets were generated for a son-father-mother trio of Han Chinese descent that is part of the Genome in a Bottle (GIAB) consortium portfolio. The dataset was generated using the Pacific Biosciences Sequel System. The son and each parent were sequenced to an average coverage of 60 and 30, respectively, with N50 subread lengths between 16 and 18?kb. Raw reads and reads aligned to both the GRCh37 and GRCh38 are available at the NCBI GIAB ftp site (ftp://ftp-trace.ncbi.nlm.nih.gov/giab/ftp/data/ChineseTrio/). The GRCh38 aligned read data are archived in NCBI SRA (SRX4739017, SRX4739121, and SRX4739122). This dataset is available for anyone to develop and evaluate long-read bioinformatics methods.

Auto Tag: Data management

The “Art” of shotgun sequencing

Multiplexed complete microbial genomes on the Sequel System

Single chromosomal genome assemblies on the Sequel System with Circulomics high molecular weight DNA extraction for microbes

Tutorial: SMRT Link overview

Tutorial: SMRT Link v5.0 overview

Webinar: Beginner’s guide to PacBio SMRT Sequencing data analysis

PAG Conference: The impact of highly accurate PacBio sequence data on the assembly of a tetraploid rose

The comparative genomics and complex population history of Papio baboons.

Characterizing the major structural variant alleles of the human genome.

Rapid and Focused Maturation of a VRC01-Class HIV Broadly Neutralizing Antibody Lineage Involves Both Binding and Accommodation of the N276-Glycan.

Genomic Characterization of a Newly Isolated Rhizobacteria Sphingomonas panacis Reveals Plant Growth Promoting Effect to Rice

Copy-number variants in clinical genome sequencing: deployment and interpretation for rare and undiagnosed disease.

Complete genome sequence analysis of the thermoacidophilic verrucomicrobial methanotroph “Candidatus Methylacidiphilum kamchatkense” strain Kam1 and comparison with its closest relatives.

High-coverage, long-read sequencing of Han Chinese trio reference samples.

Subscribe for blog updates:

Filter by topic

Talk with an expert

ALS case study

Subscribe for blog updates:

Filter by topic

Talk with an expert