Although the accuracy of the human reference genome is critical for basic and clinical research, structural variants (SVs) have been difficult to assess because data capable of resolving them have been limited. To address potential bias, we sequenced a diversity panel of nine human genomes to high depth using long-read, single-molecule, real-time sequencing data. Systematically identifying and merging SVs =50 bp in length for these nine and one public genome yielded 83,909 sequence-resolved insertions, deletions, and inversions. Among these, 2,839 (2.0 Mbp) are shared among all discovery genomes with an additional 13,349 (6.9 Mbp) present in the majority of humans,…
Strain level microbiome profiling is needed for a full understanding of how microbial communities influence human health. Microbiome profiling of rRNA gene amplicons is a well-understood method that is rapid and inexpensive, but standard 16S rRNA gene methods generally cannot differentiate closely related strains. Whole genome/shotgun microbiome profiling is considered a higher-resolution alternative, but with decreased throughput and significantly increased sequencing costs and analysis burden. With both methods there are also challenges with microbial lysis, DNA preparation, and taxonomic analysis. Specialized microbiome-focused protocols were developed to achieve strain-level taxonomic differentiation using a rapid, high throughput rRNA gene assay. The protocol…
Euan Ashley from Stanford University started with the premise that while current efforts in the field of genomics medicine address 30% of patient cases, there’s a need for new approaches to make sense of the remaining 70%. Toward that end, he said that accurately calling structural variants is a major need. In one translational research example, Ashley said that SMRT Sequencing with the Sequel System allowed his team to identify six potentially causative genes in an individual with complex and varied symptoms; one gene was associated with Carney syndrome, which was a match for the person’s physiology and was later…
Jay Shendure, a Professor in the Department of Genome Sciences at the University of Washington School of Medicine explores the role of exome sequencing in clinical genomics. In this Podcast he discusses his views on the current and future roles of sequencing in diagnosing Mendelian disorders and investigation of complex regions of the genome.
In this ASHG 2017 presentation, Charles Lee of The Jackson Laboratory for Genomic Medicine presented work from the Human Genome Structural Variation Consortium. He shared data from efforts to utilize multiple platforms for the comprehensive discovery of structural variations—including insertions, deletions, inversions and mobile element insertions—in individual genomes. By combining various technologies, this research identified 7 times more structural variation per person than was previously known to exist.
Howard Jacob, Chief Genomics Officer at the HudsonAlpha Institute for Biotechnology, explored the role of genomics in diagnosing rare diseases. In this podcast he shared his views on the economics of clinical sequencing and how long-read sequencing is advancing the ability to sequence an individual’s genome –de novo– and use structural variant calling to make clinical diagnoses. He concluded with the hurdles limiting adoption of clinical sequencing and his vision for the future of genomic medicine.
To start Day 1 of the PacBio User Group Meeting, Jonas Korlach, PacBio CSO, provides an update on the latest releases and performance metrics for the Sequel II System. The longest reads generated on this system with the SMRT Cell 8M now go beyond 175,000 bases, while maintaining extremely high accuracy. HiFi mode, for example, uses circular consensus sequencing to achieve accuracy of Q40 or even Q50.
In this webinar, Dr. Ashby gives attendees a brief update on PacBio’s metagenomics solutions on the Sequel II System. Then, Dr. Ma, University of Maryland School of Medicine, discusses her work using long read sequencing to identify high-resolution microbial biomarkers associated with leaky gut syndrome in premature infants. Finally, Dr. Weinstock, The Jackson Laboratory, talks about the potential of highly accurate long reads to enable strain-level resolution of the human gut microbiome by resolving intraspecies variation in multiple copies of the 16S gene.
In this ASHG 2020 PacBio Workshop Emily Farrow of Children’s Mercy Kansas City, shares how the incorporation of long-read sequencing into the Genomic Answers for Kids research study is increasing diagnostic yields through the identification of novel genetic variation. Emily highlights several cases in which PacBio HiFi sequencing was able to provide insights where short-read sequencing alone was inconclusive, due to limitations stemming from repetitive regions and large structural variants.
In this ASHG 2020 PacBio Workshop Jonas Korlach, CSO, shares how the new PacBio Sequel IIe System makes highly accurate long-read sequencing easy and affordable so?all scientists can gain comprehensive views of human genomes and transcriptomes. He goes on to provide updates on the applications including human WGS for variant detection, de novo genome assembly, single-cell full-length RNA sequencing, and targeted sequencing using PCR and No-Amp methods.
Our studies reveal that the oral colonizer and cause of infective endocarditis Streptococcus oralis subsp. dentisani displays a striking monolateral distribution of surface fibrils. Furthermore, our data suggest that these fibrils impact the structure of adherent bacterial chains. Mutagenesis studies indicate that these fibrils are dependent on three serine-rich repeat proteins (SRRPs), here named fibril-associated protein A (FapA), FapB, and FapC, and that each SRRP forms a different fibril with a distinct distribution. SRRPs are a family of bacterial adhesins that have diverse roles in adhesion and that can bind to different receptors through modular nonrepeat region domains. Amino acid…
We have developed a computational method based on polyploid phasing of long sequence reads to resolve collapsed regions of segmental duplications within genome assemblies. Segmental Duplication Assembler (SDA; https://github.com/mvollger/SDA ) constructs graphs in which paralogous sequence variants define the nodes and long-read sequences provide attraction and repulsion edges, enabling the partition and assembly of long reads corresponding to distinct paralogs. We apply it to single-molecule, real-time sequence data from three human genomes and recover 33-79 megabase pairs (Mb) of duplications in which approximately half of the loci are diverged (99.9%) and that the diverged sequence corresponds to copy-number-variable paralogs that…
In the past several years, single-molecule sequencing platforms, such as those by Pacific Biosciences and Oxford Nanopore Technologies, have become available to researchers and are currently being tested for clinical applications. They offer exceptionally long reads that permit direct sequencing through regions of the genome inaccessible or difficult to analyze by short-read platforms. This includes disease-causing long repetitive elements, extreme GC content regions, and complex gene loci. Similarly, these platforms enable structural variation characterization at previously unparalleled resolution and direct detection of epigenetic marks in native DNA. Here, we review how these technologies are opening up new clinical avenues that…
In order to provide a comprehensive resource for human structural variants (SVs), we generated long-read sequence data and analyzed SVs for fifteen human genomes. We sequence resolved 99,604 insertions, deletions, and inversions including 2,238 (1.6 Mbp) that are shared among all discovery genomes with an additional 13,053 (6.9 Mbp) present in the majority, indicating minor alleles or errors in the reference. Genotyping in 440 additional genomes confirms the most common SVs in unique euchromatin are now sequence resolved. We report a ninefold SV bias toward the last 5 Mbp of human chromosomes with nearly 55% of all VNTRs (variable number…
The antibody repertoire of Bos taurus is characterized by a subset of variable heavy (VH) chain regions with ultralong third complementarity determining regions (CDR3) which, compared to other species, can provide a potent response to challenging antigens like HIV env. These unusual CDR3 can range to over seventy highly diverse amino acids in length and form unique ß-ribbon ‘stalk’ and disulfide bonded ‘knob’ structures, far from the typical antigen binding site. The genetic components and processes for forming these unusual cattle antibody VH CDR3 are not well understood. Here we analyze sequences of Bos taurus antibody VH domains and find…