Although the accuracy of the human reference genome is critical for basic and clinical research, structural variants (SVs) have been difficult to assess because data capable of resolving them have been limited. To address potential bias, we sequenced a diversity panel of nine human genomes to high depth using long-read, single-molecule, real-time sequencing data. Systematically identifying and merging SVs =50 bp in length for these nine and one public genome yielded 83,909 sequence-resolved insertions, deletions, and inversions. Among these, 2,839 (2.0 Mbp) are shared among all discovery genomes with an additional 13,349 (6.9 Mbp) present in the majority of humans, indicating minor alleles or errors in the reference, which is partially explained by an enrichment for GC-content and repetitive DNA. Genotyping 83% of these in 290 additional genomes confirms that at least one allele of the most common SVs in unique euchromatin are now sequence-resolved. We observe a 9-fold increase within 5 Mbp of chromosome telomeric ends and correlation with de novo single-nucleotide variant mutations showing that such variation is nonrandomly distributed defining potential hotspots of mutation. We identify SVs affecting coding and noncoding regulatory loci improving annotation and interpretation of functional variation. To illustrate the utility of sequence-resolved SVs in resequencing experiments, we mapped 30 diverse high-coverage Illumina-sequenced samples to GRCh38 with and without contigs containing SV insertions as alternate sequences, and we found these additional sequences recover 6.4% of unmapped reads. For reads mapped within the SV insertion, 25.7% have a better mapping quality, and 18.7% improved by 1,000-fold or more. We reveal 72,964 occurrences of 15,814 unique variants that were not discoverable with the reference sequence alone, and we note that 7% of the insertions contain an SV in at least one sample indicating that there are additional alleles in the population that remain to be discovered. These data provide the framework to construct a canonical human reference and a resource for developing advanced representations capable of capturing allelic diversity. We present a summary of our findings and discuss ideas for revealing variation that was once difficult to ascertain.
ASHG PacBio Workshop: SMRT Sequencing as a translational research tool to investigate germline, somatic and infectious diseases
Melissa Laird Smith discussed how the Icahn School of Medicine at Mount Sinai uses long-read sequencing for translational research. She gave several examples of targeted sequencing projects run on the…
In this ASHG 2020 PacBio Workshop Jonas Korlach, CSO, shares how the new PacBio Sequel IIe System makes highly accurate long-read sequencing easy and affordable so?all scientists can gain comprehensive…
Genetic variation in the conjugative plasmidome of a hospital effluent multidrug resistant Escherichia coli strain.
Bacteria harboring conjugative plasmids have the potential for spreading antibiotic resistance through horizontal gene transfer. It is described that the selection and dissemination of antibiotic resistance is enhanced by stressors, like metals or antibiotics, which can occur as environmental contaminants. This study aimed at unveiling the composition of the conjugative plasmidome of a hospital effluent multidrug resistant Escherichia coli strain (H1FC54) under different mating conditions. To meet this objective, plasmid pulsed field gel electrophoresis, optical mapping analyses and DNA sequencing were used in combination with phenotype analysis. Strain H1FC54 was observed to harbor five plasmids, three of which were conjugative and two of these, pH1FC54_330 and pH1FC54_140, contained metal and antibiotic resistance genes. Transconjugants obtained in the absence or presence of tellurite (0.5?µM or 5?µM), arsenite (0.5?µM, 5?µM or 15?µM) or ceftazidime (10?mg/L) and selected in the presence of sodium azide (100?mg/L) and tetracycline (16?mg/L) presented distinct phenotypes, associated with the acquisition of different plasmid combinations, including two co-integrate plasmids, of 310 kbp and 517 kbp. The variable composition of the conjugative plasmidome, the formation of co-integrates during conjugation, as well as the transfer of non-transferable plasmids via co-integration, and the possible association between antibiotic, arsenite and tellurite tolerance was demonstrated. These evidences bring interesting insights into the comprehension of the molecular and physiological mechanisms that underlie antibiotic resistance propagation in the environment. Copyright © 2019 Elsevier Ltd. All rights reserved.
Long-read sequencing, CENP-A ChIP, and chromatin fiber imaging reveal the composition and organization of Drosophila melanogaster centromeres, which have long remained elusive despite the high quality of this species’ genome. assembly.