Menu
April 21, 2020  |  

Discovery of tandem and interspersed segmental duplications using high-throughput sequencing.

Several algorithms have been developed that use high-throughput sequencing technology to characterize structural variations (SVs). Most of the existing approaches focus on detecting relatively simple types of SVs such as insertions, deletions and short inversions. In fact, complex SVs are of crucial importance and several have been associated with genomic disorders. To better understand the contribution of complex SVs to human disease, we need new algorithms to accurately discover and genotype such variants. Additionally, due to similar sequencing signatures, inverted duplications or gene conversion events that include inverted segmental duplications are often characterized as simple inversions, likewise, duplications and gene conversions in direct orientation may be called as simple deletions. Therefore, there is still a need for accurate algorithms to fully characterize complex SVs and thus improve calling accuracy of more simple variants.We developed novel algorithms to accurately characterize tandem, direct and inverted interspersed segmental duplications using short read whole genome sequencing datasets. We integrated these methods to our TARDIS tool, which is now capable of detecting various types of SVs using multiple sequence signatures such as read pair, read depth and split read. We evaluated the prediction performance of our algorithms through several experiments using both simulated and real datasets. In the simulation experiments, using a 30× coverage TARDIS achieved 96% sensitivity with only 4% false discovery rate. For experiments that involve real data, we used two haploid genomes (CHM1 and CHM13) and one human genome (NA12878) from the Illumina Platinum Genomes set. Comparison of our results with orthogonal PacBio call sets from the same genomes revealed higher accuracy for TARDIS than state-of-the-art methods. Furthermore, we showed a surprisingly low false discovery rate of our approach for discovery of tandem, direct and inverted interspersed segmental duplications prediction on CHM1 (<5% for the top 50 predictions).TARDIS source code is available at https://github.com/BilkentCompGen/tardis, and a corresponding Docker image is available at https://hub.docker.com/r/alkanlab/tardis/.Supplementary data are available at Bioinformatics online. © The Author(s) 2019. Published by Oxford University Press. All rights reserved. For permissions, please e-mail: journals.permissions@oup.com.


April 21, 2020  |  

Genome-Scale Sequence Disruption Following Biolistic Transformation in Rice and Maize.

Biolistic transformation delivers nucleic acids into plant cells by bombarding the cells with microprojectiles, which are micron-scale, typically gold particles. Despite the wide use of this technique, little is known about its effect on the cell’s genome. We biolistically transformed linear 48-kb phage lambda and two different circular plasmids into rice (Oryza sativa) and maize (Zea mays) and analyzed the results by whole genome sequencing and optical mapping. Although some transgenic events showed simple insertions, others showed extreme genome damage in the form of chromosome truncations, large deletions, partial trisomy, and evidence of chromothripsis and breakage-fusion bridge cycling. Several transgenic events contained megabase-scale arrays of introduced DNA mixed with genomic fragments assembled by nonhomologous or microhomology-mediated joining. Damaged regions of the genome, assayed by the presence of small fragments displaced elsewhere, were often repaired without a trace, presumably by homology-dependent repair (HDR). The results suggest a model whereby successful biolistic transformation relies on a combination of end joining to insert foreign DNA and HDR to repair collateral damage caused by the microprojectiles. The differing levels of genome damage observed among transgenic events may reflect the stage of the cell cycle and the availability of templates for HDR. © 2019 American Society of Plant Biologists. All rights reserved.


April 21, 2020  |  

A 12-kb structural variation in progressive myoclonic epilepsy was newly identified by long-read whole-genome sequencing.

We report a family with progressive myoclonic epilepsy who underwent whole-exome sequencing but was negative for pathogenic variants. Similar clinical courses of a devastating neurodegenerative phenotype of two affected siblings were highly suggestive of a genetic etiology, which indicates that the survey of genetic variation by whole-exome sequencing was not comprehensive. To investigate the presence of a variant that remained unrecognized by standard genetic testing, PacBio long-read sequencing was performed. Structural variant (SV) detection using low-coverage (6×) whole-genome sequencing called 17,165 SVs (7,216 deletions and 9,949 insertions). Our SV selection narrowed down potential candidates to only five SVs (two deletions and three insertions) on the genes tagged with autosomal recessive phenotypes. Among them, a 12.4-kb deletion involving the CLN6 gene was the top candidate because its homozygous abnormalities cause neuronal ceroid lipofuscinosis. This deletion included the initiation codon and was found in a GC-rich region containing multiple repetitive elements. These results indicate the presence of a causal variant in a difficult-to-sequence region and suggest that such variants that remain enigmatic after the application of current whole-exome sequencing technology could be uncovered by unbiased application of long-read whole-genome sequencing.


April 21, 2020  |  

Fast and accurate genomic analyses using genome graphs.

The human reference genome serves as the foundation for genomics by providing a scaffold for alignment of sequencing reads, but currently only reflects a single consensus haplotype, thus impairing analysis accuracy. Here we present a graph reference genome implementation that enables read alignment across 2,800 diploid genomes encompassing 12.6 million SNPs and 4.0 million insertions and deletions (indels). The pipeline processes one whole-genome sequencing sample in 6.5?h using a system with 36?CPU cores. We show that using a graph genome reference improves read mapping sensitivity and produces a 0.5% increase in variant calling recall, with unaffected specificity. Structural variations incorporated into a graph genome can be genotyped accurately under a unified framework. Finally, we show that iterative augmentation of graph genomes yields incremental gains in variant calling accuracy. Our implementation is an important advance toward fulfilling the promise of graph genomes to radically enhance the scalability and accuracy of genomic analyses.


April 21, 2020  |  

A prophage and two ICESa2603-family integrative and conjugative elements (ICEs) carrying optrA in Streptococcus suis.

To investigate the presence and transfer of the oxazolidinone/phenicol resistance gene optrA and identify the genetic elements involved in the horizontal transfer of the optrA gene in Streptococcus suis.A total of 237 S. suis isolates were screened for the presence of the optrA gene by PCR. Whole-genome DNA of three optrA-positive strains was completely sequenced using the Illumina MiSeq and Pacbio RSII platforms. MICs were determined by broth microdilution. Transferability of the optrA gene in S. suis was investigated by conjugation. The presence of circular intermediates was examined by inverse PCR.The optrA gene was present in 11.8% (28/237) of the S. suis strains. In three strains, the optrA gene was flanked by two copies of IS1216 elements in the same orientation, located either on a prophage or on ICESa2603-family integrative and conjugative elements (ICEs), including one tandem ICE. In one isolate, the optrA-carrying ICE transferred with a frequency of 2.1?×?10-8. After the transfer, the transconjugant displayed elevated MICs of the respective antimicrobial agents. Inverse PCRs revealed that circular intermediates of different sizes were formed in the three optrA-carrying strains, containing one copy of the IS1216E element and the optrA gene alone or in combination with other resistance genes.A prophage and two ICESa2603-family ICEs (including one tandem ICE) associated with the optrA gene were identified in S. suis. The association of the optrA gene with the IS1216E elements and its location on either a prophage or ICEs will aid its horizontal transfer. © The Author(s) 2019. Published by Oxford University Press on behalf of the British Society for Antimicrobial Chemotherapy. All rights reserved. For permissions, please email: journals.permissions@oup.com.


April 21, 2020  |  

Retrospective whole-genome sequencing analysis distinguished PFGE and drug-resistance-matched retail meat and clinical Salmonella isolates.

Non-typhoidal Salmonella is a leading cause of outbreak and sporadic-associated foodborne illnesses in the United States. These infections have been associated with a range of foods, including retail meats. Traditionally, pulsed-field gel electrophoresis (PFGE) and antibiotic susceptibility testing (AST) have been used to facilitate public health investigations of Salmonella infections. However, whole-genome sequencing (WGS) has emerged as an alternative tool that can be routinely implemented. To assess its potential in enhancing integrated surveillance in Pennsylvania, USA, WGS was used to directly compare the genetic characteristics of 7 retail meat and 43 clinical historic Salmonella isolates, subdivided into 3 subsets based on PFGE and AST results, to retrospectively resolve their genetic relatedness and identify antimicrobial resistance (AMR) determinants. Single nucleotide polymorphism (SNP) analyses revealed that the retail meat isolates within S. Heidelberg, S. Typhimurium var. O5- subset 1 and S. Typhimurium var. O5- subset 2 were separated from each primary PFGE pattern-matched clinical isolate by 6-12, 41-96 and 21-81 SNPs, respectively. Fifteen resistance genes were identified across all isolates, including fosA7, a gene only recently found in a limited number of Salmonella and a =95?%?phenotype to genotype correlation was observed for all tested antimicrobials. Moreover, AMR was primarily plasmid-mediated in S. Heidelberg and S. Typhimurium var. O5- subset 2, whereas AMR was chromosomally carried in S. Typhimurium var. O5- subset 1. Similar plasmids were identified in both the retail meat and clinical isolates. Collectively, these data highlight the utility of WGS in retrospective analyses and enhancing integrated surveillance for Salmonella from multiple sources.


April 21, 2020  |  

Genetic characterization and potential molecular dissemination mechanism of tet(31) gene in Aeromonas caviae from an oxytetracycline wastewater treatment system.

Recently, the rarely reported tet(31) tetracycline resistance determinant was commonly found in Aeromonas salmonicida, Gallibacterium anatis, and Oblitimonas alkaliphila isolated from farming animals and related environment. However, its distribution in other bacteria and potential molecular dissemination mechanism in environment are still unknown. The purpose of this study was to investigate the potential mechanism underlying dissemination of tet(31) by analysing the tet(31)-carrying fragments in A. caviae strains isolated from an aerobic biofilm reactor treating oxytetracycline bearing wastewater. Twenty-three A. caviae strains were screened for the tet(31) gene by polymerase chain reaction (PCR). Three strains (two harbouring tet(31), one not) were subjected to whole genome sequencing using the PacBio RSII platform. Seventeen A. caviae strains carried the tet(31) gene and exhibited high resistance levels to oxytetracycline with minimum inhibitory concentrations (MICs) ranging from 256 to 512?mg/L. tet(31) was comprised of the transposon Tn6432 on the chromosome of A. caviae, and Tn6432 was also found in 15 additional tet(31)-positive A. caviae isolates by PCR. More important, Tn6432 was located on an integrative conjugative element (ICE)-like element, which could mediate the dissemination of the tet(31)-carrying transposon Tn6432 between bacteria. Comparative analysis demonstrated that Tn6432 homologs with the structure ISCR2-?phzF-tetR(31)-tet(31)-?glmM-sul2 were also carried by A. salmonicida, G. anatis, and O. alkaliphila, suggesting that this transposon can be transferred between species and even genera. This work provides the first report on the identification of the tet(31) gene in A. caviae, and will be helpful in exploring the dissemination mechanisms of tet(31) in water environment.Copyright © 2018. Published by Elsevier B.V.


April 21, 2020  |  

Confident phylogenetic identification of uncultured prokaryotes through long read amplicon sequencing of the 16S-ITS-23S rRNA operon.

Amplicon sequencing of the 16S rRNA gene is the predominant method to quantify microbial compositions and to discover novel lineages. However, traditional short amplicons often do not contain enough information to confidently resolve their phylogeny. Here we present a cost-effective protocol that amplifies a large part of the rRNA operon and sequences the amplicons with PacBio technology. We tested our method on a mock community and developed a read-curation pipeline that reduces the overall read error rate to 0.18%. Applying our method on four environmental samples, we captured near full-length rRNA operon amplicons from a large diversity of prokaryotes. The method operated at moderately high-throughput (22286-37,850 raw ccs reads) and generated a large amount of putative novel archaeal 23S rRNA gene sequences compared to the archaeal SILVA database. These long amplicons allowed for higher resolution during taxonomic classification by means of long (~1000 bp) 16S rRNA gene fragments and for substantially more confident phylogenies by means of combined near full-length 16S and 23S rRNA gene sequences, compared to shorter traditional amplicons (250 bp of the 16S rRNA gene). We recommend our method to those who wish to cost-effectively and confidently estimate the phylogenetic diversity of prokaryotes in environmental samples at high throughput. © 2019 The Authors. Environmental Microbiology published by Society for Applied Microbiology and John Wiley & Sons Ltd.


April 21, 2020  |  

Genome Comparisons of Wild Isolates of Caulobacter crescentus Reveal Rates of Inversion and Horizontal Gene Transfer.

Since previous interspecies comparisons of Caulobacter genomes have revealed extensive genome rearrangements, we decided to compare the nucleotide sequences of four C. crescentus genomes, NA1000, CB1, CB2, and CB13. To accomplish this goal, we used PacBio sequencing technology to determine the nucleotide sequence of the CB1, CB2, and CB13 genomes, and obtained each genome sequence as a single contig. To correct for possible sequencing errors, each genome was sequenced twice. The only differences we observed between the two sets of independently determined sequences were random omissions of a single base in a small percentage of the homopolymer regions where a single base is repeated multiple times. Comparisons of these four genomes indicated that horizontal gene transfer events that included small numbers of genes occurred at frequencies in the range of 10-3 to 10-4 insertions per generation. Large insertions were about 100 times less frequent. Also, in contrast to previous interspecies comparisons, we found no genome rearrangements when the closely related NA1000, CB1, and CB2 genomes were compared, and only eight inversions and one translocation when the more distantly related CB13 genome was compared to the other genomes. Thus, we estimate that inversions occur at a rate of one per 10 to 12 million generations in Caulobacter genomes. The inversions seem to be complex events that include the simultaneous creation of indels.


April 21, 2020  |  

Molecular Epidemiology of Candida auris in Colombia Reveals a Highly Related, Countrywide Colonization With Regional Patterns in Amphotericin B Resistance.

Candida auris is a multidrug-resistant yeast associated with hospital outbreaks worldwide. During 2015-2016, multiple outbreaks were reported in Colombia. We aimed to understand the extent of contamination in healthcare settings and to characterize the molecular epidemiology of C. auris in Colombia.We sampled patients, patient contacts, healthcare workers, and the environment in 4 hospitals with recent C. auris outbreaks. Using standardized protocols, people were swabbed at different body sites. Patient and procedure rooms were sectioned into 4 zones and surfaces were swabbed. We performed whole-genome sequencing (WGS) and antifungal susceptibility testing (AFST) on all isolates.Seven of the 17 (41%) people swabbed were found to be colonized. Candida auris was isolated from 37 of 322 (11%) environmental samples. These were collected from a variety of items in all 4 zones. WGS and AFST revealed that although isolates were similar throughout the country, isolates from the northern region were genetically distinct and more resistant to amphotericin B (AmB) than the isolates from central Colombia. Four novel nonsynonymous mutations were found to be significantly associated with AmB resistance.Our results show that extensive C. auris contamination can occur and highlight the importance of adherence to appropriate infection control practices and disinfection strategies. Observed genetic diversity supports healthcare transmission and a recent expansion of C. auris within Colombia with divergent AmB susceptibility.


April 21, 2020  |  

A reference genome for pea provides insight into legume genome evolution.

We report the first annotated chromosome-level reference genome assembly for pea, Gregor Mendel’s original genetic model. Phylogenetics and paleogenomics show genomic rearrangements across legumes and suggest a major role for repetitive elements in pea genome evolution. Compared to other sequenced Leguminosae genomes, the pea genome shows intense gene dynamics, most likely associated with genome size expansion when the Fabeae diverged from its sister tribes. During Pisum evolution, translocation and transposition differentially occurred across lineages. This reference sequence will accelerate our understanding of the molecular basis of agronomically important traits and support crop improvement.


April 21, 2020  |  

TSD: A Computational Tool To Study the Complex Structural Variants Using PacBio Targeted Sequencing Data.

PacBio sequencing is a powerful approach to study DNA or RNA sequences in a longer scope. It is especially useful in exploring the complex structural variants generated by random integration or multiple rearrangement of endogenous or exogenous sequences. Here, we present a tool, TSD, for complex structural variant discovery using PacBio targeted sequencing data. It allows researchers to identify and visualize the genomic structures of targeted sequences by unlimited splitting, alignment and assembly of long PacBio reads. Application to the sequencing data derived from an HBV integrated human cell line(PLC/PRF/5) indicated that TSD could recover the full profile of HBV integration events, especially for the regions with the complex human-HBV genome integrations and multiple HBV rearrangements. Compared to other long read analysis tools, TSD showed a better performance for detecting complex genomic structural variants. TSD is publicly available at: https://github.com/menggf/tsd. Copyright © 2019 Meng et al.


April 21, 2020  |  

Epidemiologic and genomic insights on mcr-1-harbouring Salmonella from diarrhoeal outpatients in Shanghai, China, 2006-2016.

Colistin resistance mediated by mcr-1-harbouring plasmids is an emerging threat in Enterobacteriaceae, like Salmonella. Based on its major contribution to the diarrhoea burden, the epidemic state and threat of mcr-1-harbouring Salmonella in community-acquired infections should be estimated.This retrospective study analysed the mcr-1 gene incidence in Salmonella strains collected from a surveillance on diarrhoeal outpatients in Shanghai Municipality, China, 2006-2016. Molecular characteristics of the mcr-1-positive strains and their plasmids were determined by genome sequencing. The transfer abilities of these plasmids were measured with various conjugation strains, species, and serotypes.Among the 12,053 Salmonella isolates, 37 mcr-1-harbouring strains, in which 35 were serovar Typhimurium, were detected first in 2012 and with increasing frequency after 2015. Most patients infected with mcr-1-harbouring strains were aged <5?years. All strains, including fluoroquinolone-resistant and/or extended-spectrum ß-lactamase-producing strains, were multi-drug resistant. S. Typhimurium had higher mcr-1 plasmid acquisition ability compared with other common serovars. Phylogeny based on the genomes combined with complete plasmid sequences revealed some clusters, suggesting the presence of mcr-1-harbouring Salmonella outbreaks in the community. Most mcr-1-positive strains were clustered together with the pork strains, strongly suggesting pork consumption as a main infection source.The mcr-1-harbouring Salmonella prevalence in community-acquired diarrhoea displays a rapid increase trend, and the ESBL-mcr-1-harbouring Salmonella poses a threat for children. These findings highlight the necessary and significance of prohibiting colistin use in animals and continuous monitoring of mcr-1-harbouring Salmonella.Copyright © 2019. Published by Elsevier B.V.


April 21, 2020  |  

Megabase Length Hypermutation Accompanies Human Structural Variation at 17p11.2.

DNA rearrangements resulting in human genome structural variants (SVs) are caused by diverse mutational mechanisms. We used long- and short-read sequencing technologies to investigate end products of de novo chromosome 17p11.2 rearrangements and query the molecular mechanisms underlying both recurrent and non-recurrent events. Evidence for an increased rate of clustered single-nucleotide variant (SNV) mutation in cis with non-recurrent rearrangements was found. Indel and SNV formation are associated with both copy-number gains and losses of 17p11.2, occur up to ~1 Mb away from the breakpoint junctions, and favor C > G transversion substitutions; results suggest that single-stranded DNA is formed during the genesis of the SV and provide compelling support for a microhomology-mediated break-induced replication (MMBIR) mechanism for SV formation. Our data show an additional mutational burden of MMBIR consisting of hypermutation confined to the locus and manifesting as SNVs and indels predominantly within genes. Copyright © 2019 Elsevier Inc. All rights reserved.


April 21, 2020  |  

Copy-number variants in clinical genome sequencing: deployment and interpretation for rare and undiagnosed disease.

Current diagnostic testing for genetic disorders involves serial use of specialized assays spanning multiple technologies. In principle, genome sequencing (GS) can detect all genomic pathogenic variant types on a single platform. Here we evaluate copy-number variant (CNV) calling as part of a clinically accredited GS test.We performed analytical validation of CNV calling on 17 reference samples, compared the sensitivity of GS-based variants with those from a clinical microarray, and set a bound on precision using orthogonal technologies. We developed a protocol for family-based analysis of GS-based CNV calls, and deployed this across a clinical cohort of 79 rare and undiagnosed cases.We found that CNV calls from GS are at least as sensitive as those from microarrays, while only creating a modest increase in the number of variants interpreted (~10 CNVs per case). We identified clinically significant CNVs in 15% of the first 79 cases analyzed, all of which were confirmed by an orthogonal approach. The pipeline also enabled discovery of a uniparental disomy (UPD) and a 50% mosaic trisomy 14. Directed analysis of select CNVs enabled breakpoint level resolution of genomic rearrangements and phasing of de novo CNVs.Robust identification of CNVs by GS is possible within a clinical testing environment.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.