Menu
July 19, 2019  |  

Long-read Single-Molecule Real-Time (SMRT) full gene sequencing of cytochrome P450-2D6 (CYP2D6).

The CYP2D6 enzyme metabolizes ~25% of common medications, yet homologous pseudogenes and copy-number variants (CNVs) make interrogating the polymorphic CYP2D6 gene with short-read sequencing challenging. Therefore, we developed a novel long-read, full gene CYP2D6 single-molecule real-time (SMRT) sequencing method using the Pacific Biosciences platform. Long-range PCR and CYP2D6 SMRT sequencing of 10 previously genotyped controls identified expected star (*) alleles, but also enabled suballele resolution, diplotype refinement, and discovery of novel alleles. Coupled with an optimized variant calling pipeline, CYP2D6 SMRT sequencing was highly reproducible as triplicate intra- and inter-run non-reference genotype results were completely concordant. Importantly, targeted SMRT sequencing of upstream and downstream CYP2D6 gene copies characterized the duplicated allele in 15 control samples with CYP2D6 CNVs. The utility of CYP2D6 SMRT sequencing was further underscored by identifying the diplotypes of 14 samples with discordant or unclear CYP2D6 configurations from previous targeted genotyping, which again included suballele resolution, duplicated allele characterization, and discovery of a novel allele and tandem arrangement (CYP2D6*36+*41). Taken together, long-read CYP2D6 SMRT sequencing is an innovative, reproducible, and validated method for full-gene characterization, duplication allele-specific analysis and novel allele discovery, which will likely improve CYP2D6 metabolizer phenotype prediction for both research and clinical testing applications. This article is protected by copyright. All rights reserved.This article is protected by copyright. All rights reserved.


July 19, 2019  |  

DNA methylation on N(6)-adenine in mammalian embryonic stem cells.

It has been widely accepted that 5-methylcytosine is the only form of DNA methylation in mammalian genomes. Here we identify N(6)-methyladenine as another form of DNA modification in mouse embryonic stem cells. Alkbh1 encodes a demethylase for N(6)-methyladenine. An increase of N(6)-methyladenine levels in Alkbh1-deficient cells leads to transcriptional silencing. N(6)-methyladenine deposition is inversely correlated with the evolutionary age of LINE-1 transposons; its deposition is strongly enriched at young (<1.5 million years old) but not old (>6 million years old) L1 elements. The deposition of N(6)-methyladenine correlates with epigenetic silencing of such LINE-1 transposons, together with their neighbouring enhancers and genes, thereby resisting the gene activation signals during embryonic stem cell differentiation. As young full-length LINE-1 transposons are strongly enriched on the X chromosome, genes located on the X chromosome are also silenced. Thus, N(6)-methyladenine developed a new role in epigenetic silencing in mammalian evolution distinct from its role in gene activation in other organisms. Our results demonstrate that N(6)-methyladenine constitutes a crucial component of the epigenetic regulation repertoire in mammalian genomes.


July 19, 2019  |  

Nested Russian doll-like genetic mobility drives rapid dissemination of the Carbapenem resistance gene blaKPC

The recent widespread emergence of carbapenem resistance in Enterobacteriaceae is a major public health concern, as carbapenems are a therapy of last resort against this family of common bacterial pathogens. Resistance genes can mobilize via various mechanisms, including conjugation and transposition; however, the importance of this mobility in short-term evolution, such as within nosocomial outbreaks, is unknown. Using a combination of short- and long-read whole-genome sequencing of 281 blaKPC-positive Enterobacteriaceae isolates from a single hospital over 5 years, we demonstrate rapid dissemination of this carbapenem resistance gene to multiple species, strains, and plasmids. Mobility of blaKPC occurs at multiple nested genetic levels, with transmission of blaKPC strains between individuals, frequent transfer of blaKPC plasmids between strains/species, and frequent transposition of blaKPC transposon Tn4401 between plasmids. We also identify a common insertion site for Tn4401 within various Tn2-like elements, suggesting that homologous recombination between Tn2-like elements has enhanced the spread of Tn4401 between different plasmid vectors. Furthermore, while short-read sequencing has known limitations for plasmid assembly, various studies have attempted to overcome this by the use of reference-based methods. We also demonstrate that, as a consequence of the genetic mobility observed in this study, plasmid structures can be extremely dynamic, and therefore these reference-based methods, as well as traditional partial typing methods, can produce very misleading conclusions. Overall, our findings demonstrate that nonclonal resistance gene dissemination can be extremely rapid, presenting significant challenges for public health surveillance and achieving effective control of antibiotic resistance. Copyright © 2016 Sheppard et al.


July 19, 2019  |  

Development and clinical application of an integrative genomic approach to personalized cancer therapy.

Personalized therapy provides the best outcome of cancer care and its implementation in the clinic has been greatly facilitated by recent convergence of enormous progress in basic cancer research, rapid advancement of new tumor profiling technologies, and an expanding compendium of targeted cancer therapeutics.We developed a personalized cancer therapy (PCT) program in a clinical setting, using an integrative genomics approach to fully characterize the complexity of each tumor. We carried out whole exome sequencing (WES) and single-nucleotide polymorphism (SNP) microarray genotyping on DNA from tumor and patient-matched normal specimens, as well as RNA sequencing (RNA-Seq) on available frozen specimens, to identify somatic (tumor-specific) mutations, copy number alterations (CNAs), gene expression changes, gene fusions, and also germline variants. To provide high sensitivity in known cancer mutation hotspots, Ion AmpliSeq Cancer Hotspot Panel v2 (CHPv2) was also employed. We integrated the resulting data with cancer knowledge bases and developed a specific workflow for each cancer type to improve interpretation of genomic data.We returned genomics findings to 46 patients and their physicians describing somatic alterations and predicting drug response, toxicity, and prognosis. Mean 17.3 cancer-relevant somatic mutations per patient were identified, 13.3-fold, 6.9-fold, and 4.7-fold more than could have been detected using CHPv2, Oncomine Cancer Panel (OCP), and FoundationOne, respectively. Our approach delineated the underlying genetic drivers at the pathway level and provided meaningful predictions of therapeutic efficacy and toxicity. Actionable alterations were found in 91 % of patients (mean 4.9 per patient, including somatic mutations, copy number alterations, gene expression alterations, and germline variants), a 7.5-fold, 2.0-fold, and 1.9-fold increase over what could have been uncovered by CHPv2, OCP, and FoundationOne, respectively. The findings altered the course of treatment in four cases. These results show that a comprehensive, integrative genomic approach as outlined above significantly enhanced genomics-based PCT strategies.


July 19, 2019  |  

Genetic stability of genome-scale deoptimized RNA virus vaccine candidates under selective pressure.

Recoding viral genomes by numerous synonymous but suboptimal substitutions provides live attenuated vaccine candidates. These vaccine candidates should have a low risk of deattenuation because of the many changes involved. However, their genetic stability under selective pressure is largely unknown. We evaluated phenotypic reversion of deoptimized human respiratory syncytial virus (RSV) vaccine candidates in the context of strong selective pressure. Codon pair deoptimized (CPD) versions of RSV were attenuated and temperature-sensitive. During serial passage at progressively increasing temperature, a CPD RSV containing 2,692 synonymous mutations in 9 of 11 ORFs did not lose temperature sensitivity, remained genetically stable, and was restricted at temperatures of 34 °C/35 °C and above. However, a CPD RSV containing 1,378 synonymous mutations solely in the polymerase L ORF quickly lost substantial attenuation. Comprehensive sequence analysis of virus populations identified many different potentially deattenuating mutations in the L ORF as well as, surprisingly, many appearing in other ORFs. Phenotypic analysis revealed that either of two competing mutations in the virus transcription antitermination factor M2-1, outside of the CPD area, substantially reversed defective transcription of the CPD L gene and substantially restored virus fitness in vitro and in case of one of these two mutations, also in vivo. Paradoxically, the introduction into Min L of one mutation each in the M2-1, N, P, and L proteins resulted in a virus with increased attenuation in vivo but increased immunogenicity. Thus, in addition to providing insights on the adaptability of genome-scale deoptimized RNA viruses, stability studies can yield improved synthetic RNA virus vaccine candidates.


July 19, 2019  |  

Chromosomal integration of the Klebsiella pneumoniae carbapenemase gene, blaKPC, in Klebsiella species is elusive but not rare.

Carbapenemase genes in Enterobacteriaceae are mostly described as being plasmid associated. However, the genetic context of carbapenemase genes is not always confirmed in epidemiological surveys, and the frequency of their chromosomal integration therefore is unknown. A previously sequenced collection of blaKPC-positive Enterobacteriaceae from a single U.S. institution (2007 to 2012; n = 281 isolates from 182 patients) was analyzed to identify chromosomal insertions of Tn4401, the transposon most frequently harboring blaKPC Using a combination of short- and long-read sequencing, we confirmed five independent chromosomal integration events from 6/182 (3%) patients, corresponding to 15/281 (5%) isolates. Three patients had isolates identified by perirectal screening, and three had infections which were all successfully treated. When a single copy of blaKPC was in the chromosome, one or both of the phenotypic carbapenemase tests were negative. All chromosomally integrated blaKPC genes were from Klebsiella spp., predominantly K. pneumoniae clonal group 258 (CG258), even though these represented only a small proportion of the isolates. Integration occurred via IS15-?I-mediated transposition of a larger, composite region encompassing Tn4401 at one locus of chromosomal integration, seen in the same strain (K. pneumoniae ST340) in two patients. In summary, we identified five independent chromosomal integrations of blaKPC in a large outbreak, demonstrating that this is not a rare event. blaKPC was more frequently integrated into the chromosome of epidemic CG258 K. pneumoniae lineages (ST11, ST258, and ST340) and was more difficult to detect by routine phenotypic methods in this context. The presence of chromosomally integrated blaKPC within successful, globally disseminated K. pneumoniae strains therefore is likely underestimated. Copyright © 2017 Mathers et al.


July 19, 2019  |  

Genomic confirmation of vancomycin-resistant Enterococcus transmission from deceased donor to liver transplant recipient.

In a liver transplant recipient with vancomycin-resistant Enterococcus (VRE) surgical site and bloodstream infection, a combination of pulsed-field gel electrophoresis, multilocus sequence typing, and whole genome sequencing identified that donor and recipient VRE isolates were highly similar when compared to time-matched hospital isolates. Comparison of de novo assembled isolate genomes was highly suggestive of transplant transmission rather than hospital-acquired transmission and also identified subtle internal rearrangements between donor and recipient missed by other genomic approaches. Given the improved resolution, whole-genome assembly of pathogen genomes is likely to become an essential tool for investigation of potential organ transplant transmissions.


July 19, 2019  |  

TAL effector driven induction of a SWEET gene confers susceptibility to bacterial blight of cotton.

Transcription activator-like (TAL) effectors from Xanthomonas citri subsp. malvacearum (Xcm) are essential for bacterial blight of cotton (BBC). Here, by combining transcriptome profiling with TAL effector-binding element (EBE) prediction, we show that GhSWEET10, encoding a functional sucrose transporter, is induced by Avrb6, a TAL effector determining Xcm pathogenicity. Activation of GhSWEET10 by designer TAL effectors (dTALEs) restores virulence of Xcm avrb6 deletion strains, whereas silencing of GhSWEET10 compromises cotton susceptibility to infections. A BBC-resistant line carrying an unknown recessive b6 gene bears the same EBE as the susceptible line, but Avrb6-mediated induction of GhSWEET10 is reduced, suggesting a unique mechanism underlying b6-mediated resistance. We show via an extensive survey of GhSWEET transcriptional responsiveness to different Xcm field isolates that additional GhSWEETs may also be involved in BBC. These findings advance our understanding of the disease and resistance in cotton and may facilitate the development cotton with improved resistance to BBC.


July 19, 2019  |  

Amplification-free, CRISPR-Cas9 targeted enrichment and SMRT Sequencing of repeat-expansion disease causative genomic regions

Targeted sequencing has proven to be an economical means of obtaining sequence information for one or more defined regions of a larger genome. However, most target enrichment methods require amplification. Some genomic regions, such as those with extreme GC content and repetitive sequences, are recalcitrant to faithful amplification. Yet, many human genetic disorders are caused by repeat expansions, including difficult to sequence tandem repeats. We have developed a novel, amplification-free enrichment technique that employs the CRISPR-Cas9 system for specific targeting multiple genomic loci. This method, in conjunction with long reads generated through Single Molecule, Real-Time (SMRT) sequencing and unbiased coverage, enables enrichment and sequencing of complex genomic regions that cannot be investigated with other technologies. Using human genomic DNA samples, we demonstrate successful targeting of causative loci for Huntingtontextquoterights disease (HTT; CAG repeat), Fragile X syndrome (FMR1; CGG repeat), amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (C9orf72; GGGGCC repeat), and spinocerebellar ataxia type 10 (SCA10) (ATXN10; variable ATTCT repeat). The method, amenable to multiplexing across multiple genomic loci, uses an amplification-free approach that facilitates the isolation of hundreds of individual on-target molecules in a single SMRT Cell and accurate sequencing through long repeat stretches, regardless of extreme GC percent or sequence complexity content. Our novel targeted sequencing method opens new doors to genomic analyses independent of PCR amplification that will facilitate the study of repeat expansion disorders.


July 19, 2019  |  

Coupling of single molecule, long read sequencing with IMGT/HighV-QUEST analysis expedites identification of SIV gp140-specific antibodies from scFv phage display libraries.

The simian immunodeficiency virus (SIV)/macaque model of human immunodeficiency virus (HIV)/acquired immunodeficiency syndrome pathogenesis is critical for furthering our understanding of the role of antibody responses in the prevention of HIV infection, and will only increase in importance as macaque immunoglobulin (IG) gene databases are expanded. We have previously reported the construction of a phage display library from a SIV-infected rhesus macaque (Macaca mulatta) using oligonucleotide primers based on human IG gene sequences. Our previous screening relied on Sanger sequencing, which was inefficient and generated only a few dozen sequences. Here, we re-analyzed this library using single molecule, real-time (SMRT) sequencing on the Pacific Biosciences (PacBio) platform to generate thousands of highly accurate circular consensus sequencing (CCS) reads corresponding to full length single chain fragment variable. CCS data were then analyzed through the international ImMunoGeneTics information system®(IMGT®)/HighV-QUEST (www.imgt.org) to identify variable genes and perform statistical analyses. Overall the library was very diverse, with 2,569 different IMGT clonotypes called for the 5,238 IGHV sequences assigned to an IMGT clonotype. Within the library, SIV-specific antibodies represented a relatively limited number of clones, with only 135 different IMGT clonotypes called from 4,594 IGHV-assigned sequences. Our data did confirm that the IGHV4 and IGHV3 gene usage was the most abundant within the rhesus antibodies screened, and that these genes were even more enriched among SIV gp140-specific antibodies. Although a broad range of VH CDR3 amino acid (AA) lengths was observed in the unpanned library, the vast majority of SIV gp140-specific antibodies demonstrated a more uniform VH CDR3 length (20 AA). This uniformity was far less apparent when VH CDR3 were classified according to their clonotype (range: 9-25 AA), which we believe is more relevant for specific antibody identification. Only 174 IGKV and 588 IGLV clonotypes were identified within the VL sequences associated with SIV gp140-specific VH. Together, these data strongly suggest that the combination of SMRT sequencing with the IMGT/HighV-QUEST querying tool will facilitate and expedite our understanding of polyclonal antibody responses during SIV infection and may serve to rapidly expand the known scope of macaque V genes utilized during these responses.


July 19, 2019  |  

Full-Length Envelope Analyzer (FLEA): A tool for longitudinal analysis of viral amplicons.

Next generation sequencing of viral populations has advanced our understanding of viral population dynamics, the development of drug resistance, and escape from host immune responses. Many applications require complete gene sequences, which can be impossible to reconstruct from short reads. HIV env, the protein of interest for HIV vaccine studies, is exceptionally challenging for long-read sequencing and analysis due to its length, high substitution rate, and extensive indel variation. While long-read sequencing is attractive in this setting, the analysis of such data is not well handled by existing methods. To address this, we introduce FLEA (Full-Length Envelope Analyzer), which performs end-to-end analysis and visualization of long-read sequencing data. FLEA consists of both a pipeline (optionally run on a high-performance cluster), and a client-side web application that provides interactive results. The pipeline transforms FASTQ reads into high-quality consensus sequences (HQCSs) and uses them to build a codon-aware multiple sequence alignment. The resulting alignment is then used to infer phylogenies, selection pressure, and evolutionary dynamics. The web application provides publication-quality plots and interactive visualizations, including an annotated viral alignment browser, time series plots of evolutionary dynamics, visualizations of gene-wide selective pressures (such as dN/dS) across time and across protein structure, and a phylogenetic tree browser. We demonstrate how FLEA may be used to process Pacific Biosciences HIV env data and describe recent examples of its use. Simulations show how FLEA dramatically reduces the error rate of this sequencing platform, providing an accurate portrait of complex and variable HIV env populations. A public instance of FLEA is hosted at http://flea.datamonkey.org. The Python source code for the FLEA pipeline can be found at https://github.com/veg/flea-pipeline. The client-side application is available at https://github.com/veg/flea-web-app. A live demo of the P018 results can be found at http://flea.murrell.group/view/P018.


July 19, 2019  |  

HIV envelope glycoform heterogeneity and localized diversity govern the initiation and maturation of a V2 apex broadly neutralizing antibody lineage.

Understanding how broadly neutralizing antibodies (bnAbs) to HIV envelope (Env) develop during natural infection can help guide the rational design of an HIV vaccine. Here, we described a bnAb lineage targeting the Env V2 apex and the Ab-Env co-evolution that led to development of neutralization breadth. The lineage Abs bore an anionic heavy chain complementarity-determining region 3 (CDRH3) of 25 amino acids, among the shortest known for this class of Abs, and achieved breadth with only 10% nucleotide somatic hypermutation and no insertions or deletions. The data suggested a role for Env glycoform heterogeneity in the activation of the lineage germline B cell. Finally, we showed that localized diversity at key V2 epitope residues drove bnAb maturation toward breadth, mirroring the Env evolution pattern described for another donor who developed V2-apex targeting bnAbs. Overall, these findings suggest potential strategies for vaccine approaches based on germline-targeting and serial immunogen design. Copyright © 2017 The Authors. Published by Elsevier Inc. All rights reserved.


July 7, 2019  |  

Klebsiella pneumoniae carbapenemase (KPC)-producing K. pneumoniae at a single institution: insights into endemicity from whole-genome sequencing.

The global emergence of Klebsiella pneumoniae carbapenemase-producing K. pneumoniae (KPC-Kp) multilocus sequence type ST258 is widely recognized. Less is known about the molecular and epidemiological details of non-ST258 K. pneumoniae in the setting of an outbreak mediated by an endemic plasmid. We describe the interplay of blaKPC plasmids and K. pneumoniae strains and their relationship to the location of acquisition in a U.S. health care institution. Whole-genome sequencing (WGS) analysis was applied to KPC-Kp clinical isolates collected from a single institution over 5 years following the introduction of blaKPC in August 2007, as well as two plasmid transformants. KPC-Kp from 37 patients yielded 16 distinct sequence types (STs). Two novel conjugative blaKPC plasmids (pKPC_UVA01 and pKPC_UVA02), carried by the hospital index case, accounted for the presence of blaKPC in 21/37 (57%) subsequent cases. Thirteen (35%) isolates represented an emergent lineage, ST941, which contained pKPC_UVA01 in 5/13 (38%) and pKPC_UVA02 in 6/13 (46%) cases. Seven (19%) isolates were the epidemic KPC-Kp strain, ST258, mostly imported from elsewhere and not carrying pKPC_UVA01 or pKPC_UVA02. Using WGS-based analysis of clinical isolates and plasmid transformants, we demonstrate the unexpected dispersal of blaKPC to many non-ST258 lineages in a hospital through spread of at least two novel blaKPC plasmids. In contrast, ST258 KPC-Kp was imported into the institution on numerous occasions, with other blaKPC plasmid vectors and without sustained transmission. Instead, a newly recognized KPC-Kp strain, ST941, became associated with both novel blaKPC plasmids and spread locally, making it a future candidate for clinical persistence and dissemination. Copyright © 2015, Mathers et al.


July 7, 2019  |  

Whole-genome sequencing identifies emergence of a quinolone resistance mutation in a case of Stenotrophomonas maltophilia bacteremia.

Whole-genome sequences for Stenotrophomonas maltophilia serial isolates from a bacteremic patient before and after development of levofloxacin resistance were assembled de novo and differed by one single-nucleotide variant in smeT, a repressor for multidrug efflux operon smeDEF. Along with sequenced isolates from five contemporaneous cases, they displayed considerable diversity compared against all published complete genomes. Whole-genome sequencing and complete assembly can conclusively identify resistance mechanisms emerging in S. maltophilia strains during clinical therapy. Copyright © 2015, American Society for Microbiology. All Rights Reserved.


July 7, 2019  |  

The complete genome sequence of the emerging pathogen Mycobacterium haemophilum explains its unique culture requirements.

Mycobacterium haemophilum is an emerging pathogen associated with a variety of clinical syndromes, most commonly skin infections in immunocompromised individuals. M. haemophilum exhibits a unique requirement for iron supplementation to support its growth in culture, but the basis for this property and how it may shape pathogenesis is unclear. Using a combination of Illumina, PacBio, and Sanger sequencing, the complete genome sequence of M. haemophilum was determined. Guided by this sequence, experiments were performed to define the basis for the unique growth requirements of M. haemophilum. We found that M. haemophilum, unlike many other mycobacteria, is unable to synthesize iron-binding siderophores known as mycobactins or to utilize ferri-mycobactins to support growth. These differences correlate with the absence of genes associated with mycobactin synthesis, secretion, and uptake. In agreement with the ability of heme to promote growth, we identified genes encoding heme uptake machinery. Consistent with its propensity to infect the skin, we show at the whole-genome level the genetic closeness of M. haemophilum with Mycobacterium leprae, an organism which cannot be cultivated in vitro, and we identify genes uniquely shared by these organisms. Finally, we identify means to express foreign genes in M. haemophilum. These data explain the unique culture requirements for this important pathogen, provide a foundation upon which the genome sequence can be exploited to improve diagnostics and therapeutics, and suggest use of M. haemophilum as a tool to elucidate functions of genes shared with M. leprae.Mycobacterium haemophilum is an emerging pathogen with an unknown natural reservoir that exhibits unique requirements for iron supplementation to grow in vitro. Understanding the basis for this iron requirement is important because it is fundamental to isolation of the organism from clinical samples and environmental sources. Defining the molecular basis for M. haemophilium’s growth requirements will also shed new light on mycobacterial strategies to acquire iron and can be exploited to define how differences in such strategies influence pathogenesis. Here, through a combination of sequencing and experimental approaches, we explain the basis for the iron requirement. We further demonstrate the genetic closeness of M. haemophilum and Mycobacterium leprae, the causative agent of leprosy which cannot be cultured in vitro, and we demonstrate methods to genetically manipulate M. haemophilum. These findings pave the way for the use of M. haemophilum as a model to elucidate functions of genes shared with M. leprae. Copyright © 2015 Tufariello et al.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.