Menu
July 19, 2019

Genome Update. Let the consumer beware: Streptomyces genome sequence quality.

A genome sequence assembly represents a model of a genome. This article explores some tools and methods for assessing the quality of an assembly, using publicly available data for Streptomyces species as the example. There is great variability in quality of assemblies deposited in GenBank. Only in a small minority of these assemblies are the raw data available, enabling full appraisal of the assembly quality. © 2015 The Authors. Microbial Biotechnology published by John Wiley & Sons Ltd and Society for Applied Microbiology.


July 19, 2019

A method for near full-length amplification and sequencing for six hepatitis C virus genotypes.

Hepatitis C virus (HCV) is a rapidly evolving RNA virus that has been classified into seven genotypes. All HCV genotypes cause chronic hepatitis, which ultimately leads to liver diseases such as cirrhosis. The genotypes are unevenly distributed across the globe, with genotypes 1 and 3 being the most prevalent. Until recently, molecular epidemiological studies of HCV evolution within the host and at the population level have been limited to the analyses of partial viral genome segments, as it has been technically challenging to amplify and sequence the full-length of the 9.6 kb HCV genome. Although recent improvements have been made in full genome sequencing methodologies, these protocols are still either limited to a specific genotype or cost-inefficient.In this study we describe a genotype-specific protocol for the amplification and sequencing of the near-full length genome of all six major HCV genotypes. We applied this protocol to 122 HCV positive clinical samples, and had a successful genome amplification rate of 90 %, when the viral load was greater than 15,000 IU/ml. The assay was shown to have a detection limit of 1-3 cDNA copies per reaction. The method was tested with both Illumina and PacBio single molecule, real-time (SMRT) sequencing technologies. Illumina sequencing resulted in deep coverage and allowed detection of rare variants as well as HCV co-infection with multiple genotypes. The application of the method with PacBio RS resulted in sequence reads greater than 9 kb that covered the near full-length HCV amplicon in a single read and enabled analysis of the near full-length quasispecies.The protocol described herein can be utilised for rapid amplification and sequencing of the near-full length HCV genome in a cost efficient manner suitable for a wide range of applications.


July 19, 2019

DNA methylation on N(6)-adenine in mammalian embryonic stem cells.

It has been widely accepted that 5-methylcytosine is the only form of DNA methylation in mammalian genomes. Here we identify N(6)-methyladenine as another form of DNA modification in mouse embryonic stem cells. Alkbh1 encodes a demethylase for N(6)-methyladenine. An increase of N(6)-methyladenine levels in Alkbh1-deficient cells leads to transcriptional silencing. N(6)-methyladenine deposition is inversely correlated with the evolutionary age of LINE-1 transposons; its deposition is strongly enriched at young (<1.5 million years old) but not old (>6 million years old) L1 elements. The deposition of N(6)-methyladenine correlates with epigenetic silencing of such LINE-1 transposons, together with their neighbouring enhancers and genes, thereby resisting the gene activation signals during embryonic stem cell differentiation. As young full-length LINE-1 transposons are strongly enriched on the X chromosome, genes located on the X chromosome are also silenced. Thus, N(6)-methyladenine developed a new role in epigenetic silencing in mammalian evolution distinct from its role in gene activation in other organisms. Our results demonstrate that N(6)-methyladenine constitutes a crucial component of the epigenetic regulation repertoire in mammalian genomes.


July 19, 2019

Radical remodeling of the Y chromosome in a recent radiation of malaria mosquitoes.

Y chromosomes control essential male functions in many species, including sex determination and fertility. However, because of obstacles posed by repeat-rich heterochromatin, knowledge of Y chromosome sequences is limited to a handful of model organisms, constraining our understanding of Y biology across the tree of life. Here, we leverage long single-molecule sequencing to determine the content and structure of the nonrecombining Y chromosome of the primary African malaria mosquito, Anopheles gambiae. We find that the An. gambiae Y consists almost entirely of a few massively amplified, tandemly arrayed repeats, some of which can recombine with similar repeats on the X chromosome. Sex-specific genome resequencing in a recent species radiation, the An. gambiae complex, revealed rapid sequence turnover within An. gambiae and among species. Exploiting 52 sex-specific An. gambiae RNA-Seq datasets representing all developmental stages, we identified a small repertoire of Y-linked genes that lack X gametologs and are not Y-linked in any other species except An. gambiae, with the notable exception of YG2, a candidate male-determining gene. YG2 is the only gene conserved and exclusive to the Y in all species examined, yet sequence similarity to YG2 is not detectable in the genome of a more distant mosquito relative, suggesting rapid evolution of Y chromosome genes in this highly dynamic genus of malaria vectors. The extensive characterization of the An. gambiae Y provides a long-awaited foundation for studying male mosquito biology, and will inform novel mosquito control strategies based on the manipulation of Y chromosomes.


July 19, 2019

Comprehensive mutagenesis of the fimS promoter regulatory switch reveals novel regulation of type 1 pili in uropathogenic Escherichia coli.

Type 1 pili (T1P) are major virulence factors for uropathogenic Escherichia coli (UPEC), which cause both acute and recurrent urinary tract infections. T1P expression therefore is of direct relevance for disease. T1P are phase variable (both piliated and nonpiliated bacteria exist in a clonal population) and are controlled by an invertible DNA switch (fimS), which contains the promoter for the fim operon encoding T1P. Inversion of fimS is stochastic but may be biased by environmental conditions and other signals that ultimately converge at fimS itself. Previous studies of fimS sequences important for T1P phase variation have focused on laboratory-adapted E. coli strains and have been limited in the number of mutations or by alteration of the fimS genomic context. We surmounted these limitations by using saturating genomic mutagenesis of fimS coupled with accurate sequencing to detect both mutations and phase status simultaneously. In addition to the sequences known to be important for biasing fimS inversion, our method also identifies a previously unknown pair of 5′ UTR inverted repeats that act by altering the relative fimA levels to control phase variation. Thus we have uncovered an additional layer of T1P regulation potentially impacting virulence and the coordinate expression of multiple pilus systems.


July 19, 2019

An incomplete understanding of human genetic variation.

Deciphering the genetic basis of human disease requires a comprehensive knowledge of genetic variants irrespective of their class or frequency. Although an impressive number of human genetic variants have been catalogued, a large fraction of the genetic difference that distinguishes two human genomes is still not understood at the base-pair level. This is because the emphasis has been on single-nucleotide variation as opposed to less tractable and more complex genetic variants, including indels and structural variants. The latter, we propose, will have a large impact on human phenotypes but require a more systematic assessment of genomes at deeper coverage and alternate sequencing and mapping technologies. Copyright © 2016 by the Genetics Society of America.


July 19, 2019

PacBio SMRT assembly of a complex multi-replicon genome reveals chlorocatechol degradative operon in a region of genome plasticity.

We have sequenced a Burkholderia genome that contains multiple replicons and large repetitive elements that would make it inherently difficult to assemble by short read sequencing technologies. We illustrate how the integrated long read correction algorithms implemented through the PacBio Single Molecule Real-Time (SMRT) sequencing technology successfully provided a de novo assembly that is a reasonable estimate of both the gene content and genome organization without making any further modifications. This assembly is comparable to related organisms assembled by more labour intensive methods. Our assembled genome revealed regions of genome plasticity for further investigation, one of which harbours a chlorocatechol degradative operon highly homologous to those previously identified on globally ubiquitous plasmids. In an ideal world, this assembly would still require experimental validation to confirm gene order and copy number of repeated elements. However, we submit that particularly in instances where a polished genome is not the primary goal of the sequencing project, PacBio SMRT sequencing provides a financially viable option for generating a biologically relevant genome estimate that can be utilized by other researchers for comparative studies. Copyright © 2016. Published by Elsevier B.V.


July 19, 2019

Polymerase specific error rates and profiles identified by single molecule sequencing.

DNA polymerases have an innate error rate which is polymerase and DNA context specific. Historically the mutational rate and profiles have been measured using a variety of methods, each with their own technical limitations. Here we used the unique properties of single molecule sequencing to evaluate the mutational rate and profiles of six DNA polymerases at the sequence level. In addition to accurately determining mutations in double strands, single molecule sequencing also captures direction specific transversions and transitions through the analysis of heteroduplexes. Not only did the error rates vary, but also the direction specific transitions differed among polymerases. Copyright © 2016 Elsevier B.V. All rights reserved.


July 19, 2019

Nested Russian doll-like genetic mobility drives rapid dissemination of the Carbapenem resistance gene blaKPC

The recent widespread emergence of carbapenem resistance in Enterobacteriaceae is a major public health concern, as carbapenems are a therapy of last resort against this family of common bacterial pathogens. Resistance genes can mobilize via various mechanisms, including conjugation and transposition; however, the importance of this mobility in short-term evolution, such as within nosocomial outbreaks, is unknown. Using a combination of short- and long-read whole-genome sequencing of 281 blaKPC-positive Enterobacteriaceae isolates from a single hospital over 5 years, we demonstrate rapid dissemination of this carbapenem resistance gene to multiple species, strains, and plasmids. Mobility of blaKPC occurs at multiple nested genetic levels, with transmission of blaKPC strains between individuals, frequent transfer of blaKPC plasmids between strains/species, and frequent transposition of blaKPC transposon Tn4401 between plasmids. We also identify a common insertion site for Tn4401 within various Tn2-like elements, suggesting that homologous recombination between Tn2-like elements has enhanced the spread of Tn4401 between different plasmid vectors. Furthermore, while short-read sequencing has known limitations for plasmid assembly, various studies have attempted to overcome this by the use of reference-based methods. We also demonstrate that, as a consequence of the genetic mobility observed in this study, plasmid structures can be extremely dynamic, and therefore these reference-based methods, as well as traditional partial typing methods, can produce very misleading conclusions. Overall, our findings demonstrate that nonclonal resistance gene dissemination can be extremely rapid, presenting significant challenges for public health surveillance and achieving effective control of antibiotic resistance. Copyright © 2016 Sheppard et al.


July 19, 2019

Next generation sequencing of Actinobacteria for the discovery of novel natural products.

Like many fields of the biosciences, actinomycete natural products research has been revolutionised by next-generation DNA sequencing (NGS). Hundreds of new genome sequences from actinobacteria are made public every year, many of them as a result of projects aimed at identifying new natural products and their biosynthetic pathways through genome mining. Advances in these technologies in the last five years have meant not only a reduction in the cost of whole genome sequencing, but also a substantial increase in the quality of the data, having moved from obtaining a draft genome sequence comprised of several hundred short contigs, sometimes of doubtful reliability, to the possibility of obtaining an almost complete and accurate chromosome sequence in a single contig, allowing a detailed study of gene clusters and the design of strategies for refactoring and full gene cluster synthesis. The impact that these technologies are having in the discovery and study of natural products from actinobacteria, including those from the marine environment, is only starting to be realised. In this review we provide a historical perspective of the field, analyse the strengths and limitations of the most relevant technologies, and share the insights acquired during our genome mining projects.


July 19, 2019

Initial assessment of the molecular epidemiology of blaNDM-1 in Colombia.

We report complete genome sequences of fourblaNDM-1-harboring Gram-negative multidrug resistant (MDR) isolates from Colombia. TheblaNDM-1genes were located 193Kb-Inc FIA, 178Kb-Inc A/C2 and 47Kb (unknown Inc type) plasmids. MLST revealed that isolates belong to ST10 (Escherichia coli), ST392 (Klebsiella pneumoniae), and ST322 and ST464 (Acinetobacter baumanniiandA. nosocomialis, respectively). Our analysis identified that the Inc A/C2 plasmid inE. colicontained a novel complex transposon (Tn125and Tn5393with 3 copies ofblaNDM-1) and a recombination “hotspot” for the acquisition of new resistance determinants. Copyright © 2016, American Society for Microbiology. All Rights Reserved.


July 19, 2019

Comparative analyses of low, medium and high-resolution HLA typing technologies for human populations

Human Leukocyte Antigen (HLA) encoding genes are part of the major histocompatibility complex (MHC) on human chromosome 6. This region is one of the most polymorphic regions in the human genome. Prior knowledge of HLA allelic polymorphisms is clinically important for matching donor and recipient during organ/tissue transplantation. HLA allelic information is also useful in predicting immune responses to various infectious diseases, genetic disorders and autoimmune conditions. India harbors over a billion people and its population is untapped for HLA allelic diversity. In this study, we explored and compared three HLA typing methods for South Indian population, using Sequence-Specific Primers (SSP), NGS (Roche/454) and single- molecule sequencing (PacBio RS II) platforms. Over 1020 DNA samples were typed at low resolution using SSP method to determine the major HLA alleles within the South Indian population. These studies were followed up with medium resolution HLA typing of 80 samples based on exonic sequences on the Roche/454 sequencing system and high-resolution (6-8 digit) typing of 8 samples for HLA alleles of class I genes (HLA-A, B and C) and class II genes (HLA-DRB1 and DQB1) using PacBio RS II platform. The long reads delivered by SMRT technology, covered the full-length class I and class II genes/alleles in contiguous reads including untranslated regions, exons and introns, which provided phased SNP information. We have identified three novel alleles from PacBio data that were verified by Roche 454 sequencing. This is the first case study of HLA typing using second and third generation NGS technologies for an Indian population. The PacBio platform is a promising platform for large-scale HLA typing for establishing an HLA database for the untapped ethnic populations of India.


July 19, 2019

Development and clinical application of an integrative genomic approach to personalized cancer therapy.

Personalized therapy provides the best outcome of cancer care and its implementation in the clinic has been greatly facilitated by recent convergence of enormous progress in basic cancer research, rapid advancement of new tumor profiling technologies, and an expanding compendium of targeted cancer therapeutics.We developed a personalized cancer therapy (PCT) program in a clinical setting, using an integrative genomics approach to fully characterize the complexity of each tumor. We carried out whole exome sequencing (WES) and single-nucleotide polymorphism (SNP) microarray genotyping on DNA from tumor and patient-matched normal specimens, as well as RNA sequencing (RNA-Seq) on available frozen specimens, to identify somatic (tumor-specific) mutations, copy number alterations (CNAs), gene expression changes, gene fusions, and also germline variants. To provide high sensitivity in known cancer mutation hotspots, Ion AmpliSeq Cancer Hotspot Panel v2 (CHPv2) was also employed. We integrated the resulting data with cancer knowledge bases and developed a specific workflow for each cancer type to improve interpretation of genomic data.We returned genomics findings to 46 patients and their physicians describing somatic alterations and predicting drug response, toxicity, and prognosis. Mean 17.3 cancer-relevant somatic mutations per patient were identified, 13.3-fold, 6.9-fold, and 4.7-fold more than could have been detected using CHPv2, Oncomine Cancer Panel (OCP), and FoundationOne, respectively. Our approach delineated the underlying genetic drivers at the pathway level and provided meaningful predictions of therapeutic efficacy and toxicity. Actionable alterations were found in 91 % of patients (mean 4.9 per patient, including somatic mutations, copy number alterations, gene expression alterations, and germline variants), a 7.5-fold, 2.0-fold, and 1.9-fold increase over what could have been uncovered by CHPv2, OCP, and FoundationOne, respectively. The findings altered the course of treatment in four cases. These results show that a comprehensive, integrative genomic approach as outlined above significantly enhanced genomics-based PCT strategies.


July 19, 2019

A comparison of tools for the simulation of genomic next-generation sequencing data.

Computer simulation of genomic data has become increasingly popular for assessing and validating biological models or for gaining an understanding of specific data sets. Several computational tools for the simulation of next-generation sequencing (NGS) data have been developed in recent years, which could be used to compare existing and new NGS analytical pipelines. Here we review 23 of these tools, highlighting their distinct functionality, requirements and potential applications. We also provide a decision tree for the informed selection of an appropriate NGS simulation tool for the specific question at hand.


July 19, 2019

Analysis of tandem gene copies in maize chromosomal regions reconstructed from long sequence reads.

Haplotype variation not only involves SNPs but also insertions and deletions, in particular gene copy number variations. However, comparisons of individual genomes have been difficult because traditional sequencing methods give too short reads to unambiguously reconstruct chromosomal regions containing repetitive DNA sequences. An example of such a case is the protein gene family in maize that acts as a sink for reduced nitrogen in the seed. Previously, 41-48 gene copies of the alpha zein gene family that spread over six loci spanning between 30- and 500-kb chromosomal regions have been described in two Iowa Stiff Stalk (SS) inbreds. Analyses of those regions were possible because of overlapping BAC clones, generated by an expensive and labor-intensive approach. Here we used single-molecule real-time (Pacific Biosciences) shotgun sequencing to assemble the six chromosomal regions from the Non-Stiff Stalk maize inbred W22 from a single DNA sequence dataset. To validate the reconstructed regions, we developed an optical map (BioNano genome map; BioNano Genomics) of W22 and found agreement between the two datasets. Using the sequences of full-length cDNAs from W22, we found that the error rate of PacBio sequencing seemed to be less than 0.1% after autocorrection and assembly. Expressed genes, some with premature stop codons, are interspersed with nonexpressed genes, giving rise to genotype-specific expression differences. Alignment of these regions with those from the previous analyzed regions of SS lines exhibits in part dramatic differences between these two heterotic groups.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.