Menu
July 19, 2019

High-quality assembly of an individual of Yoruban descent

De novo assembly of human genomes is now a tractable effort due in part to advances in sequencing and mapping technologies. We use PacBio single-molecule, real-time (SMRT) sequencing and BioNano genomic maps to construct the first de novo assembly of NA19240, a Yoruban individual from Africa. This chromosome-scaffolded assembly of 3.08 Gb with a contig N50 of 7.25 Mb and a scaffold N50 of 78.6 Mb represents one of the most contiguous high-quality human genomes. We utilize a BAC library derived from NA19240 DNA and novel haplotype-resolving sequencing technologies and algorithms to characterize regions of complex genomic architecture that are normally lost due to compression to a linear haploid assembly. Our results demonstrate that multiple technologies are still necessary for complete genomic representation, particularly in regions of highly identical segmental duplications. Additionally, we show that diploid assembly has utility in improving the quality of de novo human genome assemblies.


July 19, 2019

Emergence of a Homo sapiens-specific gene family and chromosome 16p11.2 CNV susceptibility.

Genetic differences that specify unique aspects of human evolution have typically been identified by comparative analyses between the genomes of humans and closely related primates, including more recently the genomes of archaic hominins. Not all regions of the genome, however, are equally amenable to such study. Recurrent copy number variation (CNV) at chromosome 16p11.2 accounts for approximately 1% of cases of autism and is mediated by a complex set of segmental duplications, many of which arose recently during human evolution. Here we reconstruct the evolutionary history of the locus and identify bolA family member 2 (BOLA2) as a gene duplicated exclusively in Homo sapiens. We estimate that a 95-kilobase-pair segment containing BOLA2 duplicated across the critical region approximately 282 thousand years ago (ka), one of the latest among a series of genomic changes that dramatically restructured the locus during hominid evolution. All humans examined carried one or more copies of the duplication, which nearly fixed early in the human lineage-a pattern unlikely to have arisen so rapidly in the absence of selection (P?


July 19, 2019

Towards precision medicine.

There is great potential for genome sequencing to enhance patient care through improved diagnostic sensitivity and more precise therapeutic targeting. To maximize this potential, genomics strategies that have been developed for genetic discovery – including DNA-sequencing technologies and analysis algorithms – need to be adapted to fit clinical needs. This will require the optimization of alignment algorithms, attention to quality-coverage metrics, tailored solutions for paralogous or low-complexity areas of the genome, and the adoption of consensus standards for variant calling and interpretation. Global sharing of this more accurate genotypic and phenotypic data will accelerate the determination of causality for novel genes or variants. Thus, a deeper understanding of disease will be realized that will allow its targeting with much greater therapeutic precision.


July 19, 2019

Biosynthesis and function of modified bases in bacteria and their viruses.

Naturally occurring modification of the canonical A, G, C, and T bases can be found in the DNA of cellular organisms and viruses from all domains of life. Bacterial viruses (bacteriophages) are a particularly rich but still underexploited source of such modified variant nucleotides. The modifications conserve the coding and base-pairing functions of DNA, but add regulatory and protective functions. In prokaryotes, modified bases appear primarily to be part of an arms race between bacteriophages (and other genomic parasites) and their hosts, although, as in eukaryotes, some modifications have been adapted to convey epigenetic information. The first half of this review catalogs the identification and diversity of DNA modifications found in bacteria and bacteriophages. What is known about the biogenesis, context, and function of these modifications are also described. The second part of the review places these DNA modifications in the context of the arms race between bacteria and bacteriophages. It focuses particularly on the defense and counter-defense strategies that turn on direct recognition of the presence of a modified base. Where modification has been shown to affect other DNA transactions, such as expression and chromosome segregation, that is summarized, with reference to recent reviews.


July 19, 2019

A distinct class of chromoanagenesis events characterized by focal copy number gains.

Chromoanagenesis is the process by which a single catastrophic event creates complex rearrangements confined to a single or a few chromosomes. It is usually characterized by the presence of multiple deletions and/or duplications, as well as by copy neutral rearrangements. In contrast, an array CGH screen of patients with developmental anomalies revealed three patients in which a single chromosome carries from 8 to 11 large copy number gains confined to a single chromosome or chromosomal arm, but the absence of deletions. Subsequent fluorescence in situ hybiridization and massive parallel sequencing revealed the duplicons to be clustered together in distinct locations across the altered chromosomes. Breakpoint junction sequences showed both microhomology and non-templated insertions of up to 40 bp. Hence, these patients each demonstrate a single altered chromosome of clustered insertional duplications, no deletions, and breakpoint junction sequences showing microhomology and/or non-templated insertions. These observations are difficult to reconcile with current mechanistic descriptions of chromothripsis and chromoanasynthesis. Therefore, we hypothesize those rearrangements to be of a mechanistically different origin. In addition, we suggest that large untemplated insertional sequences observed at breakpoints are driven by a non-canonical non-homologous end joining mechanism.© 2016 WILEY PERIODICALS, INC.


July 19, 2019

ARTISAN PCR: rapid identification of full-length immunoglobulin rearrangements without primer binding bias.

B cells recognize specific antigens by their membrane-bound B-cell receptor (BCR). Functional BCR genes are assembled in pre-B cells by recombination of the variable (V), diversity (D) and joining (J) genes [V(D)J recombination]. When B cells participate in germinal centre reactions, non-templated point mutations are introduced into BCR genes by somatic hypermutation (SHM) (Rajewsky, 1996). V(D)J recombination and SHM create virtually unlimited BCR repertoires.


July 19, 2019

High throughput random mutagenesis and Single Molecule Real Time Sequencing of the muscle nicotinic acetylcholine receptor.

High throughput random mutagenesis is a powerful tool to identify which residues are important for the function of a protein, and gain insight into its structure-function relation. The human muscle nicotinic acetylcholine receptor was used to test whether this technique previously used for monomeric receptors can be applied to a pentameric ligand-gated ion channel. A mutant library for the a1 subunit of the channel was generated by error-prone PCR, and full length sequences of all 2816 mutants were retrieved using single molecule real time sequencing. Each a1 mutant was co-transfected with wildtype ß1, d, and e subunits, and the channel function characterized by an ion flux assay. To test whether the strategy could map the structure-function relation of this receptor, we attempted to identify mutations that conferred resistance to competitive antagonists. Mutant hits were defined as receptors that responded to the nicotinic agonist epibatidine, but were not inhibited by either a-bungarotoxin or tubocurarine. Eight a1 subunit mutant hits were identified, six of which contained mutations at position Y233 or V275 in the transmembrane domain. Three single point mutations (Y233N, Y233H, and V275M) were studied further, and found to enhance the potencies of five channel agonists tested. This suggests that the mutations made the channel resistant to the antagonists, not by impairing antagonist binding, but rather by producing a gain-of-function phenotype, e.g. increased agonist sensitivity. Our data show that random high throughput mutagenesis is applicable to multimeric proteins to discover novel functional mutants, and outlines the benefits of using single molecule real time sequencing with regards to quality control of the mutant library as well as downstream mutant data interpretation.


July 19, 2019

De novo assembly and phasing of a Korean human genome.

Advances in genome assembly and phasing provide an opportunity to investigate the diploid architecture of the human genome and reveal the full range of structural variation across population groups. Here we report the de novo assembly and haplotype phasing of the Korean individual AK1 (ref. 1) using single-molecule real-time sequencing, next-generation mapping, microfluidics-based linked reads, and bacterial artificial chromosome (BAC) sequencing approaches. Single-molecule sequencing coupled with next-generation mapping generated a highly contiguous assembly, with a contig N50 size of 17.9?Mb and a scaffold N50 size of 44.8?Mb, resolving 8 chromosomal arms into single scaffolds. The de novo assembly, along with local assemblies and spanning long reads, closes 105 and extends into 72 out of 190 euchromatic gaps in the reference genome, adding 1.03?Mb of previously intractable sequence. High concordance between the assembly and paired-end sequences from 62,758 BAC clones provides strong support for the robustness of the assembly. We identify 18,210 structural variants by direct comparison of the assembly with the human reference, identifying thousands of breakpoints that, to our knowledge, have not been reported before. Many of the insertions are reflected in the transcriptome and are shared across the Asian population. We performed haplotype phasing of the assembly with short reads, long reads and linked reads from whole-genome sequencing and with short reads from 31,719 BAC clones, thereby achieving phased blocks with an N50 size of 11.6?Mb. Haplotigs assembled from single-molecule real-time reads assigned to haplotypes on phased blocks covered 89% of genes. The haplotigs accurately characterized the hypervariable major histocompatability complex region as well as demonstrating allele configuration in clinically relevant genes such as CYP2D6. This work presents the most contiguous diploid human genome assembly so far, with extensive investigation of unreported and Asian-specific structural variants, and high-quality haplotyping of clinically relevant alleles for precision medicine.


July 19, 2019

Extraction of high-molecular-weight genomic DNA for long-read sequencing of single molecules.

De novo sequencing of complex genomes is one of the main challenges for researchers seeking high-quality reference sequences. Many de novo assemblies are based on short reads, producing fragmented genome sequences. Third-generation sequencing, with read lengths >10 kb, will improve the assembly of complex genomes, but these techniques require high-molecular-weight genomic DNA (gDNA), and gDNA extraction protocols used for obtaining smaller fragments for short-read sequencing are not suitable for this purpose. Methods of preparing gDNA for bacterial artificial chromosome (BAC) libraries could be adapted, but these approaches are time-consuming, and commercial kits for these methods are expensive. Here, we present a protocol for rapid, inexpensive extraction of high-molecular-weight gDNA from bacteria, plants, and animals. Our technique was validated using sunflower leaf samples, producing a mean read length of 12.6 kb and a maximum read length of 80 kb.


July 19, 2019

IncFIIk plasmid harbouring an amplification of 16S rRNA methyltransferase-encoding gene rmtH associated with mobile element ISCR2.

To investigate the resistance mechanisms and genetic support underlying the high resistance level of the Klebsiella pneumoniae strain CMUL78 to aminoglycoside and ß-lactam antibiotics.Antibiotic susceptibility was assessed by the disc diffusion method and MICs were determined by the microdilution method. Antibiotic resistance genes and their genetic environment were characterized by PCR and Sanger sequencing. Plasmid contents were analysed in the clinical strain and transconjugants obtained by mating-out assays. Complete plasmid sequencing was performed with PacBio and Illumina technology.Strain CMUL78 co-produced the 16S rRNA methyltransferase (RMTase) RmtH, carbapenemase OXA-48 and ESBL SHV-12. The rmtH- and blaSHV-12-encoding genes were harboured by a novel ~115 kb IncFIIk plasmid designated pRmtH, and blaOXA-48 by a ~62 kb IncL/M plasmid related to pOXA-48a. pRmtH plasmid possessed seven different stability modules, one of which is a novel hybrid toxin-antitoxin system. Interestingly, pRmtH plasmid harboured a 4-fold amplification of an rmtH-ISCR2 unit arranged in tandem and inserted within a novel IS26-based composite transposon designated Tn6329.This is the first known report of the 16S RMTase-encoding gene rmtH in a plasmid. The rmtH-ISCR2 unit was inserted in a composite transposon as a 4-fold tandem repeat, a scarcely reported organization.© The Author 2016. Published by Oxford University Press on behalf of the British Society for Antimicrobial Chemotherapy. All rights reserved. For Permissions, please e-mail: journals.permissions@oup.com.


July 19, 2019

Full-length mitochondrial-DNA sequencing on the PacBio RSII.

Conventional mitochondrial-DNA (MT DNA) sequencing approaches use Sanger sequencing of 20-40 partially overlapping PCR fragments per individual, which is a time- and resource-consuming process. We have developed a high-throughput, accurate, fast, and cost-effective human MT DNA sequencing approach. In this setup we first generate long-range PCR products for two partially overlapping 7.7 and 9.2 kb MT DNA-specific amplicons, add sample-specific barcodes, and sequence these on the PacBio RSII system to obtain full-length MT DNA sequences for genotyping/haplotyping purposes.


July 19, 2019

Ribbon: Visualizing complex genome alignments and structural variation

Visualization has played an extremely important role in the current genomic revolution to inspect and understand variants, expression patterns, evolutionary changes, and a number of other relationships. However, most of the information in read-to-reference or genome-genome alignments is lost for structural variations in the one-dimensional views of most genome browsers showing only reference coordinates. Instead, structural variations captured by long reads or assembled contigs often need more context to understand, including alignments and other genomic information from multiple chromosomes. We have addressed this problem by creating Ribbon (genomeribbon.com) an interactive online visualization tool that displays alignments along both reference and query sequences, along with any associated variant calls in the sample. This way Ribbon shows patterns in alignments of many reads across multiple chromosomes, while allowing detailed inspection of individual reads (Supplementary Note 1). For example, here we show a gene fusion in the SK-BR-3 breast cancer cell line linking the genes CYTH1 and EIF3H. While it has been found in the transcriptome previously, genome sequencing did not identify a direct chromosomal fusion between these two genes. After SMRT sequencing, Ribbon shows that there are indeed long reads that span from one gene to the other, going through not one but two variants, for the first time showing the genomic link between these two genes (Figure 1a). More gene fusions of this cancer cell line are investigated in Supplementary Note 2. Figure 1b shows another complex event in this sample made simple in Ribbon: the translocation of a 4.4 kb sequence deleted from chr19 and inserted into chr16 (Figure 1b). Thus, Ribbon enables understanding of complex variants, and it may also help in the detection of sequencing and sample preparation issues, testing of aligners and variant-callers, and rapid curation of structural variant candidates (Supplementary Note 3). In addition to SAM and BAM files with long, short, or paired-end reads, Ribbon can also load coordinate files from whole genome aligners such as MUMmer. Therefore, Ribbon can be used to test assembly algorithms or inspect the similarity between species. Supplementary Note 4 shows a comparison of gorilla and human genomes using Ribbon, highlighting major structural differences. In conclusion, Ribbon is a powerful interactive web tool for viewing complex genomic alignments.


July 19, 2019

Comparative DNA methylation and gene expression analysis identifies novel genes for structural congenital heart diseases.

For the majority of congenital heart diseases (CHDs), the full complexity of the causative molecular network, which is driven by genetic, epigenetic, and environmental factors, is yet to be elucidated. Epigenetic alterations are suggested to play a pivotal role in modulating the phenotypic expression of CHDs and their clinical course during life. Candidate approaches implied that DNA methylation might have a developmental role in CHD and contributes to the long-term progress of non-structural cardiac diseases. The aim of the present study is to define the postnatal epigenome of two common cardiac malformations, representing epigenetic memory, and adaption to hemodynamic alterations, which are jointly relevant for the disease course.We present the first analysis of genome-wide DNA methylation data obtained from myocardial biopsies of Tetralogy of Fallot (TOF) and ventricular septal defect patients. We defined stringent sets of differentially methylated regions between patients and controls, which are significantly enriched for genomic features like promoters, exons, and cardiac enhancers. For TOF, we linked DNA methylation with genome-wide expression data and found a significant overlap for hypermethylated promoters and down-regulated genes, and vice versa. We validated and replicated the methylation of selected CpGs and performed functional assays. We identified a hypermethylated novel developmental CpG island in the promoter of SCO2 and demonstrate its functional impact. Moreover, we discovered methylation changes co-localized with novel, differential splicing events among sarcomeric genes as well as transcription factor binding sites. Finally, we demonstrated the interaction of differentially methylated and expressed genes in TOF with mutated CHD genes in a molecular network.By interrogating DNA methylation and gene expression data, we identify two novel mechanism contributing to the phenotypic expression of CHDs: aberrant methylation of promoter CpG islands and methylation alterations leading to differential splicing. Published on behalf of the European Society of Cardiology. All rights reserved. © The Author 2016. For permissions please email: journals.permissions@oup.com.


July 19, 2019

Single-molecule sequencing revealing the presence of distinct JC polyomavirus populations in patients with progressive multifocal leukoencephalopathy.

Progressive multifocal leukoencephalopathy (PML) is a fatal disease caused by reactivation of JC polyomavirus (JCPyV) in immunosuppressed individuals and lytic infection by neurotropic JCPyV in glial cells. The exact content of neurotropic mutations within individual JCPyV strains has not been studied to our knowledge.We exploited the capacity of single-molecule real-time sequencing technology to determine the sequence of complete JCPyV genomes in single reads. The method was used to precisely characterize individual neurotropic JCPyV strains of 3 patients with PML without the bias caused by assembly of short sequence reads.In the cerebrospinal fluid sample of a 73-year-old woman with rapid PML onset, 3 distinct JCPyV populations could be identified. All viral populations were characterized by rearrangements within the noncoding regulatory region (NCCR) and 1 point mutation, S267L in the VP1 gene, suggestive of neurotropic strains. One patient with PML had a single neurotropic strain with rearranged NCCR, and 1 patient had a single strain with small NCCR alterations.We report here, for the first time, full characterization of individual neurotropic JCPyV strains in the cerebrospinal fluid of patients with PML. It remains to be established whether PML pathogenesis is driven by one or several neurotropic strains in an individual.


July 19, 2019

The establishment and diversification of epidemic-associated serogroup W meningococcus in the African meningitis belt, 1994 to 2012.

Epidemics of invasive meningococcal disease (IMD) caused by meningococcal serogroup A have been eliminated from the sub-Saharan African so-called “meningitis belt” by the meningococcal A conjugate vaccine (MACV), and yet, other serogroups continue to cause epidemics. Neisseria meningitidis serogroup W remains a major cause of disease in the region, with most isolates belonging to clonal complex 11 (CC11). Here, the genetic variation within and between epidemic-associated strains was assessed by sequencing the genomes of 92 N. meningitidis serogroup W isolates collected between 1994 and 2012 from both sporadic and epidemic IMD cases, 85 being from selected meningitis belt countries. The sequenced isolates belonged to either CC175 (n = 9) or CC11 (n = 83). The CC11 N. meningitidis serogroup W isolates belonged to a single lineage comprising four major phylogenetic subclades. Separate CC11 N. meningitidis serogroup W subclades were associated with the 2002 and 2012 Burkina Faso epidemics. The subclade associated with the 2012 epidemic included isolates found in Burkina Faso and Mali during 2011 and 2012, which descended from a strain very similar to the Hajj (Islamic pilgrimage to Mecca)-related Saudi Arabian outbreak strain from 2000. The phylogeny of isolates from 2012 reflected their geographic origin within Burkina Faso, with isolates from the Malian border region being closely related to the isolates from Mali. Evidence of ongoing evolution, international transmission, and strain replacement stresses the importance of maintaining N. meningitidis surveillance in Africa following the MACV implementation. IMPORTANCE Meningococcal disease (meningitis and bloodstream infections) threatens millions of people across the meningitis belt of sub-Saharan Africa. A vaccine introduced in 2010 protects against Africa’s then-most common cause of meningococcal disease, N. meningitidis serogroup A. However, other serogroups continue to cause epidemics in the region-including serogroup W. The rapid identification of strains that have been associated with prior outbreaks can improve the assessment of outbreak risk and enable timely preparation of public health responses, including vaccination. Phylogenetic analysis of newly sequenced serogroup W strains isolated from 1994 to 2012 identified two groups of strains linked to large epidemics in Burkina Faso, one being descended from a strain that caused an outbreak during the Hajj pilgrimage in 2000. We find that applying whole-genome sequencing to meningococcal disease surveillance collections improves the discrimination among strains, even within a single nation-wide epidemic, which can be used to better understand pathogen spread.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.