Menu
July 19, 2019

Dual redundant sequencing strategy: Full-length gene characterisation of 1056 novel and confirmatory HLA alleles.

The high-throughput department of DKMS Life Science Lab encounters novel human leukocyte antigen (HLA) alleles on a daily basis. To characterise these alleles, we have developed a system to sequence the whole gene from 5′- to 3′-UTR for the HLA loci A, B, C, DQB1 and DPB1 for submission to the European Molecular Biology Laboratory – European Nucleotide Archive (EMBL-ENA) and the IPD-IMGT/HLA Database. Our workflow is based on a dual redundant sequencing strategy. Using shotgun sequencing on an Illumina MiSeq instrument and single molecule real-time (SMRT) sequencing on a PacBio RS II instrument, we are able to achieve highly accurate HLA full-length consensus sequences. Remaining conflicts are resolved using the R package DR2S (Dual Redundant Reference Sequencing). Given the relatively high throughput of this strategy, we have developed the semi-automated web service TypeLoader, to aid in the submission of sequences to the EMBL-ENA and the IPD-IMGT/HLA Database. In the IPD-IMGT/HLA Database release 3.24.0 (April 2016; prior to the submission of the sequences described here), only 5.2% of all known HLA alleles have been fully characterised together with intronic and UTR sequences. So far, we have applied our strategy to characterise and submit 1056 HLA alleles, thereby more than doubling the number of fully characterised alleles. Given the increasing application of next generation sequencing (NGS) for full gene characterisation in clinical practice, extending the HLA database concomitantly is highly desirable. Therefore, we propose this dual redundant sequencing strategy as a workflow for submission of novel full-length alleles and characterisation of sequences that are as yet incomplete. This would help to mitigate the predominance of partially known alleles in the database.© 2017 John Wiley & Sons A/S. Published by John Wiley & Sons Ltd.


July 19, 2019

Comparative genomics of two sequential Candida glabrata clinical isolates.

Candida glabrata is an important fungal pathogen which develops rapid antifungal resistance in treated patients. It is known that azole treatments lead to antifungal resistance in this fungal species and that multidrug efflux transporters are involved in this process. Specific mutations in the transcriptional regulator PDR1 result in upregulation of the transporters. In addition, we showed that the PDR1 mutations can contribute to enhance virulence in animal models. In this study, we were interested to compare genomes of two specific C. glabrata-related isolates, one of which was azole susceptible (DSY562) while the other was azole resistant (DSY565). DSY565 contained a PDR1 mutation (L280F) and was isolated after a time-lapse of 50 d of azole therapy. We expected that genome comparisons between both isolates could reveal additional mutations reflecting host adaptation or even additional resistance mechanisms. The PacBio technology used here yielded 14 major contigs (sizes 0.18-1.6 Mb) and mitochondrial genomes from both DSY562 and DSY565 isolates that were highly similar to each other. Comparisons of the clinical genomes with the published CBS138 genome indicated important genome rearrangements, but not between the clinical strains. Among the unique features, several retrotransposons were identified in the genomes of the investigated clinical isolates. DSY562 and DSY565 each contained a large set of adhesin-like genes (101 and 107, respectively), which exceed by far the number of reported adhesins (63) in the CBS138 genome. Comparison between DSY562 and DSY565 yielded 17 nonsynonymous SNPs (among which the was the expected PDR1 mutation) as well as small size indels in coding regions (11) but mainly in adhesin-like genes. The genomes contained a DNA mismatch repair allele of MSH2 known to be involved in the so-called hyper-mutator phenotype of this yeast species and the number of accumulated mutations between both clinical isolates is consistent with the presence of a MSH2 defect. In conclusion, this study is the first to compare genomes of C. glabrata sequential clinical isolates using the PacBio technology as an approach. The genomes of these isolates taken in the same patient at two different time points exhibited limited variations, even if submitted to the host pressure. Copyright © 2017 Vale-Silva et al.


July 19, 2019

Comparative analysis of extended-spectrum-ß-lactamase CTX-M-65-producing Salmonella enterica serovar Infantis isolates from humans, food animals, and retail chickens in the United States.

We sequenced the genomes of ten Salmonella enterica serovar Infantis containing blaCTX-M-65 isolated from chicken, cattle, and human sources collected between 2012 and 2015 in the United States through routine NARMS surveillance and product sampling programs. We also completely assembled the plasmids from four of the isolates. All isolates had a D87Y mutation in the gyrA gene and harbored between 7 and 10 resistance genes (aph (4)-Ia, aac (3)-IVa, aph(3′ )-Ic, blaCTX-M-65, fosA3, floR, dfrA14, sul1, tetA, aadA1) located in two distinct sites of a megaplasmid (~316-323kb) similar to that described in a blaCTX-M-65-positive S. Infantis isolated from a patient in Italy. High-quality single nucleotide polymorphism (hqSNP) analysis revealed that all U.S. isolates were closely related, separated by only 1 to 38 pairwise high quality SNPs, indicating a high likelihood that strains from humans, chicken, and cattle recently evolved from a common ancestor. The U.S. isolates were genetically similar to the blaCTX-M-65-positive S. Infantis isolate from Italy, with a separation of 34 to 47 SNPs. This is the first report of the blaCTX-M-65 gene and the pESI-like megaplasmid from S. Infantis in the United States, and illustrates the importance of applying a global One Health, human and animal perspective to combat antimicrobial resistance. Copyright © 2017 American Society for Microbiology.


July 19, 2019

IG and TR single chain fragment variable (scFv) sequence analysis: a new advanced functionality of IMGT/V-QUEST and IMGT/HighV-QUEST.

IMGT®, the international ImMunoGeneTics information system® ( http://www.imgt.org ), was created in 1989 in Montpellier, France (CNRS and Montpellier University) to manage the huge and complex diversity of the antigen receptors, and is at the origin of immunoinformatics, a science at the interface between immunogenetics and bioinformatics. Immunoglobulins (IG) or antibodies and T cell receptors (TR) are managed and described in the IMGT® databases and tools at the level of receptor, chain and domain. The analysis of the IG and TR variable (V) domain rearranged nucleotide sequences is performed by IMGT/V-QUEST (online since 1997, 50 sequences per batch) and, for next generation sequencing (NGS), by IMGT/HighV-QUEST, the high throughput version of IMGT/V-QUEST (portal begun in 2010, 500,000 sequences per batch). In vitro combinatorial libraries of engineered antibody single chain Fragment variable (scFv) which mimic the in vivo natural diversity of the immune adaptive responses are extensively screened for the discovery of novel antigen binding specificities. However the analysis of NGS full length scFv (~850 bp) represents a challenge as they contain two V domains connected by a linker and there is no tool for the analysis of two V domains in a single chain.The functionality “Analyis of single chain Fragment variable (scFv)” has been implemented in IMGT/V-QUEST and, for NGS, in IMGT/HighV-QUEST for the analysis of the two V domains of IG and TR scFv. It proceeds in five steps: search for a first closest V-REGION, full characterization of the first V-(D)-J-REGION, then search for a second V-REGION and full characterization of the second V-(D)-J-REGION, and finally linker delimitation.For each sequence or NGS read, positions of the 5’V-DOMAIN, linker and 3’V-DOMAIN in the scFv are provided in the ‘V-orientated’ sense. Each V-DOMAIN is fully characterized (gene identification, sequence description, junction analysis, characterization of mutations and amino changes). The functionality is generic and can analyse any IG or TR single chain nucleotide sequence containing two V domains, provided that the corresponding species IMGT reference directory is available.The “Analysis of single chain Fragment variable (scFv)” implemented in IMGT/V-QUEST and, for NGS, in IMGT/HighV-QUEST provides the identification and full characterization of the two V domains of full-length scFv (~850 bp) nucleotide sequences from combinatorial libraries. The analysis can also be performed on concatenated paired chains of expressed antigen receptor IG or TR repertoires.


July 19, 2019

Detecting AGG interruptions in male and female FMR1 premutation carriers by single-molecule sequencing.

The FMR1 gene contains an unstable CGG repeat in its 5′ untranslated region. Premutation alleles range between 55 and 200 repeat units and confer a risk for developing fragile X-associated tremor/ataxia syndrome or fragile X-associated primary ovarian insufficiency. Furthermore, the premutation allele often expands to a full mutation during female germline transmission giving rise to the fragile X syndrome. The risk for a premutation to expand depends mainly on the number of CGG units and the presence of AGG interruptions in the CGG repeat. Unfortunately, the detection of AGG interruptions is hampered by technical difficulties. Here, we demonstrate that single-molecule sequencing enables the determination of not only the repeat size, but also the complete repeat sequence including AGG interruptions in male and female alleles with repeats ranging from 45 to 100 CGG units. We envision this method will facilitate research and diagnostic analysis of the FMR1 repeat expansion. © 2016 WILEY PERIODICALS, INC.


July 19, 2019

Characterization of a large antibiotic resistance plasmid found in enteropathogenic Escherichia coli strain B171 and its relatedness to plasmids of diverse E. coli and Shigella.

Enteropathogenic Escherichia coli (EPEC) is a leading cause of severe infantile diarrhea in developing countries. Previous research has focused on the diversity of the EPEC virulence plasmid, whereas less is known regarding the genetic content and distribution of antibiotic resistance plasmids carried by EPEC. A previous study demonstrated that in addition to the virulence plasmid, reference EPEC strain B171 harbors a second, larger plasmid that confers antibiotic resistance. To further understand the genetic diversity and dissemination of antibiotic resistance plasmids among EPEC strains, we describe the complete sequence of an antibiotic resistance plasmid from EPEC strain B171. The resistance plasmid, pB171_90, has a completed sequence length of 90,229 bp, a GC content of 54.55%, and carries protein-encoding genes involved in conjugative transfer, resistance to tetracycline (tetA), sulfonamides (sulI), and mercury, as well as several virulence-associated genes, including the transcriptional regulator hha and the putative calcium sequestration inhibitor (csi). In silico detection of the pB171_90 genes among 4,798 publicly available E. coli genome assemblies indicates that the unique genes of pB171_90 (csi and traI) are primarily restricted to genomes identified as EPEC or enterotoxigenic E. coli However, conserved regions of the pB171_90 plasmid containing genes involved in replication, stability, and antibiotic resistance were identified among diverse E. coli pathotypes. Interestingly, pB171_90 also exhibited significant similarity with a sequenced plasmid from Shigella dysenteriae type I. Our findings demonstrate the mosaic nature of EPEC antibiotic resistance plasmids and highlight the need for additional sequence-based characterization of antibiotic resistance plasmids harbored by pathogenic E. coli. Copyright © 2017 American Society for Microbiology.


July 19, 2019

Sequencing the CYP2D6 gene: from variant allele discovery to clinical pharmacogenetic testing.

CYP2D6 is one of the most studied enzymes in the field of pharmacogenetics. The CYP2D6 gene is highly polymorphic with over 100 catalogued star (*) alleles, and clinical CYP2D6 testing is increasingly accessible and supported by practice guidelines. However, the degree of variation at the CYP2D6 locus and homology with its pseudogenes make interrogating CYP2D6 by short-read sequencing challenging. Moreover, accurate prediction of CYP2D6 metabolizer status necessitates analysis of duplicated alleles when an increased copy number is detected. These challenges have recently been overcome by long-read CYP2D6 sequencing; however, such platforms are not widely available. This review highlights the genomic complexities of CYP2D6, current sequencing methods and the evolution of CYP2D6 from allele discovery to clinical pharmacogenetic testing.


July 19, 2019

Genomic epidemiology of global Klebsiella pneumoniae carbapenemase (KPC)-producing Escherichia coli.

The dissemination of carbapenem resistance in Escherichia coli has major implications for the management of common infections. bla KPC, encoding a transmissible carbapenemase (KPC), has historically largely been associated with Klebsiella pneumoniae, a predominant plasmid (pKpQIL), and a specific transposable element (Tn4401, ~10?kb). Here we characterize the genetic features of bla KPC emergence in global E. coli, 2008-2013, using both long- and short-read whole-genome sequencing. Amongst 43/45 successfully sequenced bla KPC-E. coli strains, we identified substantial strain diversity (n?=?21 sequence types, 18% of annotated genes in the core genome); substantial plasmid diversity (=9 replicon types); and substantial bla KPC-associated, mobile genetic element (MGE) diversity (50% not within complete Tn4401 elements). We also found evidence of inter-species, regional and international plasmid spread. In several cases bla KPC was found on high copy number, small Col-like plasmids, previously associated with horizontal transmission of resistance genes in the absence of antimicrobial selection pressures. E. coli is a common human pathogen, but also a commensal in multiple environmental and animal reservoirs, and easily transmissible. The association of bla KPC with a range of MGEs previously linked to the successful spread of widely endemic resistance mechanisms (e.g. bla TEM, bla CTX-M) suggests that it may become similarly prevalent.


July 19, 2019

Polylox barcoding reveals haematopoietic stem cell fates realized in vivo.

Developmental deconvolution of complex organs and tissues at the level of individual cells remains challenging. Non-invasive genetic fate mapping has been widely used, but the low number of distinct fluorescent marker proteins limits its resolution. Much higher numbers of cell markers have been generated using viral integration sites, viral barcodes, and strategies based on transposons and CRISPR-Cas9 genome editing; however, temporal and tissue-specific induction of barcodes in situ has not been achieved. Here we report the development of an artificial DNA recombination locus (termed Polylox) that enables broadly applicable endogenous barcoding based on the Cre-loxP recombination system. Polylox recombination in situ reaches a practical diversity of several hundred thousand barcodes, allowing tagging of single cells. We have used this experimental system, combined with fate mapping, to assess haematopoietic stem cell (HSC) fates in vivo. Classical models of haematopoietic lineage specification assume a tree with few major branches. More recently, driven in part by the development of more efficient single-cell assays and improved transplantation efficiencies, different models have been proposed, in which unilineage priming may occur in mice and humans at the level of HSCs. We have introduced barcodes into HSC progenitors in embryonic mice, and found that the adult HSC compartment is a mosaic of embryo-derived HSC clones, some of which are unexpectedly large. Most HSC clones gave rise to multilineage or oligolineage fates, arguing against unilineage priming, and suggesting coherent usage of the potential of cells in a clone. The spreading of barcodes, both after induction in embryos and in adult mice, revealed a basic split between common myeloid-erythroid development and common lymphocyte development, supporting the long-held but contested view of a tree-like haematopoietic structure.


July 19, 2019

SplitThreader: Exploration and analysis of rearrangements in cancer genomes

Genomic rearrangements and associated copy number changes are important drivers in cancer as they can alter the expression of oncogenes and tumor suppressors, create gene fusions, and misregulate gene expression. Here we present SplitThreader (http://splitthreader.com), an open- source interactive web application for analysis and visualization of genomic rearrangements and copy number variation in cancer genomes. SplitThreader constructs a sequence graph of genomic rearrangements in the sample and uses a priority queue breadth-first search algorithm on the graph to search for novel interactions. This is applied to detect gene fusions and other novel sequences, as well as to evaluate distances in the rearranged genome between any genomic regions of interest, especially the repositioning of regulatory elements and their target genes. SplitThreader also analyzes each variant to categorize it by its relation to other variants and by its copy number concordance. This identifies balanced translocations, identifies simple and complex variants, and suggests likely false positives when copy number is not concordant across a candidate breakpoint. It also provides explanations when multiple variants affect the copy number state and obscure the contribution of a single variant, such as a deletion within a region that is overall amplified. Together, these categories triage the variants into groups and provide a starting point for further systematic analysis and manual curation. To demonstrate its utility, we apply SplitThreader to three cancer cell lines, MCF-7 and A549 with Illumina paired- end sequencing, and SK-BR-3, with long-read PacBio sequencing. Using SplitThreader, we examine the genomic rearrangements responsible for previously observed gene fusions in SK-BR-3 and MCF-7, and discover many of the fusions involved a complex series of multiple genomic rearrangements. We also find notable differences in the types of variants between the three cell lines, in particular a much higher proportion of reciprocal variants in SK-BR-3 and a distinct clustering of interchromosomal variants in SK-BR-3 and MCF-7 that is absent in A549.


July 19, 2019

Parkinson’s disease associated with pure ATXN10 repeat

Large, non-coding pentanucleotide repeat expansions of ATTCT in intron 9 of the ATXN10 gene typically cause progressive spinocerebellar ataxia with or without seizures and present neuropathologically with Purkinje cell loss resulting in symmetrical cerebellar atrophy. These ATXN10 repeat expansions can be interrupted by sequence motifs which have been attributed to seizures and are likely to act as genetic modifiers. We identified a Mexican kindred with multiple affected family members with ATXN10 expansions. Four affected family members showed clinical features of spinocerebellar ataxia type 10 (SCA10). However, one affected individual presented with early-onset levodopa-responsive parkinsonism, and one family member carried a large repeat ATXN10 expansion, but was clinically unaffected. To characterize the ATXN10 repeat, we used a novel technology of single-molecule real-time (SMRT) sequencing and CRISPR/Cas9-based capture. We sequenced the entire span of ~5.3–7.0kb repeat expansions. The Parkinson’s patient carried an ATXN10 expansion with no repeat interruption motifs as well as an unaffected sister. In the siblings with typical SCA10, we found a repeat pattern of ATTCC repeat motifs that have not been associated with seizures previously. Our data suggest that the absence of repeat interruptions is likely a genetic modifier for the clinical presentation of L-Dopa responsive parkinsonism, whereas repeat interruption motifs contribute clinically to epilepsy. Repeat interruptions are important genetic modifiers of the clinical phenotype in SCA10. Advanced sequencing techniques now allow to better characterize the underlying genetic architecture for determining accurate phenotype–genotype correlations.


July 19, 2019

Monitoring microevolution of OXA-48-producing Klebsiella pneumoniae ST147 in a hospital setting by SMRT sequencing.

Carbapenemase-producing Klebsiella pneumoniae pose an increasing risk for healthcare facilities worldwide. A continuous monitoring of ST distribution and its association with resistance and virulence genes is required for early detection of successful K. pneumoniae lineages. In this study, we used WGS to characterize MDR blaOXA-48-positive K. pneumoniae isolated from inpatients at the University Medical Center Göttingen, Germany, between March 2013 and August 2014.Closed genomes for 16 isolates of carbapenemase-producing K. pneumoniae were generated by single molecule real-time technology using the PacBio RSII platform.Eight of the 16 isolates showed identical XbaI macrorestriction patterns and shared the same MLST, ST147. The eight ST147 isolates differed by only 1-25 SNPs of their core genome, indicating a clonal origin. Most of the eight ST147 isolates carried four plasmids with sizes of 246.8, 96.1, 63.6 and 61.0?kb and a novel linear plasmid prophage, named pKO2, of 54.6?kb. The blaOXA-48 gene was located on a 63.6?kb IncL plasmid and is part of composite transposon Tn1999.2. The ST147 isolates expressed the yersinabactin system as a major virulence factor. The comparative whole-genome analysis revealed several rearrangements of mobile genetic elements and losses of chromosomal and plasmidic regions in the ST147 isolates.Single molecule real-time sequencing allowed monitoring of the genetic and epigenetic microevolution of MDR OXA-48-producing K. pneumoniae and revealed in addition to SNPs, complex rearrangements of genetic elements.© The Author 2017. Published by Oxford University Press on behalf of the British Society for Antimicrobial Chemotherapy. All rights reserved. For Permissions, please email: journals.permissions@oup.com.


July 19, 2019

A novel approach using long-read sequencing and ddPCR to investigate gonadal mosaicism and estimate recurrence risk in two families with developmental disorders.

De novo mutations contribute significantly to severe early-onset genetic disorders. Even if the mutation is apparently de novo, there is a recurrence risk due to parental germ line mosaicism, depending on in which gonadal generation the mutation occurred.We demonstrate the power of using SMRT sequencing and ddPCR to determine parental origin and allele frequencies of de novo mutations in germ cells in two families whom had undergone assisted reproduction.In the first family, a TCOF1 variant c.3156C>T was identified in the proband with Treacher Collins syndrome. The variant affects splicing and was determined to be of paternal origin. It was present in <1% of the paternal germ cells, suggesting a very low recurrence risk. In the second family, the couple had undergone several unsuccessful pregnancies where a de novo mutation PTPN11 c.923A>C causing Noonan syndrome was identified. The variant was present in 40% of the paternal germ cells suggesting a high recurrence risk.Our findings highlight a successful strategy to identify the parental origin of mutations and to investigate the recurrence risk in couples that have undergone assisted reproduction with an unknown donor or in couples with gonadal mosaicism that will undergo preimplantation genetic diagnosis.© 2017 The Authors Prenatal Diagnosis published by John Wiley & Sons Ltd.


July 19, 2019

Exonization of an intronic LINE-1 element causing Becker muscular dystrophy as a novel mutational mechanism in dystrophin gene.

A broad mutational spectrum in the dystrophin (DMD) gene, from large deletions/duplications to point mutations, causes Duchenne/Becker muscular dystrophy (D/BMD). Comprehensive genotyping is particularly relevant considering the mutation-centered therapies for dystrophinopathies. We report the genetic characterization of a patient with disease onset at age 13 years, elevated creatine kinase levels and reduced dystrophin labeling, where multiplex-ligation probe amplification (MLPA) and genomic sequencing failed to detect pathogenic variants. Bioinformatic, transcriptomic (real time PCR, RT-PCR), and genomic approaches (Southern blot, long-range PCR, and single molecule real-time sequencing) were used to characterize the mutation. An aberrant transcript was identified, containing a 103-nucleotide insertion between exons 51 and 52, with no similarity with the DMD gene. This corresponded to the partial exonization of a long interspersed nuclear element (LINE-1), disrupting the open reading frame. Further characterization identified a complete LINE-1 (~6 kb with typical hallmarks) deeply inserted in intron 51. Haplotyping and segregation analysis demonstrated that the mutation had a de novo origin. Besides underscoring the importance of mRNA studies in genetically unsolved cases, this is the first report of a disease-causing fully intronic LINE-1 element in DMD, adding to the diversity of mutational events that give rise to D/BMD.


July 19, 2019

Amplification-free, CRISPR-Cas9 targeted enrichment and SMRT Sequencing of repeat-expansion disease causative genomic regions

Targeted sequencing has proven to be an economical means of obtaining sequence information for one or more defined regions of a larger genome. However, most target enrichment methods require amplification. Some genomic regions, such as those with extreme GC content and repetitive sequences, are recalcitrant to faithful amplification. Yet, many human genetic disorders are caused by repeat expansions, including difficult to sequence tandem repeats. We have developed a novel, amplification-free enrichment technique that employs the CRISPR-Cas9 system for specific targeting multiple genomic loci. This method, in conjunction with long reads generated through Single Molecule, Real-Time (SMRT) sequencing and unbiased coverage, enables enrichment and sequencing of complex genomic regions that cannot be investigated with other technologies. Using human genomic DNA samples, we demonstrate successful targeting of causative loci for Huntingtontextquoterights disease (HTT; CAG repeat), Fragile X syndrome (FMR1; CGG repeat), amyotrophic lateral sclerosis (ALS) and frontotemporal dementia (C9orf72; GGGGCC repeat), and spinocerebellar ataxia type 10 (SCA10) (ATXN10; variable ATTCT repeat). The method, amenable to multiplexing across multiple genomic loci, uses an amplification-free approach that facilitates the isolation of hundreds of individual on-target molecules in a single SMRT Cell and accurate sequencing through long repeat stretches, regardless of extreme GC percent or sequence complexity content. Our novel targeted sequencing method opens new doors to genomic analyses independent of PCR amplification that will facilitate the study of repeat expansion disorders.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.