Menu
September 22, 2019

Discovery of gorilla MHC-C expressing C1 ligand for KIR.

In comparison to humans and chimpanzees, gorillas show low diversity at MHC class I genes (Gogo), as reflected by an overall reduced level of allelic variation as well as the absence of a functionally important sequence motif that interacts with killer cell immunoglobulin-like receptors (KIR). Here, we use recently generated large-scale genomic sequence data for a reassessment of allelic diversity at Gogo-C, the gorilla orthologue of HLA-C. Through the combination of long-range amplifications and long-read sequencing technology, we obtained, among the 35 gorillas reanalyzed, three novel full-length genomic sequences including a coding region sequence that has not been previously described. The newly identified Gogo-C*03:01 allele has a divergent recombinant structure that sets it apart from other Gogo-C alleles. Domain-by-domain phylogenetic analysis shows that Gogo-C*03:01 has segments in common with Gogo-B*07, the additional B-like gene that is present on some gorilla MHC haplotypes. Identified in ~ 50% of the gorillas analyzed, the Gogo-C*03:01 allele exclusively encodes the C1 epitope among Gogo-C allotypes, indicating its important function in controlling natural killer cell (NK cell) responses via KIR. We further explored the hypothesis whether gorillas experienced a selective sweep which may have resulted in a general reduction of the gorilla MHC class I repertoire. Our results provide little support for a selective sweep but rather suggest that the overall low Gogo class I diversity can be best explained by drastic demographic changes gorillas experienced in the ancient and recent past.


September 22, 2019

Double insertion of transposable elements provides a substrate for the evolution of satellite DNA.

Eukaryotic genomes are replete with repeated sequences in the form of transposable elements (TEs) dispersed across the genome or as satellite arrays, large stretches of tandemly repeated sequences. Many satellites clearly originated as TEs, but it is unclear how mobile genetic parasites can transform into megabase-sized tandem arrays. Comprehensive population genomic sampling is needed to determine the frequency and generative mechanisms of tandem TEs, at all stages from their initial formation to their subsequent expansion and maintenance as satellites. The best available population resources, short-read DNA sequences, are often considered to be of limited utility for analyzing repetitive DNA due to the challenge of mapping individual repeats to unique genomic locations. Here we develop a new pipeline called ConTExt that demonstrates that paired-end Illumina data can be successfully leveraged to identify a wide range of structural variation within repetitive sequence, including tandem elements. By analyzing 85 genomes from five populations of Drosophila melanogaster, we discover that TEs commonly form tandem dimers. Our results further suggest that insertion site preference is the major mechanism by which dimers arise and that, consequently, dimers form rapidly during periods of active transposition. This abundance of TE dimers has the potential to provide source material for future expansion into satellite arrays, and we discover one such copy number expansion of the DNA transposon hobo to approximately 16 tandem copies in a single line. The very process that defines TEs-transposition-thus regularly generates sequences from which new satellites can arise.© 2018 McGurk and Barbash; Published by Cold Spring Harbor Laboratory Press.


September 22, 2019

Mutant phenotypes for thousands of bacterial genes of unknown function.

One-third of all protein-coding genes from bacterial genomes cannot be annotated with a function. Here, to investigate the functions of these genes, we present genome-wide mutant fitness data from 32 diverse bacteria across dozens of growth conditions. We identified mutant phenotypes for 11,779 protein-coding genes that had not been annotated with a specific function. Many genes could be associated with a specific condition because the gene affected fitness only in that condition, or with another gene in the same bacterium because they had similar mutant phenotypes. Of the poorly annotated genes, 2,316 had associations that have high confidence because they are conserved in other bacteria. By combining these conserved associations with comparative genomics, we identified putative DNA repair proteins; in addition, we propose specific functions for poorly annotated enzymes and transporters and for uncharacterized protein families. Our study demonstrates the scalability of microbial genetics and its utility for improving gene annotations.


September 22, 2019

The complete chloroplast genome sequence of Actinidia arguta using the PacBio RS II platform.

Actinidia arguta is the most basal species in a phylogenetically and economically important genus in the family Actinidiaceae. To better understand the molecular basis of the Actinidia arguta chloroplast (cp), we sequenced the complete cp genome from A. arguta using Illumina and PacBio RS II sequencing technologies. The cp genome from A. arguta was 157,611 bp in length and composed of a pair of 24,232 bp inverted repeats (IRs) separated by a 20,463 bp small single copy region (SSC) and an 88,684 bp large single copy region (LSC). Overall, the cp genome contained 113 unique genes. The cp genomes from A. arguta and three other Actinidia species from GenBank were subjected to a comparative analysis. Indel mutation events and high frequencies of base substitution were identified, and the accD and ycf2 genes showed a high degree of variation within Actinidia. Forty-seven simple sequence repeats (SSRs) and 155 repetitive structures were identified, further demonstrating the rapid evolution in Actinidia. The cp genome analysis and the identification of variable loci provide vital information for understanding the evolution and function of the chloroplast and for characterizing Actinidia population genetics.


September 22, 2019

Multiplex assessment of protein variant abundance by massively parallel sequencing.

Determining the pathogenicity of genetic variants is a critical challenge, and functional assessment is often the only option. Experimentally characterizing millions of possible missense variants in thousands of clinically important genes requires generalizable, scalable assays. We describe variant abundance by massively parallel sequencing (VAMP-seq), which measures the effects of thousands of missense variants of a protein on intracellular abundance simultaneously. We apply VAMP-seq to quantify the abundance of 7,801 single-amino-acid variants of PTEN and TPMT, proteins in which functional variants are clinically actionable. We identify 1,138 PTEN and 777 TPMT variants that result in low protein abundance, and may be pathogenic or alter drug metabolism, respectively. We observe selection for low-abundance PTEN variants in cancer, and show that p.Pro38Ser, which accounts for ~10% of PTEN missense variants in melanoma, functions via a dominant-negative mechanism. Finally, we demonstrate that VAMP-seq is applicable to other genes, highlighting its generalizability.


September 22, 2019

Directed evolution of multiple genomic loci allows the prediction of antibiotic resistance.

Antibiotic development is frequently plagued by the rapid emergence of drug resistance. However, assessing the risk of resistance development in the preclinical stage is difficult. Standard laboratory evolution approaches explore only a small fraction of the sequence space and fail to identify exceedingly rare resistance mutations and combinations thereof. Therefore, new rapid and exhaustive methods are needed to accurately assess the potential of resistance evolution and uncover the underlying mutational mechanisms. Here, we introduce directed evolution with random genomic mutations (DIvERGE), a method that allows an up to million-fold increase in mutation rate along the full lengths of multiple predefined loci in a range of bacterial species. In a single day, DIvERGE generated specific mutation combinations, yielding clinically significant resistance against trimethoprim and ciprofloxacin. Many of these mutations have remained previously undetected or provide resistance in a species-specific manner. These results indicate pathogen-specific resistance mechanisms and the necessity of future narrow-spectrum antibacterial treatments. In contrast to prior claims, we detected the rapid emergence of resistance against gepotidacin, a novel antibiotic currently in clinical trials. Based on these properties, DIvERGE could be applicable to identify less resistance-prone antibiotics at an early stage of drug development. Finally, we discuss potential future applications of DIvERGE in synthetic and evolutionary biology. Copyright © 2018 the Author(s). Published by PNAS.


September 22, 2019

Genome Assembly.

Genome assembly uses sequence similarity to go from sequencing reads to longer contiguous sequences (contigs). Scaffolds are contigs linked together by gaps where the order and orientation of the contigs is known but the exact sequence connecting two contigs is unknown, represented by Ns which estimate the gap length. Here we describe recommendations for genome assembly for different sequencing technologies, describe organelle assembly, and review how to perform assembly quality control.


September 22, 2019

Diversity of hepatitis E virus genotype 3

Summary Hepatitis E virus genotype 3 (HEV-3) can lead to chronic infection in immunocompromised patients, and ribavirin is the treatment of choice. Recently, mutations in the polymerase gene have been associated with ribavirin failure but their frequency before treatment according to HEV-3 subtypes has not been studied on a large data set. We used single-molecule real-time sequencing technology to sequence 115 new complete genomes of HEV-3 infecting French patients. We analyzed phylogenetic relationships, the length of the polyproline region, and mutations in the HEV polymerase gene. Eighty-five (74%) were in the clade HEV-3efg, 28 (24%) in HEV-3chi clade, and 2 (2%) in HEV-3ra clade. Using automated partitioning of maximum likelihood phylogenetic trees, complete genomes were classified into subtypes. Polyproline region length differs within HEV-3 clades (from 189 to 315 nt). Investigating mutations in the polymerase gene, distinct polymorphisms between HEV-3 subtypes were found (G1634R in 95% of HEV-3e, G1634K in 56% of HEV-3ra, and V1479I in all HEV-3efg, clade HEV-3ra, and HEV-3k strains). Subtype-specific polymorphisms in the HEV-3 polymerase have been identified. Our study provides new complete genome sequences of HEV-3 that could be useful for comparing strains circulating in humans and the animal reservoir.


September 22, 2019

Tumor-specific mitochondrial DNA variants are rarely detected in cell-free DNA.

The use of blood-circulating cell-free DNA (cfDNA) as a “liquid biopsy” in oncology is being explored for its potential as a cancer biomarker. Mitochondria contain their own circular genomic entity (mitochondrial DNA, mtDNA), up to even thousands of copies per cell. The mutation rate of mtDNA is several orders of magnitude higher than that of the nuclear DNA. Tumor-specific variants have been identified in tumors along the entire mtDNA, and their number varies among and within tumors. The high mtDNA copy number per cell and the high mtDNA mutation rate make it worthwhile to explore the potential of tumor-specific cf-mtDNA variants as cancer marker in the blood of cancer patients. We used single-molecule real-time (SMRT) sequencing to profile the entire mtDNA of 19 tissue specimens (primary tumor and/or metastatic sites, and tumor-adjacent normal tissue) and 9 cfDNA samples, originating from 8 cancer patients (5 breast, 3 colon). For each patient, tumor-specific mtDNA variants were detected and traced in cfDNA by SMRT sequencing and/or digital PCR to explore their feasibility as cancer biomarker. As a reference, we measured other blood-circulating biomarkers for these patients, including driver mutations in nuclear-encoded cfDNA and cancer-antigen levels or circulating tumor cells. Four of the 24 (17%) tumor-specific mtDNA variants were detected in cfDNA, however at much lower allele frequencies compared to mutations in nuclear-encoded driver genes in the same samples. Also, extensive heterogeneity was observed among the heteroplasmic mtDNA variants present in an individual. We conclude that there is limited value in tracing tumor-specific mtDNA variants in blood-circulating cfDNA with the current methods available. Copyright © 2018 The Authors. Published by Elsevier Inc. All rights reserved.


September 22, 2019

A molecular window into the biology and epidemiology of Pneumocystis spp.

Pneumocystis, a unique atypical fungus with an elusive lifestyle, has had an important medical history. It came to prominence as an opportunistic pathogen that not only can cause life-threatening pneumonia in patients with HIV infection and other immunodeficiencies but also can colonize the lungs of healthy individuals from a very early age. The genus Pneumocystis includes a group of closely related but heterogeneous organisms that have a worldwide distribution, have been detected in multiple mammalian species, are highly host species specific, inhabit the lungs almost exclusively, and have never convincingly been cultured in vitro, making Pneumocystis a fascinating but difficult-to-study organism. Improved molecular biologic methodologies have opened a new window into the biology and epidemiology of Pneumocystis. Advances include an improved taxonomic classification, identification of an extremely reduced genome and concomitant inability to metabolize and grow independent of the host lungs, insights into its transmission mode, recognition of its widespread colonization in both immunocompetent and immunodeficient hosts, and utilization of strain variation to study drug resistance, epidemiology, and outbreaks of infection among transplant patients. This review summarizes these advances and also identifies some major questions and challenges that need to be addressed to better understand Pneumocystis biology and its relevance to clinical care. Copyright © 2018 American Society for Microbiology.


September 22, 2019

A graph-based approach to diploid genome assembly.

Constructing high-quality haplotype-resolved de novo assemblies of diploid genomes is important for revealing the full extent of structural variation and its role in health and disease. Current assembly approaches often collapse the two sequences into one haploid consensus sequence and, therefore, fail to capture the diploid nature of the organism under study. Thus, building an assembler capable of producing accurate and complete diploid assemblies, while being resource-efficient with respect to sequencing costs, is a key challenge to be addressed by the bioinformatics community.We present a novel graph-based approach to diploid assembly, which combines accurate Illumina data and long-read Pacific Biosciences (PacBio) data. We demonstrate the effectiveness of our method on a pseudo-diploid yeast genome and show that we require as little as 50× coverage Illumina data and 10× PacBio data to generate accurate and complete assemblies. Additionally, we show that our approach has the ability to detect and phase structural variants.https://github.com/whatshap/whatshap.Supplementary data are available at Bioinformatics online.


September 22, 2019

Genome analysis of the ancient tracheophyte Selaginella tamariscina reveals evolutionary features relevant to the acquisition of desiccation tolerance.

Resurrection plants, which are the “gifts” of natural evolution, are ideal models for studying the genetic basis of plant desiccation tolerance. Here, we report a high-quality genome assembly of 301 Mb for the diploid spike moss Selaginella tamariscina, a primitive vascular resurrection plant. We predicated 27 761 protein-coding genes from the assembled S. tamariscina genome, 11.38% (2363) of which showed significant expression changes in response to desiccation. Approximately 60.58% of the S. tamariscina genome was annotated as repetitive DNA, which is an almost 2-fold increase of that in the genome of desiccation-sensitive Selaginella moellendorffii. Genomic and transcriptomic analyses highlight the unique evolution and complex regulations of the desiccation response in S. tamariscina, including species-specific expansion of the oleosin and pentatricopeptide repeat gene families, unique genes and pathways for reactive oxygen species generation and scavenging, and enhanced abscisic acid (ABA) biosynthesis and potentially distinct regulation of ABA signaling and response. Comparative analysis of chloroplast genomes of several Selaginella species revealed a unique structural rearrangement and the complete loss of chloroplast NAD(P)H dehydrogenase (NDH) genes in S. tamariscina, suggesting a link between the absence of the NDH complex and desiccation tolerance. Taken together, our comparative genomic and transcriptomic analyses reveal common and species-specific desiccation tolerance strategies in S. tamariscina, providing significant insights into the desiccation tolerance mechanism and the evolution of resurrection plants. Copyright © 2018 The Author. Published by Elsevier Inc. All rights reserved.


September 22, 2019

Heterogeneous and flexible transmission of mcr-1 in hospital-associated Escherichia coli.

The recent emergence of a transferable colistin resistance mechanism, MCR-1, has gained global attention because of its threat to clinical treatment of infections caused by multidrug-resistant Gram-negative bacteria. However, the possible transmission route of mcr-1 among Enterobacteriaceae species in clinical settings is largely unknown. Here, we present a comprehensive genomic analysis of Escherichia coli isolates collected in a hospital in Hangzhou, China. We found that mcr-1-carrying isolates from clinical infections and feces of inpatients and healthy volunteers were genetically diverse and were not closely related phylogenetically, suggesting that clonal expansion is not involved in the spread of mcr-1 The mcr-1 gene was found on either chromosomes or plasmids, but in most of the E. coli isolates, mcr-1 was carried on plasmids. The genetic context of the plasmids showed considerable diversity as evidenced by the different functional insertion sequence (IS) elements, toxin-antitoxin (TA) systems, heavy metal resistance determinants, and Rep proteins of broad-host-range plasmids. Additionally, the genomic analysis revealed nosocomial transmission of mcr-1 and the coexistence of mcr-1 with other genes encoding ß-lactamases and fluoroquinolone resistance in the E. coli isolates. These findings indicate that mcr-1 is heterogeneously disseminated in both commensal and pathogenic strains of E. coli, suggest the high flexibility of this gene in its association with diverse genetic backgrounds of the hosts, and provide new insights into the genome epidemiology of mcr-1 among hospital-associated E. coli strains. IMPORTANCE Colistin represents one of the very few available drugs for treating infections caused by extensively multidrug-resistant Gram-negative bacteria. The recently emergent mcr-1 colistin resistance gene threatens the clinical utility of colistin and has gained global attention. How mcr-1 spreads in hospital settings remains unknown and was investigated by whole-genome sequencing of mcr-1-carrying Escherichia coli in this study. The findings revealed extraordinary flexibility of mcr-1 in its spread among genetically diverse E. coli hosts and plasmids, nosocomial transmission of mcr-1-carrying E. coli, and the continuous emergence of novel Inc types of plasmids carrying mcr-1 and new mcr-1 variants. Additionally, mcr-1 was found to be frequently associated with other genes encoding ß-lactams and fluoroquinolone resistance. These findings provide important information on the transmission and epidemiology of mcr-1 and are of significant public health importance as the information is expected to facilitate the control of this significant antibiotic resistance threat. Copyright © 2018 Shen et al.


September 22, 2019

A rapid method for directed gene knockout for screening in G0 zebrafish.

Zebrafish is a powerful model for forward genetics. Reverse genetic approaches are limited by the time required to generate stable mutant lines. We describe a system for gene knockout that consistently produces null phenotypes in G0 zebrafish. Yolk injection of sets of four CRISPR/Cas9 ribonucleoprotein complexes redundantly targeting a single gene recapitulated germline-transmitted knockout phenotypes in >90% of G0 embryos for each of 8 test genes. Early embryonic (6 hpf) and stable adult phenotypes were produced. Simultaneous multi-gene knockout was feasible but associated with toxicity in some cases. To facilitate use, we generated a lookup table of four-guide sets for 21,386 zebrafish genes and validated several. Using this resource, we targeted 50 cardiomyocyte transcriptional regulators and uncovered a role of zbtb16a in cardiac development. This system provides a platform for rapid screening of genes of interest in development, physiology, and disease models in zebrafish. Copyright © 2018 Elsevier Inc. All rights reserved.


September 22, 2019

Pol V-mediated translesion synthesis elicits localized untargeted mutagenesis during post-replicative gap repair.

In vivo, replication forks proceed beyond replication-blocking lesions by way of downstream repriming, generating daughter strand gaps that are subsequently processed by post-replicative repair pathways such as homologous recombination and translesion synthesis (TLS). The way these gaps are filled during TLS is presently unknown. The structure of gap repair synthesis was assessed by sequencing large collections of single DNA molecules that underwent specific TLS events in vivo. The higher error frequency of specialized relative to replicative polymerases allowed us to visualize gap-filling events at high resolution. Unexpectedly, the data reveal that a specialized polymerase, Pol V, synthesizes stretches of DNA both upstream and downstream of a site-specific DNA lesion. Pol V-mediated untargeted mutations are thus spread over several hundred nucleotides, strongly eliciting genetic instability on either side of a given lesion. Consequently, post-replicative gap repair may be a source of untargeted mutations critical for gene diversification in adaptation and evolution. Copyright © 2018 The Authors. Published by Elsevier Inc. All rights reserved.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.