Menu
September 22, 2019  |  

Comparison of phasing strategies for whole human genomes.

Humans are a diploid species that inherit one set of chromosomes paternally and one homologous set of chromosomes maternally. Unfortunately, most human sequencing initiatives ignore this fact in that they do not directly delineate the nucleotide content of the maternal and paternal copies of the 23 chromosomes individuals possess (i.e., they do not ‘phase’ the genome) often because of the costs and complexities of doing so. We compared 11 different widely-used approaches to phasing human genomes using the publicly available ‘Genome-In-A-Bottle’ (GIAB) phased version of the NA12878 genome as a gold standard. The phasing strategies we compared included laboratory-based assays that prepare DNA in unique ways to facilitate phasing as well as purely computational approaches that seek to reconstruct phase information from general sequencing reads and constructs or population-level haplotype frequency information obtained through a reference panel of haplotypes. To assess the performance of the 11 approaches, we used metrics that included, among others, switch error rates, haplotype block lengths, the proportion of fully phase-resolved genes, phasing accuracy and yield between pairs of SNVs. Our comparisons suggest that a hybrid or combined approach that leverages: 1. population-based phasing using the SHAPEIT software suite, 2. either genome-wide sequencing read data or parental genotypes, and 3. a large reference panel of variant and haplotype frequencies, provides a fast and efficient way to produce highly accurate phase-resolved individual human genomes. We found that for population-based approaches, phasing performance is enhanced with the addition of genome-wide read data; e.g., whole genome shotgun and/or RNA sequencing reads. Further, we found that the inclusion of parental genotype data within a population-based phasing strategy can provide as much as a ten-fold reduction in phasing errors. We also considered a majority voting scheme for the construction of a consensus haplotype combining multiple predictions for enhanced performance and site coverage. Finally, we also identified DNA sequence signatures associated with the genomic regions harboring phasing switch errors, which included regions of low polymorphism or SNV density.


September 22, 2019  |  

Discovery of gorilla MHC-C expressing C1 ligand for KIR.

In comparison to humans and chimpanzees, gorillas show low diversity at MHC class I genes (Gogo), as reflected by an overall reduced level of allelic variation as well as the absence of a functionally important sequence motif that interacts with killer cell immunoglobulin-like receptors (KIR). Here, we use recently generated large-scale genomic sequence data for a reassessment of allelic diversity at Gogo-C, the gorilla orthologue of HLA-C. Through the combination of long-range amplifications and long-read sequencing technology, we obtained, among the 35 gorillas reanalyzed, three novel full-length genomic sequences including a coding region sequence that has not been previously described. The newly identified Gogo-C*03:01 allele has a divergent recombinant structure that sets it apart from other Gogo-C alleles. Domain-by-domain phylogenetic analysis shows that Gogo-C*03:01 has segments in common with Gogo-B*07, the additional B-like gene that is present on some gorilla MHC haplotypes. Identified in ~ 50% of the gorillas analyzed, the Gogo-C*03:01 allele exclusively encodes the C1 epitope among Gogo-C allotypes, indicating its important function in controlling natural killer cell (NK cell) responses via KIR. We further explored the hypothesis whether gorillas experienced a selective sweep which may have resulted in a general reduction of the gorilla MHC class I repertoire. Our results provide little support for a selective sweep but rather suggest that the overall low Gogo class I diversity can be best explained by drastic demographic changes gorillas experienced in the ancient and recent past.


September 22, 2019  |  

Investigating the central metabolism of Clostridium thermosuccinogenes.

Clostridium thermosuccinogenes is a thermophilic anaerobic bacterium able to convert various carbohydrates to succinate and acetate as main fermentation products. Genomes of the four publicly available strains have been sequenced, and the genome of the type strain has been closed. The annotated genomes were used to reconstruct the central metabolism, and enzyme assays were used to validate annotations and to determine cofactor specificity. The genes were identified for the pathways to all fermentation products, as well as for the Embden-Meyerhof-Parnas pathway and the pentose phosphate pathway. Notably, a candidate transaldolase was lacking, and transcriptomics during growth on glucose versus that on xylose did not provide any leads to potential transaldolase genes or alternative pathways connecting the C5 with the C3/C6 metabolism. Enzyme assays showed xylulokinase to prefer GTP over ATP, which could be of importance for engineering xylose utilization in related thermophilic species of industrial relevance. Furthermore, the gene responsible for malate dehydrogenase was identified via heterologous expression in Escherichia coli and subsequent assays with the cell extract, which has proven to be a simple and powerful method for the basal characterization of thermophilic enzymes.IMPORTANCE Running industrial fermentation processes at elevated temperatures has several advantages, including reduced cooling requirements, increased reaction rates and solubilities, and a possibility to perform simultaneous saccharification and fermentation of a pretreated biomass. Most studies with thermophiles so far have focused on bioethanol production. Clostridium thermosuccinogenes seems an attractive production organism for organic acids, succinic acid in particular, from lignocellulosic biomass-derived sugars. This study provides valuable insights into its central metabolism and GTP and PPi cofactor utilization. Copyright © 2018 American Society for Microbiology.


September 22, 2019  |  

A graph-based approach to diploid genome assembly.

Constructing high-quality haplotype-resolved de novo assemblies of diploid genomes is important for revealing the full extent of structural variation and its role in health and disease. Current assembly approaches often collapse the two sequences into one haploid consensus sequence and, therefore, fail to capture the diploid nature of the organism under study. Thus, building an assembler capable of producing accurate and complete diploid assemblies, while being resource-efficient with respect to sequencing costs, is a key challenge to be addressed by the bioinformatics community.We present a novel graph-based approach to diploid assembly, which combines accurate Illumina data and long-read Pacific Biosciences (PacBio) data. We demonstrate the effectiveness of our method on a pseudo-diploid yeast genome and show that we require as little as 50× coverage Illumina data and 10× PacBio data to generate accurate and complete assemblies. Additionally, we show that our approach has the ability to detect and phase structural variants.https://github.com/whatshap/whatshap.Supplementary data are available at Bioinformatics online.


September 22, 2019  |  

A rapid method for directed gene knockout for screening in G0 zebrafish.

Zebrafish is a powerful model for forward genetics. Reverse genetic approaches are limited by the time required to generate stable mutant lines. We describe a system for gene knockout that consistently produces null phenotypes in G0 zebrafish. Yolk injection of sets of four CRISPR/Cas9 ribonucleoprotein complexes redundantly targeting a single gene recapitulated germline-transmitted knockout phenotypes in >90% of G0 embryos for each of 8 test genes. Early embryonic (6 hpf) and stable adult phenotypes were produced. Simultaneous multi-gene knockout was feasible but associated with toxicity in some cases. To facilitate use, we generated a lookup table of four-guide sets for 21,386 zebrafish genes and validated several. Using this resource, we targeted 50 cardiomyocyte transcriptional regulators and uncovered a role of zbtb16a in cardiac development. This system provides a platform for rapid screening of genes of interest in development, physiology, and disease models in zebrafish. Copyright © 2018 Elsevier Inc. All rights reserved.


September 22, 2019  |  

The Chara genome: Secondary complexity and implications for plant terrestrialization.

Land plants evolved from charophytic algae, among which Charophyceae possess the most complex body plans. We present the genome of Chara braunii; comparison of the genome to those of land plants identified evolutionary novelties for plant terrestrialization and land plant heritage genes. C. braunii employs unique xylan synthases for cell wall biosynthesis, a phragmoplast (cell separation) mechanism similar to that of land plants, and many phytohormones. C. braunii plastids are controlled via land-plant-like retrograde signaling, and transcriptional regulation is more elaborate than in other algae. The morphological complexity of this organism may result from expanded gene families, with three cases of particular note: genes effecting tolerance to reactive oxygen species (ROS), LysM receptor-like kinases, and transcription factors (TFs). Transcriptomic analysis of sexual reproductive structures reveals intricate control by TFs, activity of the ROS gene network, and the ancestral use of plant-like storage and stress protection proteins in the zygote. Copyright © 2018 Elsevier Inc. All rights reserved.


September 22, 2019  |  

Identification of the DNA methyltransferases establishing the methylome of the cyanobacterium Synechocystis sp. PCC 6803.

DNA methylation in bacteria is important for defense against foreign DNA, but is also involved in DNA repair, replication, chromosome partitioning, and regulatory processes. Thus, characterization of the underlying DNA methyltransferases in genetically tractable bacteria is of paramount importance. Here, we characterized the methylome and orphan methyltransferases in the model cyanobacterium Synechocystis sp. PCC 6803. Single molecule real-time (SMRT) sequencing revealed four DNA methylation recognition sequences in addition to the previously known motif m5CGATCG, which is recognized by M.Ssp6803I. For three of the new recognition sequences, we identified the responsible methyltransferases. M.Ssp6803II, encoded by the sll0729 gene, modifies GGm4CC, M.Ssp6803III, encoded by slr1803, represents the cyanobacterial dam-like methyltransferase modifying Gm6ATC, and M.Ssp6803V, encoded by slr6095 on plasmid pSYSX, transfers methyl groups to the bipartite motif GGm6AN7TTGG/CCAm6AN7TCC. The remaining methylation recognition sequence GAm6AGGC is probably recognized by methyltransferase M.Ssp6803IV encoded by slr6050. M.Ssp6803III and M.Ssp6803IV were essential for the viability of Synechocystis, while the strains lacking M.Ssp6803I and M.Ssp6803V showed growth similar to the wild type. In contrast, growth was strongly diminished of the ?sll0729 mutant lacking M.Ssp6803II. These data provide the basis for systematic studies on the molecular mechanisms impacted by these methyltransferases.


September 22, 2019  |  

The landscape of repetitive elements in the refined genome of chilli anthracnose fungus Colletotrichum truncatum.

The ascomycete fungus Colletotrichum truncatum is a major phytopathogen with a broad host range which causes anthracnose disease of chilli. The genome sequencing of this fungus led to the discovery of functional categories of genes that may play important roles in fungal pathogenicity. However, the presence of gaps in C. truncatum draft assembly prevented the accurate prediction of repetitive elements, which are the key players to determine the genome architecture and drive evolution and host adaptation. We re-sequenced its genome using single-molecule real-time (SMRT) sequencing technology to obtain a refined assembly with lesser and smaller gaps and ambiguities. This enabled us to study its genome architecture by characterising the repetitive sequences like transposable elements (TEs) and simple sequence repeats (SSRs), which constituted 4.9 and 0.38% of the assembled genome, respectively. The comparative analysis among different Colletotrichum species revealed the extensive repeat rich regions, dominated by Gypsy superfamily of long terminal repeats (LTRs), and the differential composition of SSRs in their genomes. Our study revealed a recent burst of LTR amplification in C. truncatum, C. higginsianum, and C. scovillei. TEs in C. truncatum were significantly associated with secretome, effectors and genes in secondary metabolism clusters. Some of the TE families in C. truncatum showed cytosine to thymine transitions indicative of repeat-induced point mutation (RIP). C. orbiculare and C. graminicola showed strong signatures of RIP across their genomes and “two-speed” genomes with extensive AT-rich and gene-sparse regions. Comparative genomic analyses of Colletotrichum species provided an insight into the species-specific SSR profiles. The SSRs in the coding and non-coding regions of the genome revealed the composition of trinucleotide repeat motifs in exons with potential to alter the translated protein structure through amino acid repeats. This is the first genome-wide study of TEs and SSRs in C. truncatum and their comparative analysis with six other Colletotrichum species, which would serve as a useful resource for future research to get insights into the potential role of TEs in genome expansion and evolution of Colletotrichum fungi and for development of SSR-based molecular markers for population genomic studies.


September 22, 2019  |  

Recovery of novel association loci in Arabidopsis thaliana and Drosophila melanogaster through leveraging INDELs association and integrated burden test.

Short insertions, deletions (INDELs) and larger structural variants have been increasingly employed in genetic association studies, but few improvements over SNP-based association have been reported. In order to understand why this might be the case, we analysed two publicly available datasets and observed that 63% of INDELs called in A. thaliana and 64% in D. melanogaster populations are misrepresented as multiple alleles with different functional annotations, i.e. where the same underlying variant is represented by inconsistent alignments leading to different variant calls. To address this issue, we have developed the software Irisas to reclassify and re-annotate these variants, which we then used for single-locus tests of association. We also integrated them to predict the functional impact of SNPs, INDELs, and structural variants for burden testing. Using both approaches, we re-analysed the genetic architecture of complex traits in A. thaliana and D. melanogaster. Heritability analysis using SNPs alone explained on average 27% and 19% of phenotypic variance for A. thaliana and D. melanogaster respectively. Our method explained an additional 11% and 3%, respectively. We also identified novel trait loci that previous SNP-based association studies failed to map, and which contain established candidate genes. Our study shows the value of the association test with INDELs and integrating multiple types of variants in association studies in plants and animals.


September 22, 2019  |  

The genomic basis of color pattern polymorphism in the Harlequin ladybird.

Many animal species comprise discrete phenotypic forms. A common example in natural populations of insects is the occurrence of different color patterns, which has motivated a rich body of ecological and genetic research [1-6]. The occurrence of dark, i.e., melanic, forms displaying discrete color patterns is found across multiple taxa, but the underlying genomic basis remains poorly characterized. In numerous ladybird species (Coccinellidae), the spatial arrangement of black and red patches on adult elytra varies wildly within species, forming strikingly different complex color patterns [7, 8]. In the harlequin ladybird, Harmonia axyridis, more than 200 distinct color forms have been described, which classic genetic studies suggest result from allelic variation at a single, unknown, locus [9, 10]. Here, we combined whole-genome sequencing, population-based genome-wide association studies, gene expression, and functional analyses to establish that the transcription factor Pannier controls melanic pattern polymorphism in H. axyridis. We show that pannier is necessary for the formation of melanic elements on the elytra. Allelic variation in pannier leads to protein expression in distinct domains on the elytra and thus determines the distinct color patterns in H. axyridis. Recombination between pannier alleles may be reduced by a highly divergent sequence of ~170 kb in the cis-regulatory regions of pannier, with a 50 kb inversion between color forms. This most likely helps maintain the distinct alleles found in natural populations. Thus, we propose that highly variable discrete color forms can arise in natural populations through cis-regulatory allelic variation of a single gene. Copyright © 2018 The Authors. Published by Elsevier Ltd.. All rights reserved.


September 22, 2019  |  

A complete Cannabis chromosome assembly and adaptive admixture for elevated cannabidiol (CBD) content

Cannabis has been cultivated for millennia with distinct cultivars providing either fiber and grain or tetrahydrocannabinol. Recent demand for cannabidiol rather than tetrahydrocannabinol has favored the breeding of admixed cultivars with extremely high cannabidiol content. Despite several draft Cannabis genomes, the genomic structure of cannabinoid synthase loci has remained elusive. A genetic map derived from a tetrahydrocannabinol/cannabidiol segregating population and a complete chromosome assembly from a high-cannabidiol cultivar together resolve the linkage of cannabidiolic and tetrahydrocannabinolic acid synthase gene clusters which are associated with transposable elements. High-cannabidiol cultivars appear to have been generated by integrating hemp-type cannabidiolic acid synthase gene clusters into a background of marijuana-type cannabis. Quantitative trait locus mapping suggests that overall drug potency, however, is associated with other genomic regions needing additional study.


September 22, 2019  |  

pYR4 from a Norwegian isolate of Yersinia ruckeri is a putative virulence plasmid encoding both a type IV pilus and a type IV secretion system

Enteric redmouth disease caused by the pathogen Yersinia ruckeri is a significant problem for fish farming around the world. Despite its importance, only a few virulence factors of Y. ruckeri have been identified and studied in detail. Here, we report and analyze the complete DNA sequence of pYR4, a plasmid from a highly pathogenic Norwegian Y. ruckeri isolate, sequenced using PacBio SMRT technology. Like the well-known pYV plasmid of human pathogenic Yersiniae, pYR4 is a member of the IncFII family. Thirty-one percent of the pYR4 sequence is unique compared to other Y. ruckeri plasmids. The unique regions contain, among others genes, a large number of mobile genetic elements and two partitioning systems. The G+C content of pYR4 is higher than that of the Y. ruckeri NVH_3758 genome, indicating its relatively recent horizontal acquisition. pYR4, as well as the related plasmid pYR3, comprises operons that encode for type IV pili and for a conjugation system (tra). In contrast to other Yersinia plasmids, pYR4 cannot be cured at elevated temperatures. Our study highlights the power of PacBio sequencing technology for identifying mis-assembled segments of genomic sequences. Comparative analysis of pYR4 and other Y. ruckeri plasmids and genomes, which were sequenced by second and the third generation sequencing technologies, showed errors in second generation sequencing assemblies. Specifically, in the Y. ruckeri 150 and Y. ruckeri ATCC29473 genome assemblies, we mapped the entire pYR3 plasmid sequence. Placing plasmid sequences on the chromosome can result in erroneous biological conclusions. Thus, PacBio sequencing or similar long-read methods should always be preferred for de novo genome sequencing. As the tra operons of pYR3, although misplaced on the chromosome during the genome assembly process, were demonstrated to have an effect on virulence, and type IV pili are virulence factors in many bacteria, we suggest that pYR4 directly contributes to Y. ruckeri virulence.


September 22, 2019  |  

Alpha- and beta-mannan utilization by marine Bacteroidetes.

Marine microscopic algae carry out about half of the global carbon dioxide fixation into organic matter. They provide organic substrates for marine microbes such as members of the Bacteroidetes that degrade algal polysaccharides using carbohydrate-active enzymes (CAZymes). In Bacteroidetes genomes CAZyme encoding genes are mostly grouped in distinct regions termed polysaccharide utilization loci (PULs). While some studies have shown involvement of PULs in the degradation of algal polysaccharides, the specific substrates are for the most part still unknown. We investigated four marine Bacteroidetes isolated from the southern North Sea that harbour putative mannan-specific PULs. These PULs are similarly organized as PULs in human gut Bacteroides that digest a- and ß-mannans from yeasts and plants respectively. Using proteomics and defined growth experiments with polysaccharides as sole carbon sources we could show that the investigated marine Bacteroidetes express the predicted functional proteins required for a- and ß-mannan degradation. Our data suggest that algal mannans play an as yet unknown important role in the marine carbon cycle, and that biochemical principles established for gut or terrestrial microbes also apply to marine bacteria, even though their PULs are evolutionarily distant.© 2018 The Authors. Environmental Microbiology published by Society for Applied Microbiology and John Wiley & Sons Ltd.


September 22, 2019  |  

Phenotypic and genomic comparison of Photorhabdus luminescens subsp. laumondii TT01 and a widely used rifampicin-resistant Photorhabdus luminescens laboratory strain.

Photorhabdus luminescens is an enteric bacterium, which lives in mutualistic association with soil nematodes and is highly pathogenic for a broad spectrum of insects. A complete genome sequence for the type strain P. luminescens subsp. laumondii TT01, which was originally isolated in Trinidad and Tobago, has been described earlier. Subsequently, a rifampicin resistant P. luminescens strain has been generated with superior possibilities for experimental characterization. This strain, which is widely used in research, was described as a spontaneous rifampicin resistant mutant of TT01 and is known as TT01-RifR.Unexpectedly, upon phenotypic comparison between the rifampicin resistant strain and its presumed parent TT01, major differences were found with respect to bioluminescence, pigmentation, biofilm formation, haemolysis as well as growth. Therefore, we renamed the strain TT01-RifR to DJC. To unravel the genomic basis of the observed differences, we generated a complete genome sequence for strain DJC using the PacBio long read technology. As strain DJC was supposed to be a spontaneous mutant, only few sequence differences were expected. In order to distinguish these from potential sequencing errors in the published TT01 genome, we re-sequenced a derivative of strain TT01 in parallel, also using the PacBio technology. The two TT01 genomes differed at only 30 positions. In contrast, the genome of strain DJC varied extensively from TT01, showing 13,000 point mutations, 330 frameshifts, and 220 strain-specific regions with a total length of more than 300 kb in each of the compared genomes.According to the major phenotypic and genotypic differences, the rifampicin resistant P. luminescens strain, now named strain DJC, has to be considered as an independent isolate rather than a derivative of strain TT01. Strains TT01 and DJC both belong to P. luminescens subsp. laumondii.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.