Menu
September 22, 2019

Impact of index hopping and bias towards the reference allele on accuracy of genotype calls from low-coverage sequencing.

Inherent sources of error and bias that affect the quality of sequence data include index hopping and bias towards the reference allele. The impact of these artefacts is likely greater for low-coverage data than for high-coverage data because low-coverage data has scant information and many standard tools for processing sequence data were designed for high-coverage data. With the proliferation of cost-effective low-coverage sequencing, there is a need to understand the impact of these errors and bias on resulting genotype calls from low-coverage sequencing.We used a dataset of 26 pigs sequenced both at 2× with multiplexing and at 30× without multiplexing to show that index hopping and bias towards the reference allele due to alignment had little impact on genotype calls. However, pruning of alternative haplotypes supported by a number of reads below a predefined threshold, which is a default and desired step of some variant callers for removing potential sequencing errors in high-coverage data, introduced an unexpected bias towards the reference allele when applied to low-coverage sequence data. This bias reduced best-guess genotype concordance of low-coverage sequence data by 19.0 absolute percentage points.We propose a simple pipeline to correct the preferential bias towards the reference allele that can occur during variant discovery and we recommend that users of low-coverage sequence data be wary of unexpected biases that may be produced by bioinformatic tools that were designed for high-coverage sequence data.


September 22, 2019

The Genome of Opium Poppy Reveals Evolutionary History of Morphinan Pathway.

Plants, as primary producers, have been playing an indispensable role in other organisms’ survival and the balance of whole ecosystem on Earth. Especially, they provide the main source of energy, food, and medicine for human beings, some of which are derived from the primary or secondary metabolites [1]. Angiosperms, with more than 300,000 species on Earth, are the largest group of land plants by far. Most agricultural crops, fruits, ornamental plants, and medicinal herbs belong to this group. The medicinal herbs are usually rich in specialized metabolites that could provide safe and valuable resources for pharmaceutical development.


September 22, 2019

Whole-genome landscape of Medicago truncatula symbiotic genes.

Advances in deciphering the functional architecture of eukaryotic genomes have been facilitated by recent breakthroughs in sequencing technologies, enabling a more comprehensive representation of genes and repeat elements in genome sequence assemblies, as well as more sensitive and tissue-specific analyses of gene expression. Here we show that PacBio sequencing has led to a substantially improved genome assembly of Medicago truncatula A17, a legume model species notable for endosymbiosis studies1, and has enabled the identification of genome rearrangements between genotypes at a near-base-pair resolution. Annotation of the new M. truncatula genome sequence has allowed for a thorough analysis of transposable elements and their dynamics, as well as the identification of new players involved in symbiotic nodule development, in particular 1,037 upregulated long non-coding RNAs (lncRNAs). We have also discovered that a substantial proportion (~35% and 38%, respectively) of the genes upregulated in nodules or expressed in the nodule differentiation zone colocalize in genomic clusters (270 and 211, respectively), here termed symbiotic islands. These islands contain numerous expressed lncRNA genes and display differentially both DNA methylation and histone marks. Epigenetic regulations and lncRNAs are therefore attractive candidate elements for the orchestration of symbiotic gene expression in the M. truncatula genome.


September 22, 2019

Density-dependent enhanced replication of a densovirus in Wolbachia-infected Aedes cells is associated with production of piRNAs and higher virus-derived siRNAs.

The endosymbiotic bacterium Wolbachia pipientis has been shown to restrict a range of RNA viruses in Drosophila melanogaster and transinfected dengue mosquito, Aedes aegypti. Here, we show that Wolbachia infection enhances replication of Aedes albopictus densovirus (AalDNV-1), a single stranded DNA virus, in Aedes cell lines in a density-dependent manner. Analysis of previously produced small RNAs of Aag2 cells showed that Wolbachia-infected cells produced greater absolute abundance of virus-derived short interfering RNAs compared to uninfected cells. Additionally, we found production of virus-derived PIWI-like RNAs (vpiRNA) produced in response to AalDNV-1 infection. Nuclear fractions of Aag2 cells produced a primary vpiRNA signature U1 bias whereas the typical “ping-pong” signature (U1 – A10) was evident in vpiRNAs from the cytoplasmic fractions. This is the first report of the density-dependent enhancement of DNA viruses by Wolbachia. Further, we report the generation of vpiRNAs in a DNA virus-host interaction for the first time. Copyright © 2018 Elsevier Inc. All rights reserved.


September 22, 2019

Desiccation Tolerance Evolved through Gene Duplication and Network Rewiring in Lindernia.

Although several resurrection plant genomes have been sequenced, the lack of suitable dehydration-sensitive outgroups has limited genomic insights into the origin of desiccation tolerance. Here, we utilized a comparative system of closely related desiccation-tolerant (Lindernia brevidens) and -sensitive (Lindernia subracemosa) species to identify gene- and pathway-level changes associated with the evolution of desiccation tolerance. The two high-quality Lindernia genomes we assembled are largely collinear, and over 90% of genes are conserved. L. brevidens and L. subracemosa have evidence of an ancient, shared whole-genome duplication event, and retained genes have neofunctionalized, with desiccation-specific expression in L. brevidens Tandem gene duplicates also are enriched in desiccation-associated functions, including a dramatic expansion of early light-induced proteins from 4 to 26 copies in L. brevidens A comparative differential gene coexpression analysis between L. brevidens and L. subracemosa supports extensive network rewiring across early dehydration, desiccation, and rehydration time courses. Many LATE EMBRYOGENESIS ABUNDANT genes show significantly higher expression in L. brevidens compared with their orthologs in L. subracemosa Coexpression modules uniquely upregulated during desiccation in L. brevidens are enriched with seed-specific and abscisic acid-associated cis-regulatory elements. These modules contain a wide array of seed-associated genes that have no expression in the desiccation-sensitive L. subracemosa Together, these findings suggest that desiccation tolerance evolved through a combination of gene duplications and network-level rewiring of existing seed desiccation pathways.© 2018 American Society of Plant Biologists. All rights reserved.


September 22, 2019

The genome of the tegu lizard Salvator merianae: combining Illumina, PacBio, and optical mapping data to generate a highly contiguous assembly.

Reptiles are a species-rich group with great phenotypic and life history diversity but are highly underrepresented among the vertebrate species with sequenced genomes.Here, we report a high-quality genome assembly of the tegu lizard, Salvator merianae, the first lacertoid with a sequenced genome. We combined 74X Illumina short-read, 29.8X Pacific Biosciences long-read, and optical mapping data to generate a high-quality assembly with a scaffold N50 value of 55.4 Mb. The contig N50 value of this assembly is 521 Kb, making it the most contiguous reptile assembly so far. We show that the tegu assembly has the highest completeness of coding genes and conserved non-exonic elements (CNEs) compared to other reptiles. Furthermore, the tegu assembly has the highest number of evolutionarily conserved CNE pairs, corroborating a high assembly contiguity in intergenic regions. As in other reptiles, long interspersed nuclear elements comprise the most abundant transposon class. We used transcriptomic data, homology- and de novo gene predictions to annotate 22,413 coding genes, of which 16,995 (76%) likely have human orthologs as inferred by CESAR-derived gene mappings. Finally, we generated a multiple genome alignment comprising 10 squamates and 7 other amniote species and identified conserved regions that are under evolutionary constraint. CNEs cover 38 Mb (1.8%) of the tegu genome, with 3.3 Mb in these elements being squamate specific. In contrast to placental mammal-specific CNEs, very few of these squamate-specific CNEs (<20 Kb) overlap transposons, highlighting a difference in how lineage-specific CNEs originated in these two clades.The tegu lizard genome together with the multiple genome alignment and comprehensive conserved element datasets provide a valuable resource for comparative genomic studies of reptiles and other amniotes.


September 22, 2019

The genomic landscape of molecular responses to natural drought stress in Panicum hallii

Environmental stress is a major driver of ecological community dynamics and agricultural productivity. This is especially true for soil water availability, because drought is the greatest abiotic inhibitor of worldwide crop yields. Here, we test the genetic basis of drought responses in the genetic model for C4perennial grasses, Panicum hallii, through population genomics, field-scale gene-expression (eQTL) analysis, and comparison of two complete genomes. While gene expression networks are dominated by local cis-regulatory elements, we observe three genomic hotspots of unlinked trans-regulatory loci. These regulatory hubs are four times more drought responsive than the genome-wide average. Additionally, cis- and trans-regulatory networks are more likely to have opposing effects than expected under neutral evolution, supporting a strong influence of compensatory evolution and stabilizing selection. These results implicate trans-regulatory evolution as a driver of drought responses and demonstrate the potential for crop improvement in drought-prone regions through modification of gene regulatory networks.


September 22, 2019

Evolution of host support for two ancient bacterial symbionts with differentially degraded genomes in a leafhopper host.

Plant sap-feeding insects (Hemiptera) rely on bacterial symbionts for nutrition absent in their diets. These bacteria experience extreme genome reduction and require genetic resources from their hosts, particularly for basic cellular processes other than nutrition synthesis. The host-derived mechanisms that complete these processes have remained poorly understood. It is also unclear how hosts meet the distinct needs of multiple bacterial partners with differentially degraded genomes. To address these questions, we investigated the cell-specific gene-expression patterns in the symbiotic organs of the aster leafhopper (ALF), Macrosteles quadrilineatus (Cicadellidae). ALF harbors two intracellular symbionts that have two of the smallest known bacterial genomes: Nasuia (112 kb) and Sulcia (190 kb). Symbionts are segregated into distinct host cell types (bacteriocytes) and vary widely in their basic cellular capabilities. ALF differentially expresses thousands of genes between the bacteriocyte types to meet the functional needs of each symbiont, including the provisioning of metabolites and support of cellular processes. For example, the host highly expresses genes in the bacteriocytes that likely complement gene losses in nucleic acid synthesis, DNA repair mechanisms, transcription, and translation. Such genes are required to function in the bacterial cytosol. Many host genes comprising these support mechanisms are derived from the evolution of novel functional traits via horizontally transferred genes, reassigned mitochondrial support genes, and gene duplications with bacteriocyte-specific expression. Comparison across other hemipteran lineages reveals that hosts generally support the incomplete symbiont cellular processes, but the origins of these support mechanisms are generally specific to the host-symbiont system.Copyright © 2018 the Author(s). Published by PNAS.


September 22, 2019

Genomic characterization of a B chromosome in Lake Malawi cichlid fishes.

B chromosomes (Bs) were discovered a century ago, and since then, most studies have focused on describing their distribution and abundance using traditional cytogenetics. Only recently have attempts been made to understand their structure and evolution at the level of DNA sequence. Many questions regarding the origin, structure, function, and evolution of B chromosomes remain unanswered. Here, we identify B chromosome sequences from several species of cichlid fish from Lake Malawi by examining the ratios of DNA sequence coverage in individuals with or without B chromosomes. We examined the efficiency of this method, and compared results using both Illumina and PacBio sequence data. The B chromosome sequences detected in 13 individuals from 7 species were compared to assess the rates of sequence replacement. B-specific sequence common to at least 12 of the 13 datasets were identified as the “Core” B chromosome. The location of B sequence homologs throughout the genome provides further support for theories of B chromosome evolution. Finally, we identified genes and gene fragments located on the B chromosome, some of which may regulate the segregation and maintenance of the B chromosome.


September 22, 2019

N6-methyladenine DNA methylation in Japonica and Indica rice genomes and its association with gene expression, plant development, and stress responses.

N6-Methyladenine (6mA) DNA methylation has recently been implicated as a potential new epigenetic marker in eukaryotes, including the dicot model Arabidopsis thaliana. However, the conservation and divergence of 6mA distribution patterns and functions in plants remain elusive. Here we report high-quality 6mA methylomes at single-nucleotide resolution in rice based on substantially improved genome sequences of two rice cultivars, Nipponbare (Nip; Japonica) and 93-11 (Indica). Analysis of 6mA genomic distribution and its association with transcription suggest that 6mA distribution and function is rather conserved between rice and Arabidopsis. We found that 6mA levels are positively correlated with the expression of key stress-related genes, which may be responsible for the difference in stress tolerance between Nip and 93-11. Moreover, we showed that mutations in DDM1 cause defects in plant growth and decreased 6mA level. Our results reveal that 6mA is a conserved DNA modification that is positively associated with gene expression and contributes to key agronomic traits in plants. Copyright © 2018 The Author. Published by Elsevier Inc. All rights reserved.


September 22, 2019

Detection and visualization of complex structural variants from long reads.

With applications in cancer, drug metabolism, and disease etiology, understanding structural variation in the human genome is critical in advancing the thrusts of individualized medicine. However, structural variants (SVs) remain challenging to detect with high sensitivity using short read sequencing technologies. This problem is exacerbated when considering complex SVs comprised of multiple overlapping or nested rearrangements. Longer reads, such as those from Pacific Biosciences platforms, often span multiple breakpoints of such events, and thus provide a way to unravel small-scale complexities in SVs with higher confidence.We present CORGi (COmplex Rearrangement detection with Graph-search), a method for the detection and visualization of complex local genomic rearrangements. This method leverages the ability of long reads to span multiple breakpoints to untangle SVs that appear very complicated with respect to a reference genome. We validated our approach against both simulated long reads, and real data from two long read sequencing technologies. We demonstrate the ability of our method to identify breakpoints inserted in synthetic data with high accuracy, and the ability to detect and plot SVs from NA12878 germline, achieving 88.4% concordance between the two sets of sequence data. The patterns of complexity we find in many NA12878 SVs match known mechanisms associated with DNA replication and structural variant formation, and highlight the ability of our method to automatically label complex SVs with an intuitive combination of adjacent or overlapping reference transformations.CORGi is a method for interrogating genomic regions suspected to contain local rearrangements using long reads. Using pairwise alignments and graph search CORGi produces labels and visualizations for local SVs of arbitrary complexity.


September 22, 2019

How resurrection plants survive being hung out to dry.

Resurrection plants have the unique ability to survive extreme dehydration (desiccation), lying dormant for months or sometimes years until rehydration is possible. This formidable survival strategy has independently evolved several times across the land plant phylogeny, and several phylogenetically diverse resurrection plant genomes have been sequenced and assembled in an attempt to understand the causal genetic mechanisms. Large-scale comparisons across each of these phylogenetically distant resurrection plant genomes reveals that some conserved molecular signatures may underlie desiccation tolerance (Illing et al., 2005; Zhang and Bartels, 2018), but overall the genes, networks, and regulatory factors that underlie desiccation tolerance remain largely unknown.


September 22, 2019

Genome of the small hive beetle (Aethina tumida, Coleoptera: Nitidulidae), a worldwide parasite of social bee colonies, provides insights into detoxification and herbivory.

The small hive beetle (Aethina tumida; ATUMI) is an invasive parasite of bee colonies. ATUMI feeds on both fruits and bee nest products, facilitating its spread and increasing its impact on honey bees and other pollinators. We have sequenced and annotated the ATUMI genome, providing the first genomic resources for this species and for the Nitidulidae, a beetle family that is closely related to the extraordinarily species-rich clade of beetles known as the Phytophaga. ATUMI thus provides a contrasting view as a neighbor for one of the most successful known animal groups.We present a robust genome assembly and a gene set possessing 97.5% of the core proteins known from the holometabolous insects. The ATUMI genome encodes fewer enzymes for plant digestion than the genomes of wood-feeding beetles but nonetheless shows signs of broad metabolic plasticity. Gustatory receptors are few in number compared to other beetles, especially receptors with known sensitivity (in other beetles) to bitter substances. In contrast, several gene families implicated in detoxification of insecticides and adaptation to diverse dietary resources show increased copy numbers. The presence and diversity of homologs involved in detoxification differ substantially from the bee hosts of ATUMI.Our results provide new insights into the genomic basis for local adaption and invasiveness in ATUMI and a blueprint for control strategies that target this pest without harming their honey bee hosts. A minimal set of gustatory receptors is consistent with the observation that, once a host colony is invaded, food resources are predictable. Unique detoxification pathways and pathway members can help identify which treatments might control this species even in the presence of honey bees, which are notoriously sensitive to pesticides.


September 22, 2019

Approaches for surveying cosmic radiation damage in large populations of Arabidopsis thaliana seeds-Antarctic balloons and particle beams.

The Cosmic Ray Exposure Sequencing Science (CRESS) payload system is a proof of concept experiment to assess the genomic impact of space radiation on seeds. CRESS was designed as a secondary payload for the December 2016 high-altitude, high-latitude, and long-duration balloon flight carrying the Boron And Carbon Cosmic Rays in the Upper Stratosphere (BACCUS) experimental hardware. Investigation of the biological effects of Galactic Cosmic Radiation (GCR), particularly those of ions with High-Z and Energy (HZE), is of interest due to the genomic damage this type of radiation inflicts. The biological effects of upper-stratospheric mixed radiation above Antarctica (ANT) were sampled using Arabidopsis thaliana seeds and were compared to those resulting from a controlled simulation of GCR at Brookhaven National Laboratory (BNL) and to laboratory control seed. The payload developed for Antarctica exposure was broadly designed to 1U CubeSat specifications (10cmx10cmx10cm, =1.33kg), maintained 1 atm internal pressure, and carried an internal cargo of four seed trays (about 580,000 seeds) and twelve CR-39 Solid-State Nuclear Track Detectors (SSNTDs). The irradiated seeds were recovered, sterilized and grown on Petri plates for phenotypic screening. BNL and ANT M0 seeds showed significantly reduced germination rates and elevated somatic mutation rates when compared to non-irradiated controls, with the BNL mutation rate also being significantly higher than that of ANT. Genomic DNA from mutants of interest was evaluated with whole-genome sequencing using PacBio SMRT technology. Sequence data revealed the presence of an array of genome structural variants in the genomes of M0 and M1 mutant plants.


September 22, 2019

Genetics and genomics of an unusual selfish sex ratio distortion in an insect.

Diverse selfish genetic elements have evolved the ability to manipulate reproduction to increase their transmission, and this can result in highly distorted sex ratios [1]. Indeed, one of the major explanations for why sex determination systems are so dynamic is because they are shaped by ongoing coevolutionary arms races between sex-ratio-distorting elements and the rest of the genome [2]. Here, we use genetic crosses and genome analysis to describe an unusual sex ratio distortion with striking consequences on genome organization in a booklouse species, Liposcelis sp. (Insecta: Psocodea), in which two types of females coexist. Distorter females never produce sons but must mate with males (the sons of nondistorting females) to reproduce [3]. Although they are diploid and express the genes inherited from their fathers in somatic tissues, distorter females only ever transmit genes inherited from their mothers. As a result, distorter females have unusual chimeric genomes, with distorter-restricted chromosomes diverging from their nondistorting counterparts and exhibiting features of a giant non-recombining sex chromosome. The distorter-restricted genome has also acquired a gene from the bacterium Wolbachia, a well-known insect reproductive manipulator; we found that this gene has independently colonized the genomes of two other insect species with unusual reproductive systems, suggesting possible roles in sex ratio distortion in this remarkable genetic system. Copyright © 2018 Elsevier Ltd. All rights reserved.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.