Menu
April 21, 2020  |  

Improved annotation of the domestic pig genome through integration of Iso-Seq and RNA-seq data.

Our understanding of the pig transcriptome is limited. RNA transcript diversity among nine tissues was assessed using poly(A) selected single-molecule long-read isoform sequencing (Iso-seq) and Illumina RNA sequencing (RNA-seq) from a single White cross-bred pig. Across tissues, a total of 67,746 unique transcripts were observed, including 60.5% predicted protein-coding, 36.2% long non-coding RNA and 3.3% nonsense-mediated decay transcripts. On average, 90% of the splice junctions were supported by RNA-seq within tissue. A large proportion (80%) represented novel transcripts, mostly produced by known protein-coding genes (70%), while 17% corresponded to novel genes. On average, four transcripts per known gene (tpg) were identified; an increase over current EBI (1.9 tpg) and NCBI (2.9 tpg) annotations and closer to the number reported in human genome (4.2 tpg). Our new pig genome annotation extended more than 6000 known gene borders (5′ end extension, 3′ end extension, or both) compared to EBI or NCBI annotations. We validated a large proportion of these extensions by independent pig poly(A) selected 3′-RNA-seq data, or human FANTOM5 Cap Analysis of Gene Expression data. Further, we detected 10,465 novel genes (81% non-coding) not reported in current pig genome annotations. More than 80% of these novel genes had transcripts detected in >?1 tissue. In addition, more than 80% of novel intergenic genes with at least one transcript detected in liver tissue had H3K4me3 or H3K36me3 peaks mapping to their promoter and gene body, respectively, in independent liver chromatin immunoprecipitation data. These validated results show significant improvement over current pig genome annotations.


April 21, 2020  |  

Systematic analysis of dark and camouflaged genes reveals disease-relevant genes hiding in plain sight.

The human genome contains “dark” gene regions that cannot be adequately assembled or aligned using standard short-read sequencing technologies, preventing researchers from identifying mutations within these gene regions that may be relevant to human disease. Here, we identify regions with few mappable reads that we call dark by depth, and others that have ambiguous alignment, called camouflaged. We assess how well long-read or linked-read technologies resolve these regions.Based on standard whole-genome Illumina sequencing data, we identify 36,794 dark regions in 6054 gene bodies from pathways important to human health, development, and reproduction. Of these gene bodies, 8.7% are completely dark and 35.2% are =?5% dark. We identify dark regions that are present in protein-coding exons across 748 genes. Linked-read or long-read sequencing technologies from 10x Genomics, PacBio, and Oxford Nanopore Technologies reduce dark protein-coding regions to approximately 50.5%, 35.6%, and 9.6%, respectively. We present an algorithm to resolve most camouflaged regions and apply it to the Alzheimer’s Disease Sequencing Project. We rescue a rare ten-nucleotide frameshift deletion in CR1, a top Alzheimer’s disease gene, found in disease cases but not in controls.While we could not formally assess the association of the CR1 frameshift mutation with Alzheimer’s disease due to insufficient sample-size, we believe it merits investigating in a larger cohort. There remain thousands of potentially important genomic regions overlooked by short-read sequencing that are largely resolved by long-read technologies.


April 21, 2020  |  

Resource Concentration Modulates the Fate of Dissimilated Nitrogen in a Dual-Pathway Actinobacterium.

Respiratory ammonification and denitrification are two evolutionarily unrelated dissimilatory nitrogen (N) processes central to the global N cycle, the activity of which is thought to be controlled by carbon (C) to nitrate (NO3-) ratio. Here we find that Intrasporangium calvum C5, a novel dual-pathway denitrifier/respiratory ammonifier, disproportionately utilizes ammonification rather than denitrification when grown under low C concentrations, even at low C:NO3- ratios. This finding is in conflict with the paradigm that high C:NO3- ratios promote ammonification and low C:NO3- ratios promote denitrification. We find that the protein atomic composition for denitrification modules (NirK) are significantly cost minimized for C and N compared to ammonification modules (NrfA), indicating that limitation for C and N is a major evolutionary selective pressure imprinted in the architecture of these proteins. The evolutionary precedent for these findings suggests ecological importance for microbial activity as evidenced by higher growth rates when I. calvum grows predominantly using its ammonification pathway and by assimilating its end-product (ammonium) for growth under ammonium-free conditions. Genomic analysis of I. calvum further reveals a versatile ecophysiology to cope with nutrient stress and redox conditions. Metabolite and transcriptional profiles during growth indicate that enzyme modules, NrfAH and NirK, are not constitutively expressed but rather induced by nitrite production via NarG. Mechanistically, our results suggest that pathway selection is driven by intracellular redox potential (redox poise), which may be lowered when resource concentrations are low, thereby decreasing catalytic activity of upstream electron transport steps (i.e., the bc1 complex) needed for denitrification enzymes. Our work advances our understanding of the biogeochemical flexibility of N-cycling organisms, pathway evolution, and ecological food-webs.


October 23, 2019  |  

Creating and evaluating accurate CRISPR-Cas9 scalpels for genomic surgery.

The simplicity of site-specific genome targeting by type II clustered, regularly interspaced, short palindromic repeat (CRISPR)-Cas9 nucleases, along with their robust activity profile, has changed the landscape of genome editing. These favorable properties have made the CRISPR-Cas9 system the technology of choice for sequence-specific modifications in vertebrate systems. For many applications, whether the focus is on basic science investigations or therapeutic efficacy, activity and precision are important considerations when one is choosing a nuclease platform, target site and delivery method. Here we review recent methods for increasing the activity and accuracy of Cas9 and assessing the extent of off-target cleavage events.


October 23, 2019  |  

CRISPR/Cas9-mediated scanning for regulatory elements required for HPRT1 expression via thousands of large, programmed genomic deletions.

The extent to which non-coding mutations contribute to Mendelian disease is a major unknown in human genetics. Relatedly, the vast majority of candidate regulatory elements have yet to be functionally validated. Here, we describe a CRISPR-based system that uses pairs of guide RNAs (gRNAs) to program thousands of kilobase-scale deletions that deeply scan across a targeted region in a tiling fashion (“ScanDel”). We applied ScanDel to HPRT1, the housekeeping gene underlying Lesch-Nyhan syndrome, an X-linked recessive disorder. Altogether, we programmed 4,342 overlapping 1 and 2 kb deletions that tiled 206 kb centered on HPRT1 (including 87 kb upstream and 79 kb downstream) with median 27-fold redundancy per base. We functionally assayed programmed deletions in parallel by selecting for loss of HPRT function with 6-thioguanine. As expected, sequencing gRNA pairs before and after selection confirmed that all HPRT1 exons are needed. However, HPRT1 function was robust to deletion of any intergenic or deeply intronic non-coding region, indicating that proximal regulatory sequences are sufficient for HPRT1 expression. Although our screen did identify the disruption of exon-proximal non-coding sequences (e.g., the promoter) as functionally consequential, long-read sequencing revealed that this signal was driven by rare, imprecise deletions that extended into exons. Our results suggest that no singular distal regulatory element is required for HPRT1 expression and that distal mutations are unlikely to contribute substantially to Lesch-Nyhan syndrome burden. Further application of ScanDel could shed light on the role of regulatory mutations in disease at other loci while also facilitating a deeper understanding of endogenous gene regulation. Copyright © 2017 American Society of Human Genetics. All rights reserved.


October 23, 2019  |  

Efficient CRISPR/Cas9-mediated editing of trinucleotide repeat expansion in myotonic dystrophy patient-derived iPS and myogenic cells.

CRISPR/Cas9 is an attractive platform to potentially correct dominant genetic diseases by gene editing with unprecedented precision. In the current proof-of-principle study, we explored the use of CRISPR/Cas9 for gene-editing in myotonic dystrophy type-1 (DM1), an autosomal-dominant muscle disorder, by excising the CTG-repeat expansion in the 3′-untranslated-region (UTR) of the human myotonic dystrophy protein kinase (DMPK) gene in DM1 patient-specific induced pluripotent stem cells (DM1-iPSC), DM1-iPSC-derived myogenic cells and DM1 patient-specific myoblasts. To eliminate the pathogenic gain-of-function mutant DMPK transcript, we designed a dual guide RNA based strategy that excises the CTG-repeat expansion with high efficiency, as confirmed by Southern blot and single molecule real-time (SMRT) sequencing. Correction efficiencies up to 90% could be attained in DM1-iPSC as confirmed at the clonal level, following ribonucleoprotein (RNP) transfection of CRISPR/Cas9 components without the need for selective enrichment. Expanded CTG repeat excision resulted in the disappearance of ribonuclear foci, a quintessential cellular phenotype of DM1, in the corrected DM1-iPSC, DM1-iPSC-derived myogenic cells and DM1 myoblasts. Consequently, the normal intracellular localization of the muscleblind-like splicing regulator 1 (MBNL1) was restored, resulting in the normalization of splicing pattern of SERCA1. This study validates the use of CRISPR/Cas9 for gene editing of repeat expansions.


October 23, 2019  |  

Function-based identification of mammalian enhancers using site-specific integration.

The accurate and comprehensive identification of functional regulatory sequences in mammalian genomes remains a major challenge. Here we describe site-specific integration fluorescence-activated cell sorting followed by sequencing (SIF-seq), an unbiased, medium-throughput functional assay for the discovery of distant-acting enhancers. Targeted single-copy genomic integration into pluripotent cells, reporter assays and flow cytometry are coupled with high-throughput DNA sequencing to enable parallel screening of large numbers of DNA sequences. By functionally interrogating >500 kilobases (kb) of mouse and human sequence in mouse embryonic stem cells for enhancer activity we identified enhancers at pluripotency loci including NANOG. In in vitro-differentiated cardiomyocytes and neural progenitor cells, we identified cardiac enhancers and neuronal enhancers, respectively. SIF-seq is a powerful and flexible method for de novo functional identification of mammalian enhancers in a potentially wide variety of cell types.


October 23, 2019  |  

Adeno-associated virus type 2 wild-type and vector-mediated genomic integration profiles of human diploid fibroblasts analyzed by third-generation PacBio DNA sequencing.

Genome-wide analysis of adeno-associated virus (AAV) type 2 integration in HeLa cells has shown that wild-type AAV integrates at numerous genomic sites, including AAVS1 on chromosome 19q13.42. Multiple GAGY/C repeats, resembling consensus AAV Rep-binding sites are preferred, whereas rep-deficient AAV vectors (rAAV) regularly show a random integration profile. This study is the first study to analyze wild-type AAV integration in diploid human fibroblasts. Applying high-throughput third-generation PacBio-based DNA sequencing, integration profiles of wild-type AAV and rAAV are compared side by side. Bioinformatic analysis reveals that both wild-type AAV and rAAV prefer open chromatin regions. Although genomic features of AAV integration largely reproduce previous findings, the pattern of integration hot spots differs from that described in HeLa cells before. DNase-Seq data for human fibroblasts and for HeLa cells reveal variant chromatin accessibility at preferred AAV integration hot spots that correlates with variant hot spot preferences. DNase-Seq patterns of these sites in human tissues, including liver, muscle, heart, brain, skin, and embryonic stem cells further underline variant chromatin accessibility. In summary, AAV integration is dependent on cell-type-specific, variant chromatin accessibility leading to random integration profiles for rAAV, whereas wild-type AAV integration sites cluster near GAGY/C repeats.Adeno-associated virus type 2 (AAV) is assumed to establish latency by chromosomal integration of its DNA. This is the first genome-wide analysis of wild-type AAV2 integration in diploid human cells and the first to compare wild-type to recombinant AAV vector integration side by side under identical experimental conditions. Major determinants of wild-type AAV integration represent open chromatin regions with accessible consensus AAV Rep-binding sites. The variant chromatin accessibility of different human tissues or cell types will have impact on vector targeting to be considered during gene therapy. Copyright © 2014, American Society for Microbiology. All Rights Reserved.


October 23, 2019  |  

Bioengineered AAV capsids with combined high human liver transduction in vivo and unique humoral seroreactivity.

Existing recombinant adeno-associated virus (rAAV) serotypes for delivering in vivo gene therapy treatments for human liver diseases have not yielded combined high-level human hepatocyte transduction and favorable humoral neutralization properties in diverse patient groups. Yet, these combined properties are important for therapeutic efficacy. To bioengineer capsids that exhibit both unique seroreactivity profiles and functionally transduce human hepatocytes at therapeutically relevant levels, we performed multiplexed sequential directed evolution screens using diverse capsid libraries in both primary human hepatocytes in vivo and with pooled human sera from thousands of patients. AAV libraries were subjected to five rounds of in vivo selection in xenografted mice with human livers to isolate an enriched human-hepatotropic library that was then used as input for a sequential on-bead screen against pooled human immunoglobulins. Evolved variants were vectorized and validated against existing hepatotropic serotypes. Two of the evolved AAV serotypes, NP40 and NP59, exhibited dramatically improved functional human hepatocyte transduction in vivo in xenografted mice with human livers, along with favorable human seroreactivity profiles, compared with existing serotypes. These novel capsids represent enhanced vector delivery systems for future human liver gene therapy applications. Copyright © 2017. Published by Elsevier Inc.


October 23, 2019  |  

Cas9-mediated allelic exchange repairs compound heterozygous recessive mutations in mice.

We report a genome-editing strategy to correct compound heterozygous mutations, a common genotype in patients with recessive genetic disorders. Adeno-associated viral vector delivery of Cas9 and guide RNA induces allelic exchange and rescues the disease phenotype in mouse models of hereditary tyrosinemia type I and mucopolysaccharidosis type I. This approach recombines non-mutated genetic information present in two heterozygous alleles into one functional allele without using donor DNA templates.


October 23, 2019  |  

The genome of common long-arm octopus Octopus minor.

The common long-arm octopus (Octopus minor) is found in mudflats of subtidal zones and faces numerous environmental challenges. The ability to adapt its morphology and behavioral repertoire to diverse environmental conditions makes the species a promising model for understanding genomic adaptation and evolution in cephalopods.The final genome assembly of O. minor is 5.09 Gb, with a contig N50 size of 197 kb and longest size of 3.027 Mb, from a total of 419 Gb raw reads generated using the Pacific Biosciences RS II platform. We identified 30,010 genes; 44.43% of the genome is composed of repeat elements. The genome-wide phylogenetic tree indicated the divergence time between O. minor and Octopus bimaculoides was estimated to be 43 million years ago based on single-copy orthologous genes. In total, 178 gene families are expanded in O. minor in the 14 bilaterian species.We found that the O. minor genome was larger than that of closely related O. bimaculoides, and this difference could be explained by enlarged introns and recently diversified transposable elements. The high-quality O. minor genome assembly provides a valuable resource for understanding octopus genome evolution and the molecular basis of adaptations to mudflats.


September 22, 2019  |  

Transcriptome profiling in the spathe of Anthurium andraeanum ‘Albama’ and its anthocyanin-loss mutant ‘Xueyu’.

Anthurium andraeanum is a popular tropical ornamental plant. Its spathes are brilliantly coloured due to variable anthocyanin contents. To examine the mechanisms that control anthocyanin biosynthesis, we sequenced the spathe transcriptomes of ‘Albama’, a red-spathed cultivar of A. andraeanum, and ‘Xueyu’, its anthocyanin-loss mutant. Both long reads and short reads were sequenced. Long read sequencing produced 805,869 raw reads, resulting in 83,073 high-quality transcripts. Short read sequencing produced 347.79?M reads, and the subsequent assembly resulted in 111,674 unigenes. High-quality transcripts and unigenes were quantified using the short reads, and differential expression analysis was performed between ‘Albama’ and ‘Xueyu’. Obtaining high-quality, full-length transcripts enabled the detection of long transcript structures and transcript variants. These data provide a foundation to elucidate the mechanisms regulating the biosynthesis of anthocyanin in A. andraeanum.


September 22, 2019  |  

Membrane attack complex-associated molecules from redlip mullet (Liza haematocheila): Molecular characterization and transcriptional evidence of C6, C7, C8ß, and C9 in innate immunity.

The redlip mullet (Liza haematocheila) is one of the most economically important fish in Korea and other East Asian countries; it is susceptible to infections by pathogens such as Lactococcus garvieae, Argulus spp., Trichodina spp., and Vibrio spp. Learning about the mechanisms of the complement system of the innate immunity of redlip mullet is important for efforts towards eradicating pathogens. Here, we report a comprehensive study of the terminal complement complex (TCC) components that form the membrane attack complex (MAC) through in-silico characterization and comparative spatial and temporal expression profiling. Five conserved domains (TSP1, LDLa, MACPF, CCP, and FIMAC) were detected in the TCC components, but the CCP and FIMAC domains were absent in MuC8ß and MuC9. Expression analysis of four TCC genes from healthy redlip mullets showed the highest expression levels in the liver, whereas limited expression was observed in other tissues; immune-induced expression in the head kidney and spleen revealed significant responses against Lactococcus garvieae and poly I:C injection, suggesting their involvement in MAC formation in response to harmful pathogenic infections. Furthermore, the response to poly I:C may suggest the role of TCC components in the breakdown of the membrane of enveloped viruses. These findings may help to elucidate the mechanisms behind the complement system of the teleosts innate immunity. Copyright © 2018 Elsevier Ltd. All rights reserved.


September 22, 2019  |  

Comparison of the mitochondrial genomes and steady state transcriptomes of two strains of the trypanosomatid parasite, Leishmania tarentolae.

U-insertion/deletion RNA editing is a post-transcriptional mitochondrial RNA modification phenomenon required for viability of trypanosomatid parasites. Small guide RNAs encoded mainly by the thousands of catenated minicircles contain the information for this editing. We analyzed by NGS technology the mitochondrial genomes and transcriptomes of two strains, the old lab UC strain and the recently isolated LEM125 strain. PacBio sequencing provided complete minicircle sequences which avoided the assembly problem of short reads caused by the conserved regions. Minicircles were identified by a characteristic size, the presence of three short conserved sequences, a region of inherently bent DNA and the presence of single gRNA genes at a fairly defined location. The LEM125 strain contained over 114 minicircles encoding different gRNAs and the UC strain only ~24 minicircles. Some LEM125 minicircles contained no identifiable gRNAs. Approximate copy numbers of the different minicircle classes in the network were determined by the number of PacBio CCS reads that assembled to each class. Mitochondrial RNA libraries from both strains were mapped against the minicircle and maxicircle sequences. Small RNA reads mapped to the putative gRNA genes but also to multiple regions outside the genes on both strands and large RNA reads mapped in many cases over almost the entire minicircle on both strands. These data suggest that minicircle transcription is complete and bidirectional, with 3′ processing yielding the mature gRNAs. Steady state RNAs in varying abundances are derived from all maxicircle genes, including portions of the repetitive divergent region. The relative extents of editing in both strains correlated with the presence of a cascade of cognate gRNAs. These data should provide the foundation for a deeper understanding of this dynamic genetic system as well as the evolutionary variation of editing in different strains.


September 22, 2019  |  

Two phospholipid scramblase 1-related proteins (PLSCR1like-a & -b) from Liza haematocheila: Molecular and transcriptional features and expression analysis after immune stimulation.

Phospholipid scramblases (PLSCRs) are a family of transmembrane proteins known to be responsible for Ca2+-mediated bidirectional phospholipid translocation in the plasma membrane. Apart from the scrambling activity of PLSCRs, recent studies revealed their diverse other roles, including antiviral defense, tumorigenesis, protein-DNA interactions, apoptosis regulation, and cell activation. Nonetheless, the biological and transcriptional functions of PLSCRs in fish have not been discovered to date. Therefore, in this study, two new members related to the PLSCR1 family were identified in the red lip mullet (Liza haematocheila) as MuPLSCR1like-a and MuPLSCR1like-b, and their characteristics were studied at molecular and transcriptional levels. Sequence analysis revealed that MuPLSCR1like-a and MuPLSCR1like-b are composed of 245 and 228 amino acid residues (aa) with the predicted molecular weights of 27.82 and 25.74?kDa, respectively. A constructed phylogenetic tree showed that MuPLSCR1like-a and MuPLSCR1like-b are clustered together with other known PLSCR1 and -2 orthologues, thus pointing to the relatedness to both PLSCR1 and PLSCR2 families. Two-dimensional (2D) and 3D graphical representations illustrated the well-known 12-stranded ß-barrel structure of MuPLSCR1like-a and MuPLSCR1like-b with transmembrane orientation toward the phospholipid bilayer. In analysis of tissue-specific expression, the highest expression of MuPLSCR1like-a was observed in the intestine, whereas MuPLSCR1like-b was highly expressed in the brain, indicating isoform specificity. Of note, we found that the transcription of MuPLSCR1like-a and MuPLSCR1like-b was significantly upregulated when the fish were stimulated with poly(I:C), suggesting that such immune responses target viral infections. Overall, this study provides the first experimental insight into the characteristics and immune-system relevance of PLSCR1-related genes in red lip mullets. Copyright © 2018 Elsevier Ltd. All rights reserved.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.