Menu
September 22, 2019

Computational identification of novel genes: current and future perspectives.

While it has long been thought that all genomic novelties are derived from the existing material, many genes lacking homology to known genes were found in recent genome projects. Some of these novel genes were proposed to have evolved de novo, ie, out of noncoding sequences, whereas some have been shown to follow a duplication and divergence process. Their discovery called for an extension of the historical hypotheses about gene origination. Besides the theoretical breakthrough, increasing evidence accumulated that novel genes play important roles in evolutionary processes, including adaptation and speciation events. Different techniques are available to identify genes and classify them as novel. Their classification as novel is usually based on their similarity to known genes, or lack thereof, detected by comparative genomics or against databases. Computational approaches are further prime methods that can be based on existing models or leveraging biological evidences from experiments. Identification of novel genes remains however a challenging task. With the constant software and technologies updates, no gold standard, and no available benchmark, evaluation and characterization of genomic novelty is a vibrant field. In this review, the classical and state-of-the-art tools for gene prediction are introduced. The current methods for novel gene detection are presented; the methodological strategies and their limits are discussed along with perspective approaches for further studies.


September 22, 2019

The role of MHC-E in T cell immunity is conserved among humans, rhesus macaques, and cynomolgus macaques.

MHC-E is a highly conserved nonclassical MHC class Ib molecule that predominantly binds and presents MHC class Ia leader sequence-derived peptides for NK cell regulation. However, MHC-E also binds pathogen-derived peptide Ags for presentation to CD8+ T cells. Given this role in adaptive immunity and its highly monomorphic nature in the human population, HLA-E is an attractive target for novel vaccine and immunotherapeutic modalities. Development of HLA-E-targeted therapies will require a physiologically relevant animal model that recapitulates HLA-E-restricted T cell biology. In this study, we investigated MHC-E immunobiology in two common nonhuman primate species, Indian-origin rhesus macaques (RM) and Mauritian-origin cynomolgus macaques (MCM). Compared to humans and MCM, RM expressed a greater number of MHC-E alleles at both the population and individual level. Despite this difference, human, RM, and MCM MHC-E molecules were expressed at similar levels across immune cell subsets, equivalently upregulated by viral pathogens, and bound and presented identical peptides to CD8+ T cells. Indeed, SIV-specific, Mamu-E-restricted CD8+ T cells from RM recognized antigenic peptides presented by all MHC-E molecules tested, including cross-species recognition of human and MCM SIV-infected CD4+ T cells. Thus, MHC-E is functionally conserved among humans, RM, and MCM, and both RM and MCM represent physiologically relevant animal models of HLA-E-restricted T cell immunobiology. Copyright © 2017 by The American Association of Immunologists, Inc.


September 22, 2019

Evaluating the mobility potential of antibiotic resistance genes in environmental resistomes without metagenomics.

Antibiotic resistance genes are ubiquitous in the environment. However, only a fraction of them are mobile and able to spread to pathogenic bacteria. Until now, studying the mobility of antibiotic resistance genes in environmental resistomes has been challenging due to inadequate sensitivity and difficulties in contig assembly of metagenome based methods. We developed a new cost and labor efficient method based on Inverse PCR and long read sequencing for studying mobility potential of environmental resistance genes. We applied Inverse PCR on sediment samples and identified 79 different MGE clusters associated with the studied resistance genes, including novel mobile genetic elements, co-selected resistance genes and a new putative antibiotic resistance gene. The results show that the method can be used in antibiotic resistance early warning systems. In comparison to metagenomics, Inverse PCR was markedly more sensitive and provided more data on resistance gene mobility and co-selected resistances.


September 22, 2019

Koumiss consumption alleviates symptoms of patients with chronic atrophic gastritis: A possible link To modulation of gut microbiota

Intestinal dysbiosisis closely related to a variety of medical conditions, especially gastrointestinal diseases. The present study aimed to investigate the effects of koumiss on chronic atrophic gastritis (CAG) in an out-patient clinical trial (n = 10; all female subjects aged 41-55; body mass index ranging from 19.5 to 25.8). Each patient consumed three servings of koumiss per day (i.e. 250 ml daily before each of 3 meals) for a 60-day period. The improvement of patients’ symptoms was monitored by comparing the total scores of symptoms before and after the treatment. Meanwhile, the changes in the patients’ fecal microbiota composition and specific blood parameters were determined. After the 60-day koumiss administration, significant symptom improvements were observed, as evidenced by the reduction of the total symptoms score, and changes in blood platelet and cholesterol levels. The changes in patients’ fecal microbiota composition were found. The patients’ fecal microbiota fell into two distinct enterotypes, Bacteroides dorei/ Bacteroides uniformis (BB-enterotype) and Prevotella copri (P-enterotype). Significant less Bacteroides uniformis was found in the BB-enterotype patient group, while significant more butyrate-producing bacteria (e.g. Eubacterium rectale and Faecalibacterium prausnitzii) were found in the P-enterotype patient group, following koumiss administration. After stopping koumiss consumption, the relative abundance of some biomarker taxa returned to the original level, suggesting that the gut microbiota modulatory effect was not permanent and that continuous koumiss administration was required to maintain the therapeutic effect. In conclusion, koumiss consumption could alleviate the symptoms of CAG patients. Our results may help understand the mechanism of koumiss in alleviating CAG disease symptoms, facilitating the development of such products with desired therapeutic functions.


September 22, 2019

Single-molecule long-read transcriptome profiling of Platysternon megacephalum mitochondrial genome with gene rearrangement and control region duplication.

Platysternon megacephalum is the sole living representative of the poorly studied turtle lineage Platysternidae. Their mitochondrial genome has been subject to gene rearrangement and control region duplication, resulting in a unique mitochondrial gene order in vertebrates. In this study, we sequenced the first full-length turtle (P. megacephalum) liver transcriptome using single-molecule real-time sequencing to study the transcriptional mechanisms of its mitochondrial genome. ND5 and ND6 anti-sense (ND6AS) forms a single transcript with the same expression in the human mitochondrial genome, but here we demonstrated differential expression of the rearranged ND5 and ND6AS genes in P. megacephalum. And some polycistronic transcripts were also reported in this study. Notably, we detected some novel long non-coding RNAs with alternative polyadenylation from the duplicated control region, and a novel ND6AS transcript composed of a long non-coding sequence, ND6AS, and tRNA-GluAS. These results provide the first description of a mtDNA transcriptome with gene rearrangement and control region duplication. These findings further our understanding of the fundamental concepts of mitochondrial gene transcription and RNA processing, and provide a new insight into the mechanism of transcription regulation of the mitochondrial genome.


September 22, 2019

A comprehensive fungi-specific 18S rRNA gene sequence primer toolkit suited for diverse research issues and sequencing platforms.

Several fungi-specific primers target the 18S rRNA gene sequence, one of the prominent markers for fungal classification. The design of most primers goes back to the last decades. Since then, the number of sequences in public databases increased leading to the discovery of new fungal groups and changes in fungal taxonomy. However, no reevaluation of primers was carried out and relevant information on most primers is missing. With this study, we aimed to develop an 18S rRNA gene sequence primer toolkit allowing an easy selection of the best primer pair appropriate for different sequencing platforms, research aims (biodiversity assessment versus isolate classification) and target groups.We performed an intensive literature research, reshuffled existing primers into new pairs, designed new Illumina-primers, and annealing blocking oligonucleotides. A final number of 439 primer pairs were subjected to in silico PCRs. Best primer pairs were selected and experimentally tested. The most promising primer pair with a small amplicon size, nu-SSU-1333-5’/nu-SSU-1647-3′ (FF390/FR-1), was successful in describing fungal communities by Illumina sequencing. Results were confirmed by a simultaneous metagenomics and eukaryote-specific primer approach. Co-amplification occurred in all sample types but was effectively reduced by blocking oligonucleotides.The compiled data revealed the presence of an enormous diversity of fungal 18S rRNA gene primer pairs in terms of fungal coverage, phylum spectrum and co-amplification. Therefore, the primer pair has to be carefully selected to fulfill the requirements of the individual research projects. The presented primer toolkit offers comprehensive lists of 164 primers, 439 primer combinations, 4 blocking oligonucleotides, and top primer pairs holding all relevant information including primer’s characteristics and performance to facilitate primer pair selection.


September 22, 2019

High-confidence coding and noncoding transcriptome maps.

The advent of high-throughput RNA sequencing (RNA-seq) has led to the discovery of unprecedentedly immense transcriptomes encoded by eukaryotic genomes. However, the transcriptome maps are still incomplete partly because they were mostly reconstructed based on RNA-seq reads that lack their orientations (known as unstranded reads) and certain boundary information. Methods to expand the usability of unstranded RNA-seq data by predetermining the orientation of the reads and precisely determining the boundaries of assembled transcripts could significantly benefit the quality of the resulting transcriptome maps. Here, we present a high-performing transcriptome assembly pipeline, called CAFE, that significantly improves the original assemblies, respectively assembled with stranded and/or unstranded RNA-seq data, by orienting unstranded reads using the maximum likelihood estimation and by integrating information about transcription start sites and cleavage and polyadenylation sites. Applying large-scale transcriptomic data comprising 230 billion RNA-seq reads from the ENCODE, Human BodyMap 2.0, The Cancer Genome Atlas, and GTEx projects, CAFE enabled us to predict the directions of about 220 billion unstranded reads, which led to the construction of more accurate transcriptome maps, comparable to the manually curated map, and a comprehensive lncRNA catalog that includes thousands of novel lncRNAs. Our pipeline should not only help to build comprehensive, precise transcriptome maps from complex genomes but also to expand the universe of noncoding genomes.© 2017 You et al.; Published by Cold Spring Harbor Laboratory Press.


September 22, 2019

A response to Lindsey et al. “Wolbachia pipientis should not be split into multiple species: A response to Ramírez-Puebla et al.”.

In Ramírez-Puebla et al. [18] we compared 34 Wolbachia genomes and constructed phylogenetic trees using genomic data. In general, our results were congruent with previously reported phy- logenetic trees [5,9]. Our datasets were carefully selected, checked and analyzed avoiding horizontally transferred genes. In the case of the wAna genome we did not use the raw data, but the assem- bled genome [22] and 31 genes were used to compare in a dataset of conserved proteins. To confirm our conclusions a new phyloge- nomic analysis was performed excluding the wAna strain in the dataset (Fig. 1). The same topology was obtained, therefore indi- cating that the results were not affected by the presence of this particular strain.


September 22, 2019

Recurrent structural variation, clustered sites of selection, and disease risk for the complement factor H (CFH) gene family.

Structural variation and single-nucleotide variation of the complement factor H (CFH) gene family underlie several complex genetic diseases, including age-related macular degeneration (AMD) and atypical hemolytic uremic syndrome (AHUS). To understand its diversity and evolution, we performed high-quality sequencing of this ~360-kbp locus in six primate lineages, including multiple human haplotypes. Comparative sequence analyses reveal two distinct periods of gene duplication leading to the emergence of four CFH-related (CFHR) gene paralogs (CFHR2 and CFHR4 ~25-35 Mya and CFHR1 and CFHR3 ~7-13 Mya). Remarkably, all evolutionary breakpoints share a common ~4.8-kbp segment corresponding to an ancestral CFHR gene promoter that has expanded independently throughout primate evolution. This segment is recurrently reused and juxtaposed with a donor duplication containing exons 8 and 9 from ancestral CFH, creating four CFHR fusion genes that include lineage-specific members of the gene family. Combined analysis of >5,000 AMD cases and controls identifies a significant burden of a rare missense mutation that clusters at the N terminus of CFH [P = 5.81 × 10-8, odds ratio (OR) = 9.8 (3.67-Infinity)]. A bipolar clustering pattern of rare nonsynonymous mutations in patients with AMD (P < 10-3) and AHUS (P = 0.0079) maps to functional domains that show evidence of positive selection during primate evolution. Our structural variation analysis in >2,400 individuals reveals five recurrent rearrangement breakpoints that show variable frequency among AMD cases and controls. These data suggest a dynamic and recurrent pattern of mutation critical to the emergence of new CFHR genes but also in the predisposition to complex human genetic disease phenotypes.


September 22, 2019

Proteomic detection of immunoglobulin light chain variable region peptides from amyloidosis patient biopsies.

Immunoglobulin light chain (LC) amyloidosis (AL) is caused by deposition of clonal LCs produced by an underlying plasma cell neoplasm. The clonotypic LC sequences are unique to each patient, and they cannot be reliably detected by either immunoassays or standard proteomic workflows that target the constant regions of LCs. We addressed this issue by developing a novel sequence template-based workflow to detect LC variable (LCV) region peptides directly from AL amyloid deposits. The workflow was implemented in a CAP/CLIA compliant clinical laboratory dedicated to proteomic subtyping of amyloid deposits extracted from either formalin-fixed paraffin-embedded tissues or subcutaneous fat aspirates. We evaluated the performance of the workflow on a validation cohort of 30 AL patients, whose amyloidogenic clone was identified using a novel proteogenomics method, and 30 controls. The recall and negative predictive values of the workflow, when identifying the gene family of the AL clone, were 93 and 98%, respectively. Application of the workflow on a clinical cohort of 500 AL amyloidosis samples highlighted a bias in the LCV gene families used by the AL clones. We also detected similarity between AL clones deposited in multiple organs of systemic AL patients. In summary, AL proteomic data sets are rich in LCV region peptides of potential clinical significance that are recoverable with advanced bioinformatics.


September 22, 2019

Candidatus Dactylopiibacterium carminicum, a nitrogen-fixing symbiont of Dactylopius cochineal insects (Hemiptera: Coccoidea: Dactylopiidae)

The domesticated carmine cochineal Dactylopius coccus (scale insect) has commercial value and has been used for more than 500?years for natural red pigment production. Besides the domesticated cochineal, other wild Dactylopius species such as Dactylopius opuntiae are found in the Americas, all feeding on nutrient poor sap from native cacti. To compensate nutritional deficiencies, many insects harbor symbiotic bacteria which provide essential amino acids or vitamins to their hosts. Here, we characterized a symbiont from the carmine cochineal insects, Candidatus Dactylopiibacterium carminicum (betaproteobacterium, Rhodocyclaceae family) and found it in D. coccus and in D. opuntiae ovaries by fluorescent in situ hybridization, suggesting maternal inheritance. Bacterial genomes recovered from metagenomic data derived from whole insects or tissues both from D. coccus and from D. opuntiae were around 3.6?Mb in size. Phylogenomics showed that dactylopiibacteria constituted a closely related clade neighbor to nitrogen fixing bacteria from soil or from various plants including rice and other grass endophytes. Metabolic capabilities were inferred from genomic analyses, showing a complete operon for nitrogen fixation, biosynthesis of amino acids and vitamins and putative traits of anaerobic or microoxic metabolism as well as genes for plant interaction. Dactylopiibacterium nif gene expression and acetylene reduction activity detecting nitrogen fixation were evidenced in D. coccus hemolymph and ovaries, in congruence with the endosymbiont fluorescent in situ hybridization location. Dactylopiibacterium symbionts may compensate for the nitrogen deficiency in the cochineal diet. In addition, this symbiont may provide essential amino acids, recycle uric acid, and increase the cochineal life span.


September 22, 2019

Analysis of gut microbiota – An ever changing landscape.

In the last two decades, the field of metagenomics has greatly expanded due to improvement in sequencing technologies allowing for a more comprehensive characterization of microbial communities. The use of these technologies has led to an unprecedented understanding of human, animal, and environmental microbiomes and have shown that the gut microbiota are comparable to an organ that is intrinsically linked with a variety of diseases. Characterization of microbial communities using next-generation sequencing-by-synthesis approaches have revealed important shifts in microbiota associated with debilitating diseases such as Clostridium difficile infection. But due to limitations in sequence read length, primer biases, and the quality of databases, genus- and species-level classification have been difficult. Third-generation technologies, such as Pacific Biosciences’ single molecule, real-time (SMRT) approach, allow for unbiased, more specific identification of species that are likely clinically relevant. Comparison of Illumina next-generation characterization and SMRT sequencing of samples from patients treated for C. difficile infection revealed similarities in community composition at the phylum and family levels, but SMRT sequencing further allowed for species-level characterization – permitting a better understanding of the microbial ecology of this disease. Thus, as sequencing technologies continue to advance, new species-level insights can be gained in the study of complex and clinically-relevant microbial communities.


September 22, 2019

SparseIso: a novel Bayesian approach to identify alternatively spliced isoforms from RNA-seq data.

Recent advances in high-throughput RNA sequencing (RNA-seq) technologies have made it possible to reconstruct the full transcriptome of various types of cells. It is important to accurately assemble transcripts or identify isoforms for an improved understanding of molecular mechanisms in biological systems.We have developed a novel Bayesian method, SparseIso, to reliably identify spliced isoforms from RNA-seq data. A spike-and-slab prior is incorporated into the Bayesian model to enforce the sparsity for isoform identification, effectively alleviating the problem of overfitting. A Gibbs sampling procedure is further developed to simultaneously identify and quantify transcripts from RNA-seq data. With the sampling approach, SparseIso estimates the joint distribution of all candidate transcripts, resulting in a significantly improved performance in detecting lowly expressed transcripts and multiple expressed isoforms of genes. Both simulation study and real data analysis have demonstrated that the proposed SparseIso method significantly outperforms existing methods for improved transcript assembly and isoform identification.The SparseIso package is available at http://github.com/henryxushi/SparseIso.xuan@vt.edu.Supplementary data are available at Bioinformatics online.© The Author (2017). Published by Oxford University Press. All rights reserved. For Permissions, please email: journals.permissions@oup.com


September 22, 2019

Saliva and tooth biofilm bacterial microbiota in adolescents in a low caries community.

The oral cavity harbours a complex microbiome that is linked to dental diseases and serves as a route to other parts of the body. Here, the aims were to characterize the oral microbiota by deep sequencing in a low-caries population with regular dental care since childhood and search for association with caries prevalence and incidence. Saliva and tooth biofilm from 17-year-olds and mock bacteria communities were analysed using 16S rDNA Illumina MiSeq (v3-v4) and PacBio SMRT (v1-v8) sequencing including validity and reliability estimates. Caries was scored at 17 and 19 years of age. Both sequencing platforms revealed that Firmicutes dominated in the saliva, whereas Firmicutes and Actinobacteria abundances were similar in tooth biofilm. Saliva microbiota discriminated caries-affected from caries-free adolescents, with enumeration of Scardovia wiggsiae, Streptococcus mutans, Bifidobacterium longum, Leptotrichia sp. HOT498, and Selenomonas spp. in caries-affected participants. Adolescents with B. longum in saliva had significantly higher 2-year caries increment. PacBio SMRT revealed Corynebacterium matruchotii as the most prevalent species in tooth biofilm. In conclusion, both sequencing methods were reliable and valid for oral samples, and saliva microbiota was associated with cross-sectional caries prevalence, especially S. wiggsiae, S. mutans, and B. longum; the latter also with the 2-year caries incidence.


September 22, 2019

wtf genes are prolific dual poison-antidote meiotic drivers.

Meiotic drivers are selfish genes that bias their transmission into gametes, defying Mendelian inheritance. Despite the significant impact of these genomic parasites on evolution and infertility, few meiotic drive loci have been identified or mechanistically characterized. Here, we demonstrate a complex landscape of meiotic drive genes on chromosome 3 of the fission yeasts Schizosaccharomyces kambucha and S. pombe. We identify S. kambucha wtf4 as one of these genes that acts to kill gametes (known as spores in yeast) that do not inherit the gene from heterozygotes. wtf4 utilizes dual, overlapping transcripts to encode both a gamete-killing poison and an antidote to the poison. To enact drive, all gametes are poisoned, whereas only those that inherit wtf4 are rescued by the antidote. Our work suggests that the wtf multigene family proliferated due to meiotic drive and highlights the power of selfish genes to shape genomes, even while imposing tremendous costs to fertility.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.