Menu
April 21, 2020

MSC: a metagenomic sequence classification algorithm.

Metagenomics is the study of genetic materials directly sampled from natural habitats. It has the potential to reveal previously hidden diversity of microscopic life largely due to the existence of highly parallel and low-cost next-generation sequencing technology. Conventional approaches align metagenomic reads onto known reference genomes to identify microbes in the sample. Since such a collection of reference genomes is very large, the approach often needs high-end computing machines with large memory which is not often available to researchers. Alternative approaches follow an alignment-free methodology where the presence of a microbe is predicted using the information about the unique k-mers present in the microbial genomes. However, such approaches suffer from high false positives due to trading off the value of k with the computational resources. In this article, we propose a highly efficient metagenomic sequence classification (MSC) algorithm that is a hybrid of both approaches. Instead of aligning reads to the full genomes, MSC aligns reads onto a set of carefully chosen, shorter and highly discriminating model sequences built from the unique k-mers of each of the reference sequences.Microbiome researchers are generally interested in two objectives of a taxonomic classifier: (i) to detect prevalence, i.e. the taxa present in a sample, and (ii) to estimate their relative abundances. MSC is primarily designed to detect prevalence and experimental results show that MSC is indeed a more effective and efficient algorithm compared to the other state-of-the-art algorithms in terms of accuracy, memory and runtime. Moreover, MSC outputs an approximate estimate of the abundances.The implementations are freely available for non-commercial purposes. They can be downloaded from https://drive.google.com/open?id=1XirkAamkQ3ltWvI1W1igYQFusp9DHtVl. © The Author(s) 2019. Published by Oxford University Press. All rights reserved. For Permissions, please e-mail: journals.permissions@oup.com.


April 21, 2020

Intragenomic heterogeneity of intergenic ribosomal DNA spacers in Cucurbita moschata is determined by DNA minisatellites with variable potential to form non-canonical DNA conformations.

The intergenic spacer (IGS) of rDNA is frequently built of long blocks of tandem repeats. To estimate the intragenomic variability of such knotty regions, we employed PacBio sequencing of the Cucurbita moschata genome, in which thousands of rDNA copies are distributed across a number of loci. The rRNA coding regions are highly conserved, indicating intensive interlocus homogenization and/or high selection pressure. However, the IGS exhibits high intragenomic structural diversity. Two repeated blocks, R1 (300-1250 bp) and R2 (290-643 bp), account for most of the IGS variation. They exhibit minisatellite-like features built of multiple periodically spaced short GC-rich sequence motifs with the potential to adopt non-canonical DNA conformations, G-quadruplex-folded and left-handed Z-DNA. The mutual arrangement of these motifs can be used to classify IGS variants into five structural families. Subtle polymorphisms exist within each family due to a variable number of repeats, suggesting the coexistence of an enormous number of IGS variants. The substantial length and structural heterogeneity of IGS minisatellites suggests that the tempo of their divergence exceeds the tempo of the homogenization of rDNA arrays. As frequently occurring among plants, we hypothesize that their instability may influence transcription regulation and/or destabilize rDNA units, possibly spreading them across the genome. © The Author(s) 2019. Published by Oxford University Press on behalf of Kazusa DNA Research Institute.


April 21, 2020

Confident phylogenetic identification of uncultured prokaryotes through long read amplicon sequencing of the 16S-ITS-23S rRNA operon.

Amplicon sequencing of the 16S rRNA gene is the predominant method to quantify microbial compositions and to discover novel lineages. However, traditional short amplicons often do not contain enough information to confidently resolve their phylogeny. Here we present a cost-effective protocol that amplifies a large part of the rRNA operon and sequences the amplicons with PacBio technology. We tested our method on a mock community and developed a read-curation pipeline that reduces the overall read error rate to 0.18%. Applying our method on four environmental samples, we captured near full-length rRNA operon amplicons from a large diversity of prokaryotes. The method operated at moderately high-throughput (22286-37,850 raw ccs reads) and generated a large amount of putative novel archaeal 23S rRNA gene sequences compared to the archaeal SILVA database. These long amplicons allowed for higher resolution during taxonomic classification by means of long (~1000 bp) 16S rRNA gene fragments and for substantially more confident phylogenies by means of combined near full-length 16S and 23S rRNA gene sequences, compared to shorter traditional amplicons (250 bp of the 16S rRNA gene). We recommend our method to those who wish to cost-effectively and confidently estimate the phylogenetic diversity of prokaryotes in environmental samples at high throughput. © 2019 The Authors. Environmental Microbiology published by Society for Applied Microbiology and John Wiley & Sons Ltd.


April 21, 2020

Long-read sequence capture of the haemoglobin gene clusters across codfish species.

Combining high-throughput sequencing with targeted sequence capture has become an attractive tool to study specific genomic regions of interest. Most studies have so far focused on the exome using short-read technology. These approaches are not designed to capture intergenic regions needed to reconstruct genomic organization, including regulatory regions and gene synteny. Here, we demonstrate the power of combining targeted sequence capture with long-read sequencing technology for comparative genomic analyses of the haemoglobin (Hb) gene clusters across eight species separated by up to 70 million years. Guided by the reference genome assembly of the Atlantic cod (Gadus morhua) together with genome information from draft assemblies of selected codfishes, we designed probes covering the two Hb gene clusters. Use of custom-made barcodes combined with PacBio RSII sequencing led to highly continuous assemblies of the LA (~100 kb) and MN (~200 kb) clusters, which include syntenic regions of coding and intergenic sequences. Our results revealed an overall conserved genomic organization of the Hb genes within this lineage, yet with several, lineage-specific gene duplications. Moreover, for some of the species examined, we identified amino acid substitutions at two sites in the Hbb1 gene as well as length polymorphisms in its regulatory region, which has previously been linked to temperature adaptation in Atlantic cod populations. This study highlights the use of targeted long-read capture as a versatile approach for comparative genomic studies by generation of a cross-species genomic resource elucidating the evolutionary history of the Hb gene family across the highly divergent group of codfishes. © 2018 The Authors. Molecular Ecology Resources Published by John Wiley & Sons Ltd.


April 21, 2020

Vaccine-induced protection from homologous tier 2 SHIV challenge in nonhuman primates depends on serum-neutralizing antibody titers.

Passive administration of HIV neutralizing antibodies (nAbs) can protect macaques from hard-to-neutralize (tier 2) chimeric simian-human immunodeficiency virus (SHIV) challenge. However, conditions for nAb-mediated protection after vaccination have not been established. Here, we selected groups of 6 rhesus macaques with either high or low serum nAb titers from a total of 78 animals immunized with recombinant native-like (SOSIP) Env trimers. Repeat intrarectal challenge with homologous tier 2 SHIVBG505 led to rapid infection in unimmunized and low-titer animals. High-titer animals, however, demonstrated protection that was gradually lost as nAb titers waned over time. An autologous serum ID50 nAb titer of ~1:500 afforded more than 90% protection from medium-dose SHIV infection. In contrast, antibody-dependent cellular cytotoxicity and T cell activity did not correlate with protection. Therefore, Env protein-based vaccination strategies can protect against hard-to-neutralize SHIV challenge in rhesus macaques by inducing tier 2 nAbs, provided appropriate neutralizing titers can be reached and maintained. Copyright © 2018 The Author(s). Published by Elsevier Inc. All rights reserved.


April 21, 2020

DNA Methylation at the Schizophrenia and Intelligence GWAS-Implicated MIR137HG Locus May Be Associated with Disease and Cognitive Functions

The largest genome-wide association studies have identified schizophrenia and intelligence associated variants in the MIR137HG locus containing genes encoding microRNA-137 and microRNA-2682. In the present study, we investigated DNA methylation in the MIR137HG intragenic CpG island (CGI) in the peripheral blood of 44 patients with schizophrenia and 50 healthy controls. The CGI included the entire MIR137 gene and the region adjacent to the 5′-end of MIR2682. The aim of the study was to examine the relationship of the CGI methylation with schizophrenia and cognitive functioning. The methylation level of 91 CpG located in the selected region was established for each participant by means of single-molecule real-time bisulfite sequencing. All subjects completed the battery of neuropsychological tests. We found that the CGI was hypomethylated in both groups, except for one site—CpG (chr1: 98?511?049), with significant interindividual variability in methylation. A higher level of methylation of this CpG was seen in male patients and was associated with a decrease in the cognitive index in the combined sample of patients and controls. Our data suggest that further investigation of mechanisms that regulate the MIR137 and MIR2682 genes expression might help to understand the molecular basis of cognitive deficits in schizophrenia.


April 21, 2020

Using Cre-recombinase-driven Polylox barcoding for in vivo fate mapping in mice.

Fate mapping is a powerful genetic tool for linking stem or progenitor cells with their progeny, and hence for defining cell lineages in vivo. The resolution of fate mapping depends on the numbers of distinct markers that are introduced in the beginning into stem or progenitor cells; ideally, numbers should be sufficiently large to allow the tracing of output from individual cells. Highly diverse genetic barcodes can serve this purpose. We recently developed an endogenous genetic barcoding system, termed Polylox. In Polylox, random DNA recombination can be induced by transient activity of Cre recombinase in a 2.1-kb-long artificial recombination substrate that has been introduced into a defined locus in mice (Rosa26Polylox reporter mice). Here, we provide a step-by-step protocol for the use of Polylox, including barcode induction and estimation of induction efficiency, barcode retrieval with single-molecule real-time (SMRT) DNA sequencing followed by computational barcode identification, and the calculation of barcode-generation probabilities, which is key for estimations of single-cell labeling for a given number of stem cells. Thus, Polylox barcoding enables high-resolution fate mapping in essentially all tissues in mice for which inducible Cre driver lines are available. Alternative methods include ex vivo cell barcoding, inducible transposon insertion and CRISPR-Cas9-based barcoding; Polylox currently allows combining non-invasive and cell-type-specific labeling with high label diversity. The execution time of this protocol is ~2-3 weeks for experimental data generation and typically <2 d for computational Polylox decoding and downstream analysis.


April 21, 2020

High-Resolution Evolutionary Analysis of Within-Host Hepatitis C Virus Infection.

Despite recent breakthroughs in treatment of hepatitis C virus (HCV) infection, we have limited understanding of how virus diversity generated within individuals impacts the evolution and spread of HCV variants at the population scale. Addressing this gap is important for identifying the main sources of disease transmission and evaluating the risk of drug-resistance mutations emerging and disseminating in a population.We have undertaken a high-resolution analysis of HCV within-host evolution from 4 individuals coinfected with human immunodeficiency virus 1 (HIV-1). We used long-read, deep-sequenced data of full-length HCV envelope glycoprotein, longitudinally sampled from acute to chronic HCV infection to investigate the underlying viral population and evolutionary dynamics.We found statistical support for population structure maintaining the within-host HCV genetic diversity in 3 out of 4 individuals. We also report the first population genetic estimate of the within-host recombination rate for HCV (0.28 × 10-7 recombination/site/year), which is considerably lower than that estimated for HIV-1 and the overall nucleotide substitution rate estimated during HCV infection.Our findings indicate that population structure and strong genetic linkage shapes within-host HCV evolutionary dynamics. These results will guide the future investigation of potential HCV drug resistance adaptation during infection, and at the population scale. © The Author(s) 2019. Published by Oxford University Press for the Infectious Diseases Society of America.


April 21, 2020

Characterization of Mauritian cynomolgus macaque Fc?R alleles using long-read sequencing.

The Fc?Rs are immune cell surface proteins that bind IgG and facilitate cytokine production, phagocytosis, and Ab-dependent, cell-mediated cytotoxicity. Fc?Rs play a critical role in immunity; variation in these genes is implicated in autoimmunity and other diseases. Cynomolgus macaques are an excellent animal model for many human diseases, and Mauritian cynomolgus macaques (MCMs) are particularly useful because of their restricted genetic diversity. Previous studies of MCM immune gene diversity have focused on the MHC and killer cell Ig-like receptor. In this study, we characterize Fc?R diversity in 48 MCMs using PacBio long-read sequencing to identify novel alleles of each of the four expressed MCM Fc?R genes. We also developed a high-throughput Fc?R genotyping assay, which we used to determine allele frequencies and identify Fc?R haplotypes in more than 500 additional MCMs. We found three alleles for Fc?R1A, seven each for Fc?R2A and Fc?R2B, and four for Fc?R3A; these segregate into eight haplotypes. We also assessed whether different Fc?R alleles confer different Ab-binding affinities by surface plasmon resonance and found minimal difference in binding affinities across alleles for a panel of wild type and Fc-engineered human IgG. This work suggests that although MCMs may not fully represent the diversity of Fc?R responses in humans, they may offer highly reproducible results for mAb therapy and toxicity studies. Copyright © 2018 by The American Association of Immunologists, Inc.


April 21, 2020

Detection of pretreatment minority HIV-1 reverse transcriptase inhibitor-resistant variants by ultra-deep sequencing has a limited impact on virological outcomes.

Ultra-deep sequencing (UDS) is a powerful tool for exploring the impact on virological outcome of minority variants with low frequencies, some even <1% of the virus population. Here, we compared HIV-1 minority variants at baseline, through plasma RNA and PBMC DNA analyses, and the dominant variants at the virological failure (VF) point, to evaluate the impact of minority drug-resistant variants (MDRVs) on virological outcomes.Single-molecule real-time sequencing (SMRTS) was performed on baseline RNA and DNA. The Stanford HIV-1 drug resistance database was used for the identification and evaluation of drug resistance-associated mutations (DRAMs).We classified 50 patients into virological success (VS) and VF groups. We found that the rates of reverse transcriptase inhibitor (RTI) DRAMs determined by SMRTS did not differ significantly within or between groups, whether based on RNA or DNA analyses. There was no significant difference in the level of resistance to specific drugs between groups, in either DNA or RNA analyses, except for the DNA-based analysis of lamivudine, for which there was a trend towards a higher prevalence of intermediate/high-level resistance in the VF group. The RNA MDRVs corresponded to DNA MDRVs, except for M100I and Y188H. Sequencing from DNA appeared to be more sensitive than from RNA to detect MDRVs.Detection of pretreatment minority HIV-1 RTI-resistant variants by UDS showed that MDRVs at baseline were not significantly associated with virological outcome. However, HIV-1 DNA sequencing by UDS was useful for detecting pretreatment drug resistance mutations in patients, potentially affecting virological responses, suggesting a potential clinical relevance for ultra-deep DNA sequencing. © The Author(s) 2019. Published by Oxford University Press on behalf of the British Society for Antimicrobial Chemotherapy. All rights reserved. For permissions, please email: journals.permissions@oup.com.


April 21, 2020

Low-copy nuclear sequence data confirm complex patterns of farina evolution in notholaenid ferns (Pteridaceae).

Notholaenids are an unusual group of ferns that have adapted to, and diversified within, the deserts of Mexico and the southwestern United States. With approximately 40 species, this group is noted for being desiccation-tolerant and having “farina”-powdery exudates of lipophilic flavonoid aglycones-that occur on both the gametophytic and sporophytic phases of their life cycle. The most recent circumscription of notholaenids based on plastid markers surprisingly suggests that several morphological characters, including the expression of farina, are homoplasious. In a striking case of convergence, Notholaena standleyi appears to be distantly related to core Notholaena, with several taxa not before associated with Notholaena nested between them. Such conflicts can be due to morphological homoplasy resulting from adaptive convergence or, alternatively, the plastid phylogeny itself might be misleading, diverging from the true species tree due to incomplete lineage sorting, hybridization, or other factors. In this study, we present a species phylogeny for notholaenid ferns, using four low-copy nuclear loci and concatenated data from three plastid loci. A total of 61 individuals (49 notholaenids and 12 outgroup taxa) were sampled, including 31 out of 37 recognized notholaenid species. The homeologous/allelic nuclear sequences were retrieved using PacBio sequencing and the PURC bioinformatics pipeline. Each dataset was first analyzed individually using maximum likelihood and Bayesian inference, and the species phylogeny was inferred using *BEAST. Although we observed several incongruences between the nuclear and plastid phylogenies, our principal results are broadly congruent with previous inferences based on plastid data. By mapping the presence of farina and their biochemical constitutions on our consensus phylogenetic tree, we confirmed that the characters are indeed homoplastic and have complex evolutionary histories. Hybridization among recognized species of the notholaenid clade appears to be relatively rare compared to that observed in other well-studied fern genera.Copyright © 2019 Elsevier Inc. All rights reserved.


April 21, 2020

NCF1 (p47phox)-deficient chronic granulomatous disease: comprehensive genetic and flow cytometric analysis.

Mutations in NCF1 (p47phox) cause autosomal recessive chronic granulomatous disease (CGD) with abnormal dihydrorhodamine (DHR) assay and absent p47phox protein. Genetic identification of NCF1 mutations is complicated by adjacent highly conserved (>98%) pseudogenes (NCF1B and NCF1C). NCF1 has GTGT at the start of exon 2, whereas the pseudogenes each delete 1 GT (?GT). In p47phox CGD, the most common mutation is ?GT in NCF1 (c.75_76delGT; p.Tyr26fsX26). Sequence homology between NCF1 and its pseudogenes precludes reliable use of standard Sanger sequencing for NCF1 mutations and for confirming carrier status. We first established by flow cytometry that neutrophils from p47phox CGD patients had negligible p47phox expression, whereas those from p47phox CGD carriers had ~60% of normal p47phox expression, independent of the specific mutation in NCF1 We developed a droplet digital polymerase chain reaction (ddPCR) with 2 distinct probes, recognizing either the wild-type GTGT sequence or the ?GT sequence. A second ddPCR established copy number by comparison with the single-copy telomerase reverse transcriptase gene, TERT We showed that 84% of p47phox CGD patients were homozygous for ?GT NCF1 The ddPCR assay also enabled determination of carrier status of relatives. Furthermore, only 79.2% of normal volunteers had 2 copies of GTGT per 6 total (NCF1/NCF1B/NCF1C) copies, designated 2/6; 14.7% had 3/6, and 1.6% had 4/6 GTGT copies. In summary, flow cytometry for p47phox expression quickly identifies patients and carriers of p47phox CGD, and genomic ddPCR identifies patients and carriers of ?GT NCF1, the most common mutation in p47phox CGD.


April 21, 2020

A Species-Wide Inventory of NLR Genes and Alleles in Arabidopsis thaliana.

Infectious disease is both a major force of selection in nature and a prime cause of yield loss in agriculture. In plants, disease resistance is often conferred by nucleotide-binding leucine-rich repeat (NLR) proteins, intracellular immune receptors that recognize pathogen proteins and their effects on the host. Consistent with extensive balancing and positive selection, NLRs are encoded by one of the most variable gene families in plants, but the true extent of intraspecific NLR diversity has been unclear. Here, we define a nearly complete species-wide pan-NLRome in Arabidopsis thaliana based on sequence enrichment and long-read sequencing. The pan-NLRome largely saturates with approximately 40 well-chosen wild strains, with half of the pan-NLRome being present in most accessions. We chart NLR architectural diversity, identify new architectures, and quantify selective forces that act on specific NLRs and NLR domains. Our study provides a blueprint for defining pan-NLRomes.Copyright © 2019 The Author(s). Published by Elsevier Inc. All rights reserved.


April 21, 2020

Noncoding CGG repeat expansions in neuronal intranuclear inclusion disease, oculopharyngodistal myopathy and an overlapping disease.

Noncoding repeat expansions cause various neuromuscular diseases, including myotonic dystrophies, fragile X tremor/ataxia syndrome, some spinocerebellar ataxias, amyotrophic lateral sclerosis and benign adult familial myoclonic epilepsies. Inspired by the striking similarities in the clinical and neuroimaging findings between neuronal intranuclear inclusion disease (NIID) and fragile X tremor/ataxia syndrome caused by noncoding CGG repeat expansions in FMR1, we directly searched for repeat expansion mutations and identified noncoding CGG repeat expansions in NBPF19 (NOTCH2NLC) as the causative mutations for NIID. Further prompted by the similarities in the clinical and neuroimaging findings with NIID, we identified similar noncoding CGG repeat expansions in two other diseases: oculopharyngeal myopathy with leukoencephalopathy and oculopharyngodistal myopathy, in LOC642361/NUTM2B-AS1 and LRP12, respectively. These findings expand our knowledge of the clinical spectra of diseases caused by expansions of the same repeat motif, and further highlight how directly searching for expanded repeats can help identify mutations underlying diseases.


April 21, 2020

Accurate circular consensus long-read sequencing improves variant detection and assembly of a human genome.

The DNA sequencing technologies in use today produce either highly accurate short reads or less-accurate long reads. We report the optimization of circular consensus sequencing (CCS) to improve the accuracy of single-molecule real-time (SMRT) sequencing (PacBio) and generate highly accurate (99.8%) long high-fidelity (HiFi) reads with an average length of 13.5?kilobases (kb). We applied our approach to sequence the well-characterized human HG002/NA24385 genome and obtained precision and recall rates of at least 99.91% for single-nucleotide variants (SNVs), 95.98% for insertions and deletions <50 bp (indels) and 95.99% for structural variants. Our CCS method matches or exceeds the ability of short-read sequencing to detect small variants and structural variants. We estimate that 2,434 discordances are correctable mistakes in the 'genome in a bottle' (GIAB) benchmark set. Nearly all (99.64%) variants can be phased into haplotypes, further improving variant detection. De novo genome assembly using CCS reads alone produced a contiguous and accurate genome with a contig N50 of >15?megabases (Mb) and concordance of 99.997%, substantially outperforming assembly with less-accurate long reads.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.