Menu
April 21, 2020

A robust benchmark for germline structural variant detection

New technologies and analysis methods are enabling genomic structural variants (SVs) to be detected with ever-increasing accuracy, resolution, and comprehensiveness. Translating these methods to routine research and clinical practice requires robust benchmark sets. We developed the first benchmark set for identification of both false negative and false positive germline SVs, which complements recent efforts emphasizing increasingly comprehensive characterization of SVs. To create this benchmark for a broadly consented son in a Personal Genome Project trio with broadly available cells and DNA, the Genome in a Bottle (GIAB) Consortium integrated 19 sequence-resolved variant calling methods, both alignment- and de novo assembly-based, from short-, linked-, and long-read sequencing, as well as optical and electronic mapping. The final benchmark set contains 12745 isolated, sequence-resolved insertion and deletion calls =50 base pairs (bp) discovered by at least 2 technologies or 5 callsets, genotyped as heterozygous or homozygous variants by long reads. The Tier 1 benchmark regions, for which any extra calls are putative false positives, cover 2.66 Gbp and 9641 SVs supported by at least one diploid assembly. Support for SVs was assessed using svviz with short-, linked-, and long-read sequence data. In general, there was strong support from multiple technologies for the benchmark SVs, with 90 % of the Tier 1 SVs having support in reads from more than one technology. The Mendelian genotype error rate was 0.3 %, and genotype concordance with manual curation was >98.7 %. We demonstrate the utility of the benchmark set by showing it reliably identifies both false negatives and false positives in high-quality SV callsets from short-, linked-, and long-read sequencing and optical mapping.


April 21, 2020

The Genome of the Zebra Mussel, Dreissena polymorpha: A Resource for Invasive Species Research

The zebra mussel, Dreissena polymorpha, continues to spread from its native range in Eurasia to Europe and North America, causing billions of dollars in damage and dramatically altering invaded aquatic ecosystems. Despite these impacts, there are few genomic resources for Dreissena or related bivalves, with nearly 450 million years of divergence between zebra mussels and its closest sequenced relative. Although the D. polymorpha genome is highly repetitive, we have used a combination of long-read sequencing and Hi-C-based scaffolding to generate the highest quality molluscan assembly to date. Through comparative analysis and transcriptomics experiments we have gained insights into processes that likely control the invasive success of zebra mussels, including shell formation, synthesis of byssal threads, and thermal tolerance. We identified multiple intact Steamer-Like Elements, a retrotransposon that has been linked to transmissible cancer in marine clams. We also found that D. polymorpha have an unusual 67 kb mitochondrial genome containing numerous tandem repeats, making it the largest observed in Eumetazoa. Together these findings create a rich resource for invasive species research and control efforts.


April 21, 2020

ASA3P: An automatic and scalable pipeline for the assembly, annotation and higher level analysis of closely related bacterial isolates

Whole genome sequencing of bacteria has become daily routine in many fields. Advances in DNA sequencing technologies and continuously dropping costs have resulted in a tremendous increase in the amounts of available sequence data. However, comprehensive in-depth analysis of the resulting data remains an arduous and time consuming task. In order to keep pace with these promising but challenging developments and to transform raw data into valuable information, standardized analyses and scalable software tools are needed. Here, we introduce ASA3P, a fully automatic, locally executable and scalable assembly, annotation and analysis pipeline for bacterial genomes. The pipeline automatically executes necessary data processing steps, i.e. quality clipping and assembly of raw sequencing reads, scaffolding of contigs and annotation of the resulting genome sequences. Furthermore, ASA3P conducts comprehensive genome characterizations and analyses, e.g. taxonomic classification, detection of antibiotic resistance genes and identification of virulence factors. All results are presented via an HTML5 user interface providing aggregated information, interactive visualizations and access to intermediate results in standard bioinformatics file formats. We distribute ASA3P in two versions: a locally executable Docker container for small-to-medium-scale projects and an OpenStack based cloud computing version able to automatically create and manage self-scaling compute clusters. Thus, automatic and standardized analysis of hundreds of bacterial genomes becomes feasible within hours. The software and further information is available at: http://asap.computational.bio.


April 21, 2020

Whole-genome sequence of Arthrinium phaeospermum, a globally distributed pathogenic fungus.

Arthrinium phaeospermum (Corda) M.B. Ellis is a globally distributed pathogenic fungus with a wide host range; its hosts include not only plants, but also humans and animals. This study aimed to develop genomic resources for A. phaeospermum to provide solid data and a theoretical basis for further studies of its pathogenesis, transcriptomics, proteomics, metabolomics and RNA genomics. The genome was obtained from the mycelia of the strain AP-Z13 using a combination of analyses with the high-throughput Illumina HiSeq 4000 system and PacBio RSII LongRead sequencing platform. Functional annotation was performed by BLASTing protein sequences against those in different publicly available databases to obtain their corresponding annotations. The genome is 48.45?Mb in size, with an N90 scaffold size of 1,931,147?bp, and encodes 19,836 putative predicted genes. This is the first report of the genome-scale assembly and annotation for A. phaeospermum, the first species in the genus Arthrinium to be subjected to whole genome sequencing. Copyright © 2019 Elsevier Inc. All rights reserved.


April 21, 2020

Generating amplicon reads for microbial community assessment with next-generation sequencing.

Marker gene amplicon sequencing is often preferred over whole genome sequencing for microbial community characterization, due to its lower cost while still enabling assessment of uncultivable organisms. This technique involves many experimental steps, each of which can be a source of errors and bias. We present an up-to-date overview of the whole experimental pipeline, from sampling to sequencing reads, and give information allowing for informed choices at each step of both planning and execution of a microbial community assessment study. When applicable, we also suggest ways of avoiding inherent pitfalls in amplicon sequencing. © 2019 The Society for Applied Microbiology.


April 21, 2020

Rational development of transformation in Clostridium thermocellum ATCC 27405 via complete methylome analysis and evasion of native restriction-modification systems.

A major barrier to both metabolic engineering and fundamental biological studies is the lack of genetic tools in most microorganisms. One example is Clostridium thermocellum ATCC 27405T, where genetic tools are not available to help validate decades of hypotheses. A significant barrier to DNA transformation is restriction-modification systems, which defend against foreign DNA methylated differently than the host. To determine the active restriction-modification systems in this strain, we performed complete methylome analysis via single-molecule, real-time sequencing to detect 6-methyladenine and 4-methylcytosine and the rarely used whole-genome bisulfite sequencing to detect 5-methylcytosine. Multiple active systems were identified, and corresponding DNA methyltransferases were expressed from the Escherichia coli chromosome to mimic the C. thermocellum methylome. Plasmid methylation was experimentally validated and successfully electroporated into C. thermocellum ATCC 27405. This combined approach enabled genetic modification of the C. thermocellum-type strain and acts as a blueprint for transformation of other non-model microorganisms.


April 21, 2020

Comparative genomics reveals unique wood-decay strategies and fruiting body development in the Schizophyllaceae.

Agaricomycetes are fruiting body-forming fungi that produce some of the most efficient enzyme systems to degrade wood. Despite decades-long interest in their biology, the evolution and functional diversity of both wood-decay and fruiting body formation are incompletely known. We performed comparative genomic and transcriptomic analyses of wood-decay and fruiting body development in Auriculariopsis ampla and Schizophyllum commune (Schizophyllaceae), species with secondarily simplified morphologies, an enigmatic wood-decay strategy and weak pathogenicity to woody plants. The plant cell wall-degrading enzyme repertoires of Schizophyllaceae are transitional between those of white rot species and less efficient wood-degraders such as brown rot or mycorrhizal fungi. Rich repertoires of suberinase and tannase genes were found in both species, with tannases restricted to Agaricomycetes that preferentially colonize bark-covered wood, suggesting potential complementation of their weaker wood-decaying abilities and adaptations to wood colonization through the bark. Fruiting body transcriptomes revealed a high rate of divergence in developmental gene expression, but also several genes with conserved expression patterns, including novel transcription factors and small-secreted proteins, some of the latter which might represent fruiting body effectors. Taken together, our analyses highlighted novel aspects of wood-decay and fruiting body development in an important family of mushroom-forming fungi. © 2019 The Authors. New Phytologist © 2019 New Phytologist Trust.


April 21, 2020

Complete genome sequence of Paracoccus sp. Arc7-R13, a silver nanoparticles synthesizing bacterium isolated from Arctic Ocean sediments

Paracoccus sp. Arc7-R13, a silver nanoparticles (AgNPs) synthesizing bacterium, was isolated from Arctic Ocean sediment. Here we describe the complete genome of Paracoccus sp. Arc7-R13. The complete genome contains 4,040,012?bp with 66.66?mol%?G?+?C content, including one circular chromosome of 3,231,929?bp (67.45?mol%?G?+?C content), and eight plasmids with length ranging from 24,536?bp to 199,685?bp. The genome contains 3835 protein-coding genes (CDSs), 49 tRNA genes, as well as 3 rRNA operons as 16S-23S-5S rRNA. Based on the gene annotation and Swiss-Prot analysis, a total of 15 genes belonging to 11 kinds, including silver exporting P-type ATPase (SilP), alkaline phosphatase, nitroreductase, thioredoxin reductase, NADPH dehydrogenase and glutathione peroxidase, might be related to the synthesis of AgNPs. Meanwhile, many additional genes associated with synthesis of AgNPs such as protein-disulfide isomerase, c-type cytochrome, glutathione synthase and dehydrogenase reductase were also identified.


April 21, 2020

Acquired N-Linked Glycosylation Motifs in B-Cell Receptors of Primary Cutaneous B-Cell Lymphoma and the Normal B-Cell Repertoire.

Primary cutaneous follicle center lymphoma (PCFCL) is a rare mature B-cell lymphoma with an unknown etiology. PCFCL resembles follicular lymphoma (FL) by cytomorphologic and microarchitectural criteria. FL B cells are selected for N-linked glycosylation motifs in their B-cell receptors (BCRs) that are acquired during continuous somatic hypermutation. The stimulation of mannosylated BCR by lectins on the tumor microenvironment is therefore a candidate driver in FL pathogenesis. We investigated whether the same mechanism could play a role in PCFCL pathogenesis. Full-length functional variable, diversity, and joining gene sequences of 18 PCFCL and 8 primary cutaneous diffuse large B-cell lymphoma, leg-type were identified by unbiased Anchoring Reverse Transcription of Immunoglobulin Sequences and Amplification by Nested PCR and BCR reconstruction from RNA sequencing data. Low BCR variation demonstrated negligible ongoing somatic hypermutation in PCFCL and primary cutaneous diffuse large B-cell lymphoma, leg-type, and indicated that the PCFCL microarchitecture does not act as a functional germinal center. Similar to FL but in contrast to primary cutaneous diffuse large B-cell lymphoma, leg-type, BCR genes of 15 PCFCLs (83%) had acquired N-linked glycosylation motifs. These motifs were located at the BCR positions converted to N-linked glycosylation motifs in normal B-cell repertoires with low prevalence but mostly at different positions than those found in FL. The cutaneous localization of PCFCL might suggest a role for lectins from commensal skin bacteria in PCFCL lymphomagenesis.Copyright © 2019 The Authors. Published by Elsevier Inc. All rights reserved.


April 21, 2020

Plasmid-encoded tet(X) genes that confer high-level tigecycline resistance in Escherichia coli.

Tigecycline is one of the last-resort antibiotics to treat complicated infections caused by both multidrug-resistant Gram-negative and Gram-positive bacteria1. Tigecycline resistance has sporadically occurred in recent years, primarily due to chromosome-encoding mechanisms, such as overexpression of efflux pumps and ribosome protection2,3. Here, we report the emergence of the plasmid-mediated mobile tigecycline resistance mechanism Tet(X4) in Escherichia coli isolates from China, which is capable of degrading all tetracyclines, including tigecycline and the US FDA newly approved eravacycline. The tet(X4)-harbouring IncQ1 plasmid is highly transferable, and can be successfully mobilized and stabilized in recipient clinical and laboratory strains of Enterobacteriaceae bacteria. It is noteworthy that tet(X4)-positive E.?coli strains, including isolates co-harbouring mcr-1, have been widely detected in pigs, chickens, soil and dust samples in China. In vivo murine models demonstrated that the presence of Tet(X4) led to tigecycline treatment failure. Consequently, the emergence of plasmid-mediated Tet(X4) challenges the clinical efficacy of the entire family of tetracycline antibiotics. Importantly, our study raises concern that the plasmid-mediated tigecycline resistance may further spread into various ecological niches and into clinical high-risk pathogens. Collective efforts are in urgent need to preserve the potency of these essential antibiotics.


April 21, 2020

RNA sequencing: the teenage years.

Over the past decade, RNA sequencing (RNA-seq) has become an indispensable tool for transcriptome-wide analysis of differential gene expression and differential splicing of mRNAs. However, as next-generation sequencing technologies have developed, so too has RNA-seq. Now, RNA-seq methods are available for studying many different aspects of RNA biology, including single-cell gene expression, translation (the translatome) and RNA structure (the structurome). Exciting new applications are being explored, such as spatial transcriptomics (spatialomics). Together with new long-read and direct RNA-seq technologies and better computational tools for data analysis, innovations in RNA-seq are contributing to a fuller understanding of RNA biology, from questions such as when and where transcription occurs to the folding and intermolecular interactions that govern RNA function.


April 21, 2020

Comparison of mitochondrial DNA variants detection using short- and long-read sequencing.

The recent advent of long-read sequencing technologies is expected to provide reasonable answers to genetic challenges unresolvable by short-read sequencing, primarily the inability to accurately study structural variations, copy number variations, and homologous repeats in complex parts of the genome. However, long-read sequencing comes along with higher rates of random short deletions and insertions, and single nucleotide errors. The relatively higher sequencing accuracy of short-read sequencing has kept it as the first choice of screening for single nucleotide variants and short deletions and insertions. Albeit, short-read sequencing still suffers from systematic errors that tend to occur at specific positions where a high depth of reads is not always capable to correct for these errors. In this study, we compared the genotyping of mitochondrial DNA variants in three samples using PacBio’s Sequel (Pacific Biosciences Inc., Menlo Park, CA, USA) long-read sequencing and illumina’s HiSeqX10 (illumine Inc., San Diego, CA, USA) short-read sequencing data. We concluded that, despite the differences in the type and frequency of errors in the long-reads sequencing, its accuracy is still comparable to that of short-reads for genotyping short nuclear variants; due to the randomness of errors in long reads, a lower coverage, around 37 reads, can be sufficient to correct for these random errors.


April 21, 2020

Complete genome sequence of Paenisporosarcina antarctica CGMCC 1.6503 T, a marine psychrophilic bacterium isolated from Antarctica

A marine psychrophilic bacterium _Paenisporosarcina antarctica_ CGMCC 1.6503T (= JCM 14646T) was isolated off King George Island, Antarctica (62°13’31? S 58°57’08? W). In this study, we report the complete genome sequence of _Paenisporosarcina antarctica_, which is comprised of 3,972,524?bp with a mean G?+?C content of 37.0%. By gene function and metabolic pathway analyses, studies showed that strain CGMCC 1.6503T encodes a series of genes related to cold adaptation, including encoding fatty acid desaturases, dioxygenases, antifreeze proteins and cold shock proteins, and possesses several two-component regulatory systems, which could assist this strain in responding to the cold stress, the oxygen stress and the osmotic stress in Antarctica. The complete genome sequence of _P. antarctica_ may provide further insights into the genetic mechanism of cold adaptation for Antarctic marine bacteria.


April 21, 2020

Genome sequence resource for Ilyonectria mors-panacis, causing rusty root rot of Panax notoginseng.

Ilyonectria mors-panacis is a serious disease hampering the production of Panax notoginseng, an important Chinese medicinal herb, widely used for its anti-inflammatory, anti-fatigue, hepato-protective, and coronary heart disease prevention effects. Here, we report the first Illumina-Pacbio hybrid sequenced draft genome assembly of I. mors-panacis strain G3B and its annotation. The availability of this genome sequence not only represents an important tool toward understanding the genetics behind the infection mechanism of I. mors-panacis strain G3B but also will help illuminate the complexities of the taxonomy of this species.


April 21, 2020

Decoding and analysis of organelle genomes of Indian tea (Camellia assamica) for phylogenetic confirmation.

The NCBI database has >15 chloroplast (cp) genome sequences available for different Camellia species but none for C. assamica. There is no report of any mitochondrial (mt) genome in the Camellia genus or Theaceae family. With the strong believes that these organelle genomes can play a great tool for taxonomic and phylogenetic analysis, we successfully assembled and analyzed cp and mt genome of C. assamica. We assembled the complete mt genome of C. assamica in a single circular contig of 707,441?bp length comprising of a total of 66 annotated genes, including 35 protein-coding genes, 29 tRNAs and two rRNAs. The first ever cp genome of C. assamica resulted in a circular contig of 157,353?bp length with a typical quadripartite structure. Phylogenetic analysis based on these organelle genomes showed that C. assamica was closely related to C. sinensis and C. leptophylla. It also supports Caryophyllales as Superasterids. Copyright © 2019. Published by Elsevier Inc.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.