Menu
July 7, 2019

The fast changing landscape of sequencing technologies and their impact on microbial genome assemblies and annotation.

The emergence of next generation sequencing (NGS) has provided the means for rapid and high throughput sequencing and data generation at low cost, while concomitantly creating a new set of challenges. The number of available assembled microbial genomes continues to grow rapidly and their quality reflects the quality of the sequencing technology used, but also of the analysis software employed for assembly and annotation.In this work, we have explored the quality of the microbial draft genomes across various sequencing technologies. We have compared the draft and finished assemblies of 133 microbial genomes sequenced at the Department of Energy-Joint Genome Institute and finished at the Los Alamos National Laboratory using a variety of combinations of sequencing technologies, reflecting the transition of the institute from Sanger-based sequencing platforms to NGS platforms. The quality of the public assemblies and of the associated gene annotations was evaluated using various metrics. Results obtained with the different sequencing technologies, as well as their effects on downstream processes, were analyzed. Our results demonstrate that the Illumina HiSeq 2000 sequencing system, the primary sequencing technology currently used for de novo genome sequencing and assembly at JGI, has various advantages in terms of total sequence throughput and cost, but it also introduces challenges for the downstream analyses. In all cases assembly results although on average are of high quality, need to be viewed critically and consider sources of errors in them prior to analysis.These data follow the evolution of microbial sequencing and downstream processing at the JGI from draft genome sequences with large gaps corresponding to missing genes of significant biological role to assemblies with multiple small gaps (Illumina) and finally to assemblies that generate almost complete genomes (Illumina+PacBio).


July 7, 2019

Bacteriophage P70: unique morphology and unrelatedness to other Listeria bacteriophages.

Listeria monocytogenes is an important food-borne pathogen, and its bacteriophages find many uses in detection and biocontrol of its host. The novel broad-host-range virulent phage P70 has a unique morphology with an elongated capsid. Its genome sequence was determined by a hybrid sequencing strategy employing Sanger and PacBio techniques. The P70 genome contains 67,170 bp and 119 open reading frames (ORFs). Our analyses suggest that P70 represents an archetype of virus unrelated to other known Listeria bacteriophages.


July 7, 2019

A gapless genome sequence of the fungus Botrytis cinerea.

Following earlier incomplete and fragmented versions of a genome sequence for the grey mould Botrytis cinerea, we here report a gapless, near-finished genome sequence for B. cinerea strain B05.10. The assembly comprises 18 chromosomes and was confirmed by an optical map and a genetic map based on ~75 000 SNP markers. All chromosomes contain fully assembled centromeric regions, and 10 chromosomes have telomeres on both ends. The genetic map consisted of 4153 cM and comparison of genetic distances with the physical distances identified 40 recombination hotspots. The linkage map also identified two mutations, located in the previously described genes Bos1 and BcsdhB, that confer resistance to the fungicides boscalid and iprodione. The genome was predicted to encode 11 701 proteins. RNAseq data from >20 different samples were used to validate and improve gene models. Manual curation of chromosome 1 revealed interesting features, such as the occurrence of a dicistronic transcript and fully overlapping genes in opposite orientations, as well as many spliced antisense transcripts. Manual curation also revealed that UTRs of genes can be complex and long, with many UTRs exceeding lengths of 1 kb and possessing multiple introns. Community annotation is in progress. This article is protected by copyright. All rights reserved. © 2016 BSPP AND JOHN WILEY & SONS LTD.


July 7, 2019

What distinguishes cyanobacteria able to revive after desiccation from those that cannot: the genome aspect.

Filamentous cyanobacteria are the main founders and primary producers in biological desert soil crusts (BSCs) and are likely equipped to cope with one of the harshest environmental conditions on earth including daily hydration/dehydration cycles, high irradiance and extreme temperatures. Here, we resolved and report on the genome sequence of Leptolyngbya ohadii, an important constituent of the BSC. Comparative genomics identified a set of genes present in desiccation-tolerant but not in dehydration-sensitive cyanobacteria. RT qPCR analyses showed that the transcript abundance of many of them is upregulated during desiccation in L. ohadii. In addition, we identified genes where the orthologs detected in desiccation-tolerant cyanobacteria differs substantially from that found in desiccation-sensitive cells. We present two examples, treS and fbpA (encoding trehalose synthase and fructose 1,6-bisphosphate aldolase respectively) where, in addition to the orthologs present in the desiccation-sensitive strains, the resistant cyanobacteria also possess genes with different predicted structures. We show that in both cases the two orthologs are transcribed during controlled dehydration of L. ohadii and discuss the genetic basis for the acclimation of cyanobacteria to the desiccation conditions in desert BSC.© 2016 Society for Applied Microbiology and John Wiley & Sons Ltd.


July 7, 2019

Draft genome assembly and annotation of Glycyrrhiza uralensis, a medicinal legume.

Chinese liquorice/licorice (Glycyrrhiza uralensis) is a leguminous plant species whose roots and rhizomes have been widely used as a herbal medicine and natural sweetener. Whole-genome sequencing is essential for gene discovery studies and molecular breeding in liquorice. Here, we report a draft assembly of the approximately 379-Mb whole-genome sequence of strain 308-19 of G. uralensis; this assembly contains 34 445 predicted protein-coding genes. Comparative analyses suggested well-conserved genomic components and collinearity of gene loci (synteny) between the genome of liquorice and those of other legumes such as Medicago and chickpea. We observed that three genes involved in isoflavonoid biosynthesis, namely, 2-hydroxyisoflavanone synthase (CYP93C), 2,7,4′-trihydroxyisoflavanone 4′-O-methyltransferase/isoflavone 4′-O-methyltransferase (HI4OMT) and isoflavone-7-O-methyltransferase (7-IOMT) formed a cluster on the scaffold of the liquorice genome and showed conserved microsynteny with Medicago and chickpea. Based on the liquorice genome annotation, we predicted genes in the P450 and UDP-dependent glycosyltransferase (UGT) superfamilies, some of which are involved in triterpenoid saponin biosynthesis, and characterised their gene expression with the reference genome sequence. The genome sequencing and its annotations provide an essential resource for liquorice improvement through molecular breeding and the discovery of useful genes for engineering bioactive components through synthetic biology approaches.© 2016 The Authors The Plant Journal © 2016 John Wiley & Sons Ltd.


July 7, 2019

Whole genome sequencing analysis of the cutaneous pathogenic yeast Malassezia restricta and identification of the major lipase expressed on the scalp of patients with dandruff.

Malassezia species are opportunistic pathogenic fungi that are frequently associated with seborrhoeic dermatitis, including dandruff. Most Malassezia species are lipid dependent, a property that is compensated by breaking down host sebum into fatty acids by lipases. In this study, we aimed to sequence and analyse the whole genome of Malassezia restricta KCTC 27527, a clinical isolate from a Korean patient with severe dandruff, to search for lipase orthologues and identify the lipase that is the most frequently expressed on the scalp of patients with dandruff. The genome of M. restricta KCTC 27527 was sequenced using the Illumina MiSeq and PacBio platforms. Lipase orthologues were identified by comparison with known lipase genes in the genomes of Malassezia globosa and Malassezia sympodialis. The expression of the identified lipase genes was directly evaluated in swab samples from the scalps of 56 patients with dandruff. We found that, among the identified lipase-encoding genes, the gene encoding lipase homolog MRES_03670, named LIP5 in this study, was the most frequently expressed lipase in the swab samples. Our study provides an overview of the genome of a clinical isolate of M. restricta and fundamental information for elucidating the role of lipases during fungus-host interaction.© 2016 Blackwell Verlag GmbH.


July 7, 2019

Brassica rapa genome 2.0: a reference upgrade through sequence re-assembly and gene re-annotation.

Brassica rapa includes many important crops that are cultivated as vegetables, condiments, and oilseeds. Recently, the Brassica genomes have been sequenced extensively: a B. rapa draft reference genome in 2011 (Wang et al., 2011), a Brassica oleracea in 2014 (Liu et al., 2014), a Brassica napus in 2014 (Chalhoub et al., 2014), and Brassica nigra and Brassica juncea in 2016 (Yang et al., 2016). The first released B. rapa genome reference served as a valuable resource in the genome assembly and annotation of the other Brassicas (Chalhoub et al., 2014, Liu et al., 2014, Parkin et al., 2014). B. rapa has been used widely in Brassica comparative and evolutionary genomics among the Brassicaceae (Cheng et al., 2013). However, the first B. rapa genome assembly (version 1.5) is only about 283.8 Mb, 58.52% of the estimated genome size (485 Mb) (Wang et al., 2011). Considering that much of the genome assembly is still missing (41.48%), there is a considerable possibility that important genes have been missed.


July 7, 2019

Wild tobacco genomes reveal the evolution of nicotine biosynthesis.

Nicotine, the signature alkaloid of Nicotiana species responsible for the addictive properties of human tobacco smoking, functions as a defensive neurotoxin against attacking herbivores. However, the evolution of the genetic features that contributed to the assembly of the nicotine biosynthetic pathway remains unknown. We sequenced and assembled genomes of two wild tobaccos, Nicotiana attenuata (2.5 Gb) and Nicotiana obtusifolia (1.5 Gb), two ecological models for investigating adaptive traits in nature. We show that after the Solanaceae whole-genome triplication event, a repertoire of rapidly expanding transposable elements (TEs) bloated these Nicotiana genomes, promoted expression divergences among duplicated genes, and contributed to the evolution of herbivory-induced signaling and defenses, including nicotine biosynthesis. The biosynthetic machinery that allows for nicotine synthesis in the roots evolved from the stepwise duplications of two ancient primary metabolic pathways: the polyamine and nicotinamide adenine dinucleotide (NAD) pathways. In contrast to the duplication of the polyamine pathway that is shared among several solanaceous genera producing polyamine-derived tropane alkaloids, we found that lineage-specific duplications within the NAD pathway and the evolution of root-specific expression of the duplicated Solanaceae-specific ethylene response factor that activates the expression of all nicotine biosynthetic genes resulted in the innovative and efficient production of nicotine in the genus Nicotiana Transcription factor binding motifs derived from TEs may have contributed to the coexpression of nicotine biosynthetic pathway genes and coordinated the metabolic flux. Together, these results provide evidence that TEs and gene duplications facilitated the emergence of a key metabolic innovation relevant to plant fitness.


July 7, 2019

Genomic data mining of the marine actinobacteria Streptomyces sp. H-KF8 unveils insights into multi-stress related genes and metabolic pathways involved in antimicrobial synthesis.

Streptomyces sp. H-KF8 is an actinobacterial strain isolated from marine sediments of a Chilean Patagonian fjord. Morphological characterization together with antibacterial activity was assessed in various culture media, revealing a carbon-source dependent activity mainly against Gram-positive bacteria (S. aureus and L. monocytogenes). Genome mining of this antibacterial-producing bacterium revealed the presence of 26 biosynthetic gene clusters (BGCs) for secondary metabolites, where among them, 81% have low similarities with known BGCs. In addition, a genomic search in Streptomyces sp. H-KF8 unveiled the presence of a wide variety of genetic determinants related to heavy metal resistance (49 genes), oxidative stress (69 genes) and antibiotic resistance (97 genes). This study revealed that the marine-derived Streptomyces sp. H-KF8 bacterium has the capability to tolerate a diverse set of heavy metals such as copper, cobalt, mercury, chromate and nickel; as well as the highly toxic tellurite, a feature first time described for Streptomyces. In addition, Streptomyces sp. H-KF8 possesses a major resistance towards oxidative stress, in comparison to the soil reference strain Streptomyces violaceoruber A3(2). Moreover, Streptomyces sp. H-KF8 showed resistance to 88% of the antibiotics tested, indicating overall, a strong response to several abiotic stressors. The combination of these biological traits confirms the metabolic versatility of Streptomyces sp. H-KF8, a genetically well-prepared microorganism with the ability to confront the dynamics of the fjord-unique marine environment.


July 7, 2019

De novo hybrid assembly of the rubber tree genome reveals evidence of paleotetraploidy in Hevea species.

Para rubber tree (Hevea brasiliensis) is an important economic species as it is the sole commercial producer of high-quality natural rubber. Here, we report a de novo hybrid assembly of BPM24 accession, which exhibits resistance to major fungal pathogens in Southeast Asia. Deep-coverage 454/Illumina short-read and Pacific Biosciences (PacBio) long-read sequence data were acquired to generate a preliminary draft, which was subsequently scaffolded using a long-range “Chicago” technique to obtain a final assembly of 1.26?Gb (N50?=?96.8?kb). The assembled genome contains 69.2% repetitive sequences and has a GC content of 34.31%. Using a high-density SNP-based genetic map, we were able to anchor 28.9% of the genome assembly (363?Mb) associated with over two thirds of the predicted protein-coding genes into rubber tree’s 18 linkage groups. These genetically anchored sequences allowed comparative analyses of the intragenomic homeologous synteny, providing the first concrete evidence to demonstrate the presence of paleotetraploidy in Hevea species. Additionally, the degree of macrosynteny conservation observed between rubber tree and cassava strongly supports the hypothesis that the paleotetraploidization event took place prior to the divergence of the Hevea and Manihot species.


July 7, 2019

The genome sequence of Barbarea vulgaris facilitates the study of ecological biochemistry.

The genus Barbarea has emerged as a model for evolution and ecology of plant defense compounds, due to its unusual glucosinolate profile and production of saponins, unique to the Brassicaceae. One species, B. vulgaris, includes two ‘types’, G-type and P-type that differ in trichome density, and their glucosinolate and saponin profiles. A key difference is the stereochemistry of hydroxylation of their common phenethylglucosinolate backbone, leading to epimeric glucobarbarins. Here we report a draft genome sequence of the G-type, and re-sequencing of the P-type for comparison. This enables us to identify candidate genes underlying glucosinolate diversity, trichome density, and study the genetics of biochemical variation for glucosinolate and saponins. B. vulgaris is resistant to the diamondback moth, and may be exploited for “dead-end” trap cropping where glucosinolates stimulate oviposition and saponins deter larvae to the extent that they die. The B. vulgaris genome will promote the study of mechanisms in ecological biochemistry to benefit crop resistance breeding.


July 7, 2019

Genome scaffolding and annotation for the pathogen vector Ixodes ricinus by ultra-long single molecule sequencing.

Global warming and other ecological changes have facilitated the expansion of Ixodes ricinus tick populations. Ixodes ricinus is the most important carrier of vector-borne pathogens in Europe, transmitting viruses, protozoa and bacteria, in particular Borrelia burgdorferi (sensu lato), the causative agent of Lyme borreliosis, the most prevalent vector-borne disease in humans in the Northern hemisphere. To faster control this disease vector, a better understanding of the I. ricinus tick is necessary. To facilitate such studies, we recently published the first reference genome of this highly prevalent pathogen vector. Here, we further extend these studies by scaffolding and annotating the first reference genome by using ultra-long sequencing reads from third generation single molecule sequencing. In addition, we present the first genome size estimation for I. ricinus ticks and the embryo-derived cell line IRE/CTVM19.235,953 contigs were integrated into 204,904 scaffolds, extending the currently known genome lengths by more than 30% from 393 to 516 Mb and the N50 contig value by 87% from 1643 bp to a N50 scaffold value of 3067 bp. In addition, 25,263 sequences were annotated by comparison to the tick’s North American relative Ixodes scapularis. After (conserved) hypothetical proteins, zinc finger proteins, secreted proteins and P450 coding proteins were the most prevalent protein categories annotated. Interestingly, more than 50% of the amino acid sequences matching the homology threshold had 95-100% identity to the corresponding I. scapularis gene models. The sequence information was complemented by the first genome size estimation for this species. Flow cytometry-based genome size analysis revealed a haploid genome size of 2.65Gb for I. ricinus ticks and 3.80 Gb for the cell line.We present a first draft sequence map of the I. ricinus genome based on a PacBio-Illumina assembly. The I. ricinus genome was shown to be 26% (500 Mb) larger than the genome of its American relative I. scapularis. Based on the genome size of 2.65 Gb we estimated that we covered about 67% of the non-repetitive sequences. Genome annotation will facilitate screening for specific molecular pathways in I. ricinus cells and provides an overview of characteristics and functions.


July 7, 2019

Genome sequence of the filamentous actinomycete Kitasatospora viridifaciens.

The vast majority of antibiotics are produced by filamentous soil bacteria called actinomycetes. We report here the genome sequence of the tetracycline producer “Streptomyces viridifaciens” DSM 40239. Given that this species has the hallmark signatures characteristic of the Kitasatospora genus, we previously proposed to rename this organism Kitasatospora viridifaciens. Copyright © 2017 Ramijan et al.


July 7, 2019

Draft genome sequence of Halolamina pelagica CDK2 isolated from natural salterns from Rann of Kutch, Gujarat, India.

Halolamina pelagica strain CDK2, a halophilic archaeon (growth range 1.36 to 5.12 M NaCl), was isolated from rhizosphere of wild grasses of hypersaline soil of the Rann of Kutch, Gujarat, India. Its draft genome contains 2,972,542 bp and 3,485 coding sequences, depicting genes for halophilic serine proteases and trehalose synthesis. Copyright © 2017 Gaba et al.


July 7, 2019

Potential probiotic-associated traits revealed from completed high quality genome sequence of Lactobacillus fermentum 3872.

The article provides an overview of the genomic features of Lactobacillus fermentum strain 3872. The genomic sequence reported here is one of three L. fermentum genome sequences completed to date. Comparative genomic analysis allowed the identification of genes that may be contributing to enhanced probiotic properties of this strain. In particular, the genes encoding putative mucus binding proteins, collagen-binding proteins, class III bacteriocin, as well as exopolysaccharide and prophage-related genes were identified. Genes related to bacterial aggregation and survival under harsh conditions in the gastrointestinal tract, along with the genes required for vitamin production were also found.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.