Menu
July 7, 2019

Microsatellite length scoring by Single Molecule Real Time Sequencing – Effects of sequence structure and PCR regime.

Microsatellites are DNA sequences consisting of repeated, short (1-6 bp) sequence motifs that are highly mutable by enzymatic slippage during replication. Due to their high intrinsic variability, microsatellites have important applications in population genetics, forensics, genome mapping, as well as cancer diagnostics and prognosis. The current analytical standard for microsatellites is based on length scoring by high precision electrophoresis, but due to increasing efficiency next-generation sequencing techniques may provide a viable alternative. Here, we evaluated single molecule real time (SMRT) sequencing, implemented in the PacBio series of sequencing apparatuses, as a means of microsatellite length scoring. To this end we carried out multiplexed SMRT sequencing of plasmid-carried artificial microsatellites of varying structure under different pre-sequencing PCR regimes. For each repeat structure, reads corresponding to the target length dominated. We found that pre-sequencing amplification had large effects on scoring accuracy and error distribution relative to controls, but that the effects of the number of amplification cycles were generally weak. In line with expectations enzymatic slippage decreased proportionally with microsatellite repeat unit length and increased with repetition number. Finally, we determined directional mutation trends, showing that PCR and SMRT sequencing introduced consistent but opposing error patterns in contraction and expansion of the microsatellites on the repeat motif and single nucleotide level.


July 7, 2019

Draft genome sequence of Ustilago trichophora RK089, a promising malic acid producer.

The basidiomycetous smut fungus Ustilago trichophora RK089 produces malate from glycerol. De novo genome sequencing revealed a 20.7-Mbp genome (301 gap-closed contigs, 246 scaffolds). A comparison to the genome of Ustilago maydis 521 revealed all essential genes for malate production from glycerol contributing to metabolic engineering for improving malate production. Copyright © 2016 Zambanini et al.


July 7, 2019

Structural basis for recombinatorial permissiveness in the generation of Anaplasma marginale Msp2 antigenic variants.

Sequential expression of outer membrane protein antigenic variants is an evolutionarily convergent mechanism used by bacterial pathogens to escape host immune clearance and establish persistent infection. Variants must be sufficiently structurally distinct to escape existing immune effectors yet retain core structural elements required for localization and function within the outer membrane. We examined this balance using Anaplasma marginale, which generates antigenic variants in the outer membrane protein Msp2 using gene conversion. The overwhelming majority of Msp2 variants expressed during long-term persistent infection are mosaics, derived by recombination of oligonucleotide segments from multiple alleles to form unique hypervariable regions (HVR). As a result, the mosaics are not under long-term selective pressure to encode a functional protein; consequently, we hypothesized that the Msp2 HVR is structurally permissive for mosaic expression. Using an integrated approach of predictive modeling with determination of native Msp2 protein structure and function, we demonstrate that structured elements, most notably ß-sheets, are significantly concentrated in the highly conserved N- and C-terminal domains. In contrast the HVR is overwhelmingly random coil with the structured a-helices and ß-sheets confined to the genomically defined “structural tethers” that separate the antigenically variable microdomains. This structure is supported by the surface exposure of the HVR microdomains and the slow diffusion type porin function in native Msp2. Importantly, the predominance of random coil provides plasticity for formation of functional HVR mosaics and realization of the full potential of segmental gene conversion to dramatically expand the variant repertoire. Copyright © 2016, American Society for Microbiology. All Rights Reserved.


July 7, 2019

The genomic sequence of the oral pathobiont strain NI1060 reveals unique strategies for bacterial competition and pathogenicity.

Strain NI1060 is an oral bacterium responsible for periodontitis in a murine ligature-induced disease model. To better understand its pathogenicity, we have determined the complete sequence of its 2,553,982 bp genome. Although closely related to Pasteurella pneumotropica, a pneumonia-associated rodent commensal based on its 16S rRNA, the NI1060 genomic content suggests that they are different species thriving on different energy sources via alternative metabolic pathways. Genomic and phylogenetic analyses showed that strain NI1060 is distinct from the genera currently described in the family Pasteurellaceae, and is likely to represent a novel species. In addition, we found putative virulence genes involved in lipooligosaccharide synthesis, adhesins and bacteriotoxic proteins. These genes are potentially important for host adaption and for the induction of dysbiosis through bacterial competition and pathogenicity. Importantly, strain NI1060 strongly stimulates Nod1, an innate immune receptor, but is defective in two peptidoglycan recycling genes due to a frameshift mutation. The in-depth analysis of its genome thus provides critical insights for the development of NI1060 as a prime model system for infectious disease.


July 7, 2019

Challenges, solutions, and quality metrics of personal genome assembly in advancing precision medicine.

Even though each of us shares more than 99% of the DNA sequences in our genome, there are millions of sequence codes or structure in small regions that differ between individuals, giving us different characteristics of appearance or responsiveness to medical treatments. Currently, genetic variants in diseased tissues, such as tumors, are uncovered by exploring the differences between the reference genome and the sequences detected in the diseased tissue. However, the public reference genome was derived with the DNA from multiple individuals. As a result of this, the reference genome is incomplete and may misrepresent the sequence variants of the general population. The more reliable solution is to compare sequences of diseased tissue with its own genome sequence derived from tissue in a normal state. As the price to sequence the human genome has dropped dramatically to around $1000, it shows a promising future of documenting the personal genome for every individual. However, de novo assembly of individual genomes at an affordable cost is still challenging. Thus, till now, only a few human genomes have been fully assembled. In this review, we introduce the history of human genome sequencing and the evolution of sequencing platforms, from Sanger sequencing to emerging “third generation sequencing” technologies. We present the currently available de novo assembly and post-assembly software packages for human genome assembly and their requirements for computational infrastructures. We recommend that a combined hybrid assembly with long and short reads would be a promising way to generate good quality human genome assemblies and specify parameters for the quality assessment of assembly outcomes. We provide a perspective view of the benefit of using personal genomes as references and suggestions for obtaining a quality personal genome. Finally, we discuss the usage of the personal genome in aiding vaccine design and development, monitoring host immune-response, tailoring drug therapy and detecting tumors. We believe the precision medicine would largely benefit from bioinformatics solutions, particularly for personal genome assembly.


July 7, 2019

Transposons passively and actively contribute to evolution of the two-speed genome of a fungal pathogen.

Genomic plasticity enables adaptation to changing environments, which is especially relevant for pathogens that engage in “arms races” with their hosts. In many pathogens, genes mediating virulence cluster in highly variable, transposon-rich, physically distinct genomic compartments. However, understanding of the evolution of these compartments, and the role of transposons therein, remains limited. Here, we show that transposons are the major driving force for adaptive genome evolution in the fungal plant pathogen Verticillium dahliae We show that highly variable lineage-specific (LS) regions evolved by genomic rearrangements that are mediated by erroneous double-strand repair, often utilizing transposons. We furthermore show that recent genetic duplications are enhanced in LS regions, against an older episode of duplication events. Finally, LS regions are enriched in active transposons, which contribute to local genome plasticity. Thus, we provide evidence for genome shaping by transposons, both in an active and passive manner, which impacts the evolution of pathogen virulence. © 2016 Faino et al.; Published by Cold Spring Harbor Laboratory Press.


July 7, 2019

Genomic characterization of the Atlantic cod sex-locus.

A variety of sex determination mechanisms can be observed in evolutionary divergent teleosts. Sex determination is genetic in Atlantic cod (Gadus morhua), however the genomic location or size of its sex-locus is unknown. Here, we characterize the sex-locus of Atlantic cod using whole genome sequence (WGS) data of 227 wild-caught specimens. Analyzing more than 55 million polymorphic loci, we identify 166 loci that are associated with sex. These loci are located in six distinct regions on five different linkage groups (LG) in the genome. The largest of these regions, an approximately 55?Kb region on LG11, contains the majority of genotypes that segregate closely according to a XX-XY system. Genotypes in this region can be used genetically determine sex, whereas those in the other regions are inconsistently sex-linked. The identified region on LG11 and its surrounding genes have no clear sequence homology with genes or regulatory elements associated with sex-determination or differentiation in other species. The functionality of this sex-locus therefore remains unknown. The WGS strategy used here proved adequate for detecting the small regions associated with sex in this species. Our results highlight the evolutionary flexibility in genomic architecture underlying teleost sex-determination and allow practical applications to genetically sex Atlantic cod.


July 7, 2019

The two chromosomes of the mitochondrial genome of a sugarcane cultivar: assembly and recombination analysis using long PacBio reads.

Sugarcane accounts for a large portion of the worlds sugar production. Modern commercial cultivars are complex hybrids of S. officinarum and several other Saccharum species. Historical records identify New Guinea as the origin of S. officinarum and that a small number of plants originating from there were used to generate all modern commercial cultivars. The mitochondrial genome can be a useful way to identify the maternal origin of commercial cultivars. We have used the PacBio RSII to sequence and assemble the mitochondrial genome of a South East Asian commercial cultivar, known as Khon Kaen 3. The long read length of this sequencing technology allowed for the mitochondrial genome to be assembled into two distinct circular chromosomes with all repeat sequences spanned by individual reads. Comparison of five commercial hybrids, two S. officinarum and one S. spontaneum to our assembly reveals no structural rearrangements between our assembly, the commercial hybrids and an S. officinarum from New Guinea. The S. spontaneum, from India, and one sample of S. officinarum (unknown origin) are substantially rearranged and have a large number of homozygous variants. This supports the record that S. officinarum plants from New Guinea are the maternal source of all modern commercial hybrids.


July 7, 2019

Distinct Salmonella enteritidis lineages associated with enterocolitis in high-income settings and invasive disease in low-income settings.

An epidemiological paradox surrounds Salmonella enterica serovar Enteritidis. In high-income settings, it has been responsible for an epidemic of poultry-associated, self-limiting enterocolitis, whereas in sub-Saharan Africa it is a major cause of invasive nontyphoidal Salmonella disease, associated with high case fatality. By whole-genome sequence analysis of 675 isolates of S. Enteritidis from 45 countries, we show the existence of a global epidemic clade and two new clades of S. Enteritidis that are geographically restricted to distinct regions of Africa. The African isolates display genomic degradation, a novel prophage repertoire, and an expanded multidrug resistance plasmid. S. Enteritidis is a further example of a Salmonella serotype that displays niche plasticity, with distinct clades that enable it to become a prominent cause of gastroenteritis in association with the industrial production of eggs and of multidrug-resistant, bloodstream-invasive infection in Africa.


July 7, 2019

The complete chloroplast genome sequence of the medicinal plant Swertia mussotii using the PacBio RS II platform.

Swertia mussotii is an important medicinal plant that has great economic and medicinal value and is found on the Qinghai Tibetan Plateau. The complete chloroplast (cp) genome of S. mussotii is 153,431 bp in size, with a pair of inverted repeat (IR) regions of 25,761 bp each that separate an large single-copy (LSC) region of 83,567 bp and an a small single-copy (SSC) region of 18,342 bp. The S. mussotii cp genome encodes 84 protein-coding genes, 37 transfer RNA (tRNA) genes, and eight ribosomal RNA (rRNA) genes. The identity, number, and GC content of S. mussotii cp genes were similar to those in the genomes of other Gentianales species. Via analysis of the repeat structure, 11 forward repeats, eight palindromic repeats, and one reverse repeat were detected in the S. mussotii cp genome. There are 45 SSRs in the S. mussotii cp genome, the majority of which are mononucleotides found in all other Gentianales species. An entire cp genome comparison study of S. mussotii and two other species in Gentianaceae was conducted. The complete cp genome sequence provides intragenic information for the cp genetic engineering of this medicinal plant.


July 7, 2019

Genome sequence and annotation of Colletotrichum higginsianum, a causal agent of crucifer anthracnose disease.

Colletotrichum higginsianum is an ascomycete fungus causing anthracnose disease on numerous cultivated plants in the family Brassicaceae, as well as the model plant Arabidopsis thaliana We report an assembly of the nuclear genome and gene annotation of this pathogen, which was obtained using a combination of PacBio long-read sequencing and optical mapping. Copyright © 2016 Zampounis et al.


July 7, 2019

DBG2OLC: Efficient assembly of large genomes using long erroneous reads of the third generation sequencing technologies.

The highly anticipated transition from next generation sequencing (NGS) to third generation sequencing (3GS) has been difficult primarily due to high error rates and excessive sequencing cost. The high error rates make the assembly of long erroneous reads of large genomes challenging because existing software solutions are often overwhelmed by error correction tasks. Here we report a hybrid assembly approach that simultaneously utilizes NGS and 3GS data to address both issues. We gain advantages from three general and basic design principles: (i) Compact representation of the long reads leads to efficient alignments. (ii) Base-level errors can be skipped; structural errors need to be detected and corrected. (iii) Structurally correct 3GS reads are assembled and polished. In our implementation, preassembled NGS contigs are used to derive the compact representation of the long reads, motivating an algorithmic conversion from a de Bruijn graph to an overlap graph, the two major assembly paradigms. Moreover, since NGS and 3GS data can compensate for each other, our hybrid assembly approach reduces both of their sequencing requirements. Experiments show that our software is able to assemble mammalian-sized genomes orders of magnitude more quickly than existing methods without consuming a lot of memory, while saving about half of the sequencing cost.


July 7, 2019

New Delhi metallo-ß-lactamase-1-producing Klebsiella pneumoniae, Florida, USA(1).

New Delhi metallo-ß-lactamase (NDM)–producing Enterobacteriaceae have swiftly spread worldwide since an initial report in 2008 from a patient who had been transferred from India back home to Sweden (1). Epidemiologically, the global diffusion of NDM-1 producers has been associated with the Indian subcontinent and the Balkan region, which are considered the primary and secondary reservoirs of these pathogens, respectively (1). However, recent reports suggest that countries in the Middle East may constitute another potential reservoir for NDM-1 producers (1). More than 100 NDM-producing isolates have been reported in the United States, most of which were associated with recent travel from the Indian subcontinent (2,3). We report an NDM-1–producing Klebsiella pneumoniae strain that was recovered from a patient who had been transferred from Iran to a hospital in Florida, United States.


July 7, 2019

Complete genome sequence of Bacillus oceanisediminis 2691, a reservoir of heavy-metal resistance genes.

Ocean sediments are commonly subject to the pollution of various heavy metals. Intracellular heavy metal concentrations in marine microorganisms should be kept within allowable concentrations. Here, we report redundant heavy metal resistance related genes encoding heavy metal-sensing transcriptional regulators (i.e. cadC), heavy metal efflux pumps, and detoxifying enzymes in the complete genome sequence of Bacillus oceanisediminis 2691. By comparing CadC sequences of strain 2691 with those from other bacterial genomes, we demonstrated that each cadC gene located in the chromosome or plasmid of 2691 cells are similar to those of various near or distant microbes, which might shed light on evolutionary trajectories of redundant heavy metal resistance genes. In application aspects, these diverse heavy metal sensing genes can be harnessed as synthetic biological parts, modules, and devices for the development of heavy metal-specific biosensors. Heavy metal bioremediation technologies or platform cells can be also developed based on the marine genomic information of heavy metal resistance and/or detoxification genes in a bacterial isolate from ocean sediments. Copyright © 2016 Elsevier B.V. All rights reserved.


July 7, 2019

Genomic insight into the host-endosymbiont relationship of Endozoicomonas montiporae CL-33(T) with its coral host.

The bacterial genus Endozoicomonas was commonly detected in healthy corals in many coral-associated bacteria studies in the past decade. Although, it is likely to be a core member of coral microbiota, little is known about its ecological roles. To decipher potential interactions between bacteria and their coral hosts, we sequenced and investigated the first culturable endozoicomonal bacterium from coral, the E. montiporae CL-33(T). Its genome had potential sign of ongoing genome erosion and gene exchange with its host. Testosterone degradation and type III secretion system are commonly present in Endozoicomonas and may have roles to recognize and deliver effectors to their hosts. Moreover, genes of eukaryotic ephrin ligand B2 are present in its genome; presumably, this bacterium could move into coral cells via endocytosis after binding to coral’s Eph receptors. In addition, 7,8-dihydro-8-oxoguanine triphosphatase and isocitrate lyase are possible type III secretion effectors that might help coral to prevent mitochondrial dysfunction and promote gluconeogenesis, especially under stress conditions. Based on all these findings, we inferred that E. montiporae was a facultative endosymbiont that can recognize, translocate, communicate and modulate its coral host.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.