June 1, 2021  |  

Complete resequencing of extended genomic regions using fosmid target capture and single molecule real-time (SMRT) long read sequencing technology.

A longstanding goal of genomic analysis is the identification of causal genetic factors contributing to disease. While the common disease/common variant hypothesis has been tested in many genome-wide association studies, few advancements in identifying causal variation have been realized, and instead recent findings point away from common variants towards aggregate rare variants as causal. A challenge is obtaining complete phased genomic sequences over extended genomic regions from sufficient numbers of cases and controls to identify all potential variation causal of a disease. To address this, we modified methods for targeted DNA isolation using fosmid technology and single-molecule, long-sequence-read generaton that combine for complete, haplotype-resolved resequencing across extended genomic subregions. As proof of principal, we validated the approach by resequencing four 800 kbp segments that span a major histocompatibility complex (MHC) common extended haplotype (CEH) associated with disease. The data revealed the extent of conservation exposing a near identity among four DR4 CEHs over conserved regions, detailing rare variation and measuring sequence accuracy. In a second test, we sequenced the complete KIR haplotypes from 8 individuals within a specific timeframe and cost. Single molecule long-read sequencing technology generated contiguous full-­length fosmid sequences of 30 to 40 kb in a single read, allowing assembly of resolved haplotypes with very little data processing. All of the sequences produced from these projects were contiguous, phased, with accuracy above 99.99%. The results demonstrated that cost-effective scale-­up is possible to generate scores to hundreds of phased chromosomal sequences of extended lengths that can encompass genomic regions associated with disease.

April 21, 2020  |  

The complete genome sequence and comparative genome analysis of the multi-drug resistant food-borne pathogen Bacillus cereus.

Bacillus cereus is an opportunistic human pathogen causing food-borne gastrointestinal infections and non-gastrointestinal infections worldwide. The strain B. cereus FORC_013 was isolated from fried eel. Its genome was completely sequenced by PacBio technology, analyzed and compared with other complete genome sequences of Bacillus to elucidate the distinct pathogenic features of the strain isolated in South Korea. Genomic analysis revealed pathogenesis and host immune evasion-associated genes encoding tissue-destructive exoenzymes, and pore-forming toxins. In particular, tissue-destructive (hemolysin BL, nonhaemolytic enterotoxins) and cytolytic proteins (cytolysin) were observed in the genome, which damage the plasma membrane of the epithelial cells of the small intestine causing diarrhea in humans. Capsule biosynthesis gene found in both chromosome and plasmid, which might be responsible for protecting the pathogen from the host cell immune defense system after host cell invasion. Additionally, multidrug resistance operon and efflux pumps were identified in the genome, which play a prominent role in multi-antibiotic resistance. Comparative phylogenetic tree analysis of the strain FORC_013 and other B. cereus strains revealed that the closest strains are ATCC 14579 and B4264. This genome data can be used to identify virulence factors that can be applied for the development of novel biomarkers for the rapid detection of this pathogen in foods.Copyright © 2018. Published by Elsevier Inc.

April 21, 2020  |  

Tracking short-term changes in the genetic diversity and antimicrobial resistance of OXA-232-producing Klebsiella pneumoniae ST14 in clinical settings.

To track stepwise changes in genetic diversity and antimicrobial resistance in rapidly evolving OXA-232-producing Klebsiella pneumoniae ST14, an emerging carbapenem-resistant high-risk clone, in clinical settings.Twenty-six K. pneumoniae ST14 isolates were collected by the Korean Nationwide Surveillance of Antimicrobial Resistance system over the course of 1 year. Isolates were subjected to whole-genome sequencing and MIC determinations using 33 antibiotics from 14 classes.Single-nucleotide polymorphism (SNP) typing identified 72 unique SNP sites spanning the chromosomes of the isolates, dividing them into three clusters (I, II and III). The initial isolate possessed two plasmids with 18 antibiotic-resistance genes, including blaOXA-232, and exhibited resistance to 11 antibiotic classes. Four other plasmids containing 12 different resistance genes, including blaCTX-M-15 and strA/B, were introduced over time, providing additional resistance to aztreonam and streptomycin. Moreover, chromosomal integration of insertion sequence Ecp1-blaCTX-M-15 mediated the inactivation of mgrB responsible for colistin resistance in four isolates from cluster III. To the best of our knowledge, this is the first description of K. pneumoniae ST14 resistant to both carbapenem and colistin in South Korea. Furthermore, although some acquired genes were lost over time, the retention of 12 resistance genes and inactivation of mgrB provided resistance to 13 classes of antibiotics.We describe stepwise changes in OXA-232-producing K. pneumoniae ST14 in vivo over time in terms of antimicrobial resistance. Our findings contribute to our understanding of the evolution of emerging high-risk K. pneumoniae clones and provide reference data for future outbreaks.Copyright © 2019 European Society of Clinical Microbiology and Infectious Diseases. Published by Elsevier Ltd. All rights reserved.

April 21, 2020  |  

Identification and characterization of chicken circovirus from commercial broiler chickens in China.

Circoviruses are found in many species, including mammals, birds, lower vertebrates and invertebrates. To date, there are no reports of circovirus-induced diseases in chickens. In this study, we identified a new strain of chicken circovirus (CCV) by PacBio third-generation sequencing samples from chickens with acute gastroenteritis in a Shandong commercial broiler farm in China. The complete genome of CCV was verified by inverse PCR. Genomic analysis revealed that CCV codes two inverse open reading frames (ORFs), and a potential stem-loop structure was present at the 5′ end with a structure typical of a circular virus. Phylogenetic tree analysis showed that CCV formed an independent branch between mammalian and avian circovirus, and homology analysis indicated that the homology of CCV with 21 other known circoviruses was less than 40%. Thus, this CCV strain represents a new species in the genus Circovirus. The infection rate of CCV in 12 chickens with diarrhoea was 100%, but no CCV was found in healthy chickens, thereby indicating that the novel CCV strain is highly associated with acute infectious gastroenteritis in chickens. The emergence of a novel CCV in commercial broiler chickens is highly concerning for the broiler industry. © 2019 Blackwell Verlag GmbH.

April 21, 2020  |  

Complete genome sequence of Paracoccus sp. Arc7-R13, a silver nanoparticles synthesizing bacterium isolated from Arctic Ocean sediments

Paracoccus sp. Arc7-R13, a silver nanoparticles (AgNPs) synthesizing bacterium, was isolated from Arctic Ocean sediment. Here we describe the complete genome of Paracoccus sp. Arc7-R13. The complete genome contains 4,040,012?bp with 66.66?mol%?G?+?C content, including one circular chromosome of 3,231,929?bp (67.45?mol%?G?+?C content), and eight plasmids with length ranging from 24,536?bp to 199,685?bp. The genome contains 3835 protein-coding genes (CDSs), 49 tRNA genes, as well as 3 rRNA operons as 16S-23S-5S rRNA. Based on the gene annotation and Swiss-Prot analysis, a total of 15 genes belonging to 11 kinds, including silver exporting P-type ATPase (SilP), alkaline phosphatase, nitroreductase, thioredoxin reductase, NADPH dehydrogenase and glutathione peroxidase, might be related to the synthesis of AgNPs. Meanwhile, many additional genes associated with synthesis of AgNPs such as protein-disulfide isomerase, c-type cytochrome, glutathione synthase and dehydrogenase reductase were also identified.

April 21, 2020  |  

Comparative Genomic Analysis of Virulence, Antimicrobial Resistance, and Plasmid Profiles of Salmonella Dublin Isolated from Sick Cattle, Retail Beef, and Humans in the United States.

Salmonella enterica serovar Dublin is a host-adapted serotype associated with typhoidal disease in cattle. While rare in humans, it usually causes severe illness, including bacteremia. In the United States, Salmonella Dublin has become one of the most multidrug-resistant (MDR) serotypes. To understand the genetic elements that are associated with virulence and resistance, we sequenced 61 isolates of Salmonella Dublin (49 from sick cattle and 12 from retail beef) using the Illumina MiSeq and closed 5 genomes using the PacBio sequencing platform. Genomic data of eight human isolates were also downloaded from NCBI (National Center for Biotechnology Information) for comparative analysis. Fifteen Salmonella pathogenicity islands (SPIs) and a spv operon (spvRABCD), which encodes important virulence factors, were identified in all 69 (100%) isolates. The 15 SPIs were located on the chromosome of the 5 closed genomes, with each of these isolates also carrying 1 or 2 plasmids with sizes between 36 and 329?kb. Multiple antimicrobial resistance genes (ARGs), including blaCMY-2, blaTEM-1B, aadA12, aph(3′)-Ia, aph(3′)-Ic, strA, strB, floR, sul1, sul2, and tet(A), along with spv operons were identified on these plasmids. Comprehensive antimicrobial resistance genotypes were determined, including 17 genes encoding resistance to 5 different classes of antimicrobials, and mutations in the housekeeping gene (gyrA) associated with resistance or decreased susceptibility to fluoroquinolones. Together these data revealed that this panel of Salmonella Dublin commonly carried 15 SPIs, MDR/virulence plasmids, and ARGs against several classes of antimicrobials. Such genomic elements may make important contributions to the severity of disease and treatment failures in Salmonella Dublin infections in both humans and cattle.

April 21, 2020  |  

Comparison of mitochondrial DNA variants detection using short- and long-read sequencing.

The recent advent of long-read sequencing technologies is expected to provide reasonable answers to genetic challenges unresolvable by short-read sequencing, primarily the inability to accurately study structural variations, copy number variations, and homologous repeats in complex parts of the genome. However, long-read sequencing comes along with higher rates of random short deletions and insertions, and single nucleotide errors. The relatively higher sequencing accuracy of short-read sequencing has kept it as the first choice of screening for single nucleotide variants and short deletions and insertions. Albeit, short-read sequencing still suffers from systematic errors that tend to occur at specific positions where a high depth of reads is not always capable to correct for these errors. In this study, we compared the genotyping of mitochondrial DNA variants in three samples using PacBio’s Sequel (Pacific Biosciences Inc., Menlo Park, CA, USA) long-read sequencing and illumina’s HiSeqX10 (illumine Inc., San Diego, CA, USA) short-read sequencing data. We concluded that, despite the differences in the type and frequency of errors in the long-reads sequencing, its accuracy is still comparable to that of short-reads for genotyping short nuclear variants; due to the randomness of errors in long reads, a lower coverage, around 37 reads, can be sufficient to correct for these random errors.

April 21, 2020  |  

Genomic analysis of Marinobacter sp. NP-4 and NP-6 isolated from the deep-sea oceanic crust on the western flank of the Mid-Atlantic Ridge

Two Marinobacter sp. NP-4 and NP-6 were isolated from a deep oceanic basaltic crust at North Pond, located at the western flank of the Mid-Atlantic Ridge. These two strains are capable of using multiple carbon sources such as acetate, succinate, glucose and sucrose while take oxygen as a primary electron acceptor. The strain NP-4 is also able to grow anaerobically under 20?MPa, with nitrate as the electron acceptor, thus represents a piezotolerant. To explore the metabolic potentials of Marinobacter sp. NP-4 and NP-6, the complete genome of NP-4 and close-to-complete genome of NP-6 were sequenced. The genome of NP-4 contains one chromosome and two plasmids with the size of 4.6?Mb in total, and with average GC content of 57.0%. The genome of NP-6 is 4.5?Mb and consists of 6 scaffolds, with an average GC content of 57.1%. Complete glycolysis, citrate cycle and aromatics compounds degradation pathways are identified in genomes of these two strains, suggesting that they possess a heterotrophic life style. Additionally, one plasmid of NP-4 contains genes for alkane degradation, phosphonate ABC transporter and cation efflux system, enabling NP-4 extra surviving abilities. In total, genomic information of these two strains provide insights into the physiological features and adaptation strategies of Marinobacter spp. in the deep oceanic crust biosphere.

April 21, 2020  |  

Variant Phasing and Haplotypic Expression from Single-molecule Long-read Sequencing in Maize

Haplotype phasing of genetic variants is important for interpretation of the maize genome, population genetic analysis, and functional genomic analysis of allelic activity. Accordingly, accurate methods for phasing full-length isoforms are essential for functional genomics study. In this study, we performed an isoform-level phasing study in maize, using two inbred lines and their reciprocal crosses, based on single-molecule full-length cDNA sequencing. To phase and analyze full-length transcripts between hybrids and parents, we developed a tool called IsoPhase. Using this tool, we validated the majority of SNPs called against matching short read data and identified cases of allele-specific, gene-level, and isoform-level expression. Our results revealed that maize parental and hybrid lines exhibit different splicing activities. After phasing 6,847 genes in two reciprocal hybrids using embryo, endosperm and root tissues, we annotated the SNPs and identified large-effect genes. In addition, based on single-molecule sequencing, we identified parent-of-origin isoforms in maize hybrids, different novel isoforms between maize parent and hybrid lines, and imprinted genes from different tissues. Finally, we characterized variation in cis- and trans-regulatory effects. Our study provides measures of haplotypic expression that could increase power and accuracy in studies of allelic expression.

