Menu
September 22, 2019

GMAP and GSNAP for genomic sequence alignment: enhancements to speed, accuracy, and functionality.

The programs GMAP and GSNAP, for aligning RNA-Seq and DNA-Seq datasets to genomes, have evolved along with advances in biological methodology to handle longer reads, larger volumes of data, and new types of biological assays. The genomic representation has been improved to include linear genomes that can compare sequences using single-instruction multiple-data (SIMD) instructions, compressed genomic hash tables with fast access using SIMD instructions, handling of large genomes with more than four billion bp, and enhanced suffix arrays (ESAs) with novel data structures for fast access. Improvements to the algorithms have included a greedy match-and-extend algorithm using suffix arrays, segment chaining using genomic hash tables, diagonalization using segmental hash tables, and nucleotide-level dynamic programming procedures that use SIMD instructions and eliminate the need for F-loop calculations. Enhancements to the functionality of the programs include standardization of indel positions, handling of ambiguous splicing, clipping and merging of overlapping paired-end reads, and alignments to circular chromosomes and alternate scaffolds. The programs have been adapted for use in pipelines by integrating their usage into R/Bioconductor packages such as gmapR and HTSeqGenie, and these pipelines have facilitated the discovery of numerous biological phenomena.


September 22, 2019

Direct chromosome-length haplotyping by single-cell sequencing.

Haplotypes are fundamental to fully characterize the diploid genome of an individual, yet methods to directly chart the unique genetic makeup of each parental chromosome are lacking. Here we introduce single-cell DNA template strand sequencing (Strand-seq) as a novel approach to phasing diploid genomes along the entire length of all chromosomes. We demonstrate this by building a complete haplotype for a HapMap individual (NA12878) at high accuracy (concordance 99.3%), without using generational information or statistical inference. By use of this approach, we mapped all meiotic recombination events in a family trio with high resolution (median range ~14 kb) and phased larger structural variants like deletions, indels, and balanced rearrangements like inversions. Lastly, the single-cell resolution of Strand-seq allowed us to observe loss of heterozygosity regions in a small number of cells, a significant advantage for studies of heterogeneous cell populations, such as cancer cells. We conclude that Strand-seq is a unique and powerful approach to completely phase individual genomes and map inheritance patterns in families, while preserving haplotype differences between single cells.© 2016 Porubský et al.; Published by Cold Spring Harbor Laboratory Press.


September 22, 2019

Accurate characterization of the IFITM locus using MiSeq and PacBio sequencing shows genetic variation in Galliformes.

Interferon inducible transmembrane (IFITM) proteins are effectors of the immune system widely characterized for their role in restricting infection by diverse enveloped and non-enveloped viruses. The chicken IFITM (chIFITM) genes are clustered on chromosome 5 and to date four genes have been annotated, namely chIFITM1, chIFITM3, chIFITM5 and chIFITM10. However, due to poor assembly of this locus in the Gallus Gallus v4 genome, accurate characterization has so far proven problematic. Recently, a new chicken reference genome assembly Gallus Gallus v5 was generated using Sanger, 454, Illumina and PacBio sequencing technologies identifying considerable differences in the chIFITM locus over the previous genome releases.We re-sequenced the locus using both Illumina MiSeq and PacBio RS II sequencing technologies and we mapped RNA-seq data from the European Nucleotide Archive (ENA) to this finalized chIFITM locus. Using SureSelect probes capture probes designed to the finalized chIFITM locus, we sequenced the locus of a different chicken breed, namely a White Leghorn, and a turkey.We confirmed the Gallus Gallus v5 consensus except for two insertions of 5 and 1 base pair within the chIFITM3 and B4GALNT4 genes, respectively, and a single base pair deletion within the B4GALNT4 gene. The pull down revealed a single amino acid substitution of A63V in the CIL domain of IFITM2 compared to Red Jungle fowl and 13, 13 and 11 differences between IFITM1, 2 and 3 of chickens and turkeys, respectively. RNA-seq shows chIFITM2 and chIFITM3 expression in numerous tissue types of different chicken breeds and avian cell lines, while the expression of the putative chIFITM1 is limited to the testis, caecum and ileum tissues.Locus resequencing using these capture probes and RNA-seq based expression analysis will allow the further characterization of genetic diversity within Galliformes.


September 22, 2019

Improved full-length killer cell immunoglobulin-like receptor transcript discovery in Mauritian cynomolgus macaques.

Killer cell immunoglobulin-like receptors (KIRs) modulate disease progression of pathogens including HIV, malaria, and hepatitis C. Cynomolgus and rhesus macaques are widely used as nonhuman primate models to study human pathogens, and so, considerable effort has been put into characterizing their KIR genetics. However, previous studies have relied on cDNA cloning and Sanger sequencing that lack the throughput of current sequencing platforms. In this study, we present a high throughput, full-length allele discovery method utilizing Pacific Biosciences circular consensus sequencing (CCS). We also describe a new approach to Macaque Exome Sequencing (MES) and the development of the Rhexome1.0, an adapted target capture reagent that includes macaque-specific capture probe sets. By using sequence reads generated by whole genome sequencing (WGS) and MES to inform primer design, we were able to increase the sensitivity of KIR allele discovery. We demonstrate this increased sensitivity by defining nine novel alleles within a cohort of Mauritian cynomolgus macaques (MCM), a geographically isolated population with restricted KIR genetics that was thought to be completely characterized. Finally, we describe an approach to genotyping KIRs directly from sequence reads generated using WGS/MES reads. The findings presented here expand our understanding of KIR genetics in MCM by associating new genes with all eight KIR haplotypes and demonstrating the existence of at least one KIR3DS gene associated with every haplotype.


September 22, 2019

De novo assembly of a Chinese soybean genome.

Soybean was domesticated in China and has become one of the most important oilseed crops. Due to bottlenecks in their introduction and dissemination, soybeans from different geographic areas exhibit extensive genetic diversity. Asia is the largest soybean market; therefore, a high-quality soybean reference genome from this area is critical for soybean research and breeding. Here, we report the de novo assembly and sequence analysis of a Chinese soybean genome for “Zhonghuang 13” by a combination of SMRT, Hi-C and optical mapping data. The assembled genome size is 1.025 Gb with a contig N50 of 3.46 Mb and a scaffold N50 of 51.87 Mb. Comparisons between this genome and the previously reported reference genome (cv. Williams 82) uncovered more than 250,000 structure variations. A total of 52,051 protein coding genes and 36,429 transposable elements were annotated for this genome, and a gene co-expression network including 39,967 genes was also established. This high quality Chinese soybean genome and its sequence analysis will provide valuable information for soybean improvement in the future.


September 22, 2019

KIR3DL01 upregulation on gut natural killer cells in response to SIV infection of KIR- and MHC class I-defined rhesus macaques.

Natural killer cells provide an important early defense against viral pathogens and are regulated in part by interactions between highly polymorphic killer-cell immunoglobulin-like receptors (KIRs) on NK cells and their MHC class I ligands on target cells. We previously identified MHC class I ligands for two rhesus macaque KIRs: KIR3DL01 recognizes Mamu-Bw4 molecules and KIR3DL05 recognizes Mamu-A1*002. To determine how these interactions influence NK cell responses, we infected KIR3DL01+ and KIR3DL05+ macaques with and without defined ligands for these receptors with SIVmac239, and monitored NK cell responses in peripheral blood and lymphoid tissues. NK cell responses in blood were broadly stimulated, as indicated by rapid increases in the CD16+ population during acute infection and sustained increases in the CD16+ and CD16-CD56- populations during chronic infection. Markers of proliferation (Ki-67), activation (CD69 & HLA-DR) and antiviral activity (CD107a & TNFa) were also widely expressed, but began to diverge during chronic infection, as reflected by sustained CD107a and TNFa upregulation by KIR3DL01+, but not by KIR3DL05+ NK cells. Significant increases in the frequency of KIR3DL01+ (but not KIR3DL05+) NK cells were also observed in tissues, particularly in the gut-associated lymphoid tissues, where this receptor was preferentially upregulated on CD56+ and CD16-CD56- subsets. These results reveal broad NK cell activation and dynamic changes in the phenotypic properties of NK cells in response to SIV infection, including the enrichment of KIR3DL01+ NK cells in tissues that support high levels of virus replication.


September 22, 2019

Long reads: their purpose and place.

In recent years long-read technologies have moved from being a niche and specialist field to a point of relative maturity likely to feature frequently in the genomic landscape. Analogous to next generation sequencing, the cost of sequencing using long-read technologies has materially dropped whilst the instrument throughput continues to increase. Together these changes present the prospect of sequencing large numbers of individuals with the aim of fully characterizing genomes at high resolution. In this article, we will endeavour to present an introduction to long-read technologies showing: what long reads are; how they are distinct from short reads; why long reads are useful and how they are being used. We will highlight the recent developments in this field, and the applications and potential of these technologies in medical research, and clinical diagnostics and therapeutics.


September 22, 2019

A new standard for crustacean genomes: The highly contiguous, annotated genome assembly of the clam shrimp Eulimnadia texana reveals HOX gene order and identifies the sex chromosome.

Vernal pool clam shrimp (Eulimnadia texana) are a promising model system due to their ease of lab culture, short generation time, modest sized genome, a somewhat rare stable androdioecious sex determination system, and a requirement to reproduce via desiccated diapaused eggs. We generated a highly contiguous genome assembly using 46× of PacBio long read data and 216× of Illumina short reads, and annotated using Illumina RNAseq obtained from adult males or hermaphrodites. Of the 120?Mb genome 85% is contained in the largest eight contigs, the smallest of which is 4.6?Mb. The assembly contains 98% of transcripts predicted via RNAseq. This assembly is qualitatively different from scaffolded Illumina assemblies: It is produced from long reads that contain sequence data along their entire length, and is thus gap free. The contiguity of the assembly allows us to order the HOX genes within the genome, identifying two loci that contain HOX gene orthologs, and which approximately maintain the order observed in other arthropods. We identified a partial duplication of the Antennapedia complex adjacent to the few genes homologous to the Bithorax locus. Because the sex chromosome of an androdioecious species is of special interest, we used existing allozyme and microsatellite markers to identify the E. texana sex chromosome, and find that it comprises nearly half of the genome of this species. Linkage patterns indicate that recombination is extremely rare and perhaps absent in hermaphrodites, and as a result the location of the sex determining locus will be difficult to refine using recombination mapping.© The Author 2017. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution.


September 22, 2019

Comparative genome and phenotypic analysis of three Clostridioides difficile strains isolated from a single patient provide insight into multiple infection of C. difficile.

Clostridioides difficile infections (CDI) have emerged over the past decade causing symptoms that range from mild, antibiotic-associated diarrhea (AAD) to life-threatening toxic megacolon. In this study, we describe a multiple and isochronal (mixed) CDI caused by the isolates DSM 27638, DSM 27639 and DSM 27640 that already initially showed different morphotypes on solid media.The three isolates belonging to the ribotypes (RT) 012 (DSM 27639) and 027 (DSM 27638 and DSM 27640) were phenotypically characterized and high quality closed genome sequences were generated. The genomes were compared with seven reference strains including three strains of the RT 027, two of the RT 017, and one of the RT 078 as well as a multi-resistant RT 012 strain. The analysis of horizontal gene transfer events revealed gene acquisition incidents that sort the strains within the time line of the spread of their RTs within Germany. We could show as well that horizontal gene transfer between the members of different RTs occurred within this multiple infection. In addition, acquisition and exchange of virulence-related features including antibiotic resistance genes were observed. Analysis of the two genomes assigned to RT 027 revealed three single nucleotide polymorphisms (SNPs) and apparently a regional genome modification within the flagellar switch that regulates the fli operon.Our findings show that (i) evolutionary events based on horizontal gene transfer occur within an ongoing CDI and contribute to the adaptation of the species by the introduction of new genes into the genomes, (ii) within a multiple infection of a single patient the exchange of genetic material was responsible for a much higher genome variation than the observed SNPs.


September 22, 2019

Aberration or analogy? The atypical plastomes of Geraniaceae

A number of plant groups have been proposed as ideal systems to explore plastid inheritance, plastome evolution and plastome-nuclear genome coevolution. Quick generation times and a compact nuclear genome in Arabidopsis thaliana, the relative ease of plastid isolation from Spinacia oleracea and the tractability of plastid transformation in Nicotiana tabacum are all desirable attributes in a model system; however, these and most other groups all lack novelty in terms of plastome structure and nucleotide sequence evolution. Contemporary sequencing and assembly technologies have facilitated analyses of atypical plastomes and, as predicted by early investigations, Geraniaceae plastomes have experienced unprecedented rearrangements relative to the canonical structure and exhibit remarkably high rates of synonymous and nonsynonymous nucleotide substitutions. While not the only lineage with unusual plastome features, likely no other group represents the array of aberrant phenomena recorded for the family. In this chapter, Geraniaceae plastomes will be discussed and, where possible, compared with other taxa.


September 22, 2019

Translating genomics into practice for real-time surveillance and response to carbapenemase-producing Enterobacteriaceae: evidence from a complex multi-institutional KPC outbreak.

Until recently, Klebsiella pneumoniae carbapenemase (KPC)-producing Enterobacteriaceae were rarely identified in Australia. Following an increase in the number of incident cases across the state of Victoria, we undertook a real-time combined genomic and epidemiological investigation. The scope of this study included identifying risk factors and routes of transmission, and investigating the utility of genomics to enhance traditional field epidemiology for informing management of established widespread outbreaks.All KPC-producing Enterobacteriaceae isolates referred to the state reference laboratory from 2012 onwards were included. Whole-genome sequencing was performed in parallel with a detailed descriptive epidemiological investigation of each case, using Illumina sequencing on each isolate. This was complemented with PacBio long-read sequencing on selected isolates to establish high-quality reference sequences and interrogate characteristics of KPC-encoding plasmids.Initial investigations indicated that the outbreak was widespread, with 86 KPC-producing Enterobacteriaceae isolates (K. pneumoniae 92%) identified from 35 different locations across metropolitan and rural Victoria between 2012 and 2015. Initial combined analyses of the epidemiological and genomic data resolved the outbreak into distinct nosocomial transmission networks, and identified healthcare facilities at the epicentre of KPC transmission. New cases were assigned to transmission networks in real-time, allowing focussed infection control efforts. PacBio sequencing confirmed a secondary transmission network arising from inter-species plasmid transmission. Insights from Bayesian transmission inference and analyses of within-host diversity informed the development of state-wide public health and infection control guidelines, including interventions such as an intensive approach to screening contacts following new case detection to minimise unrecognised colonisation.A real-time combined epidemiological and genomic investigation proved critical to identifying and defining multiple transmission networks of KPC Enterobacteriaceae, while data from either investigation alone were inconclusive. The investigation was fundamental to informing infection control measures in real-time and the development of state-wide public health guidelines on carbapenemase-producing Enterobacteriaceae surveillance and management.


September 22, 2019

Screening and genomic characterization of filamentous hemagglutinin-deficient Bordetella pertussis.

Despite high vaccine coverage, pertussis cases in the United States have increased over the last decade. Growing evidence suggests that disease resurgence results, in part, from genetic divergence of circulating strain populations away from vaccine references. The United States employs acellular vaccines exclusively, and current Bordetella pertussis isolates are predominantly deficient in at least one immunogen, pertactin (Prn). First detected in the United States retrospectively in a 1994 isolate, the rapid spread of Prn deficiency is likely vaccine driven, raising concerns about whether other acellular vaccine immunogens experience similar pressures, as further antigenic changes could potentially threaten vaccine efficacy. We developed an electrochemiluminescent antibody capture assay to monitor the production of the acellular vaccine immunogen filamentous hemagglutinin (Fha). Screening 722 U.S. surveillance isolates collected from 2010 to 2016 identified two that were both Prn and Fha deficient. Three additional Fha-deficient laboratory strains were also identified from a historic collection of 65 isolates dating back to 1935. Whole-genome sequencing of deficient isolates revealed putative, underlying genetic changes. Only four isolates harbored mutations to known genes involved in Fha production, highlighting the complexity of its regulation. The chromosomes of two Fha-deficient isolates included unexpected structural variation that did not appear to influence Fha production. Furthermore, insertion sequence disruption of fhaB was also detected in a previously identified pertussis toxin-deficient isolate that still produced normal levels of Fha. These results demonstrate the genetic potential for additional vaccine immunogen deficiency and underscore the importance of continued surveillance of circulating B. pertussis evolution in response to vaccine pressure. Copyright © 2018 American Society for Microbiology.


September 22, 2019

Genomes of 13 domesticated and wild rice relatives highlight genetic conservation, turnover and innovation across the genus Oryza.

The genus Oryza is a model system for the study of molecular evolution over time scales ranging from a few thousand to 15 million years. Using 13 reference genomes spanning the Oryza species tree, we show that despite few large-scale chromosomal rearrangements rapid species diversification is mirrored by lineage-specific emergence and turnover of many novel elements, including transposons, and potential new coding and noncoding genes. Our study resolves controversial areas of the Oryza phylogeny, showing a complex history of introgression among different chromosomes in the young ‘AA’ subclade containing the two domesticated species. This study highlights the prevalence of functionally coupled disease resistance genes and identifies many new haplotypes of potential use for future crop protection. Finally, this study marks a milestone in modern rice research with the release of a complete long-read assembly of IR 8 ‘Miracle Rice’, which relieved famine and drove the Green Revolution in Asia 50 years ago.


September 22, 2019

Comparative genomics reveals new single-nucleotide polymorphisms that can assist in identification of adherent-invasive Escherichia coli.

Adherent-invasive Escherichia coli (AIEC) have been involved in Crohn’s disease (CD). Currently, AIEC are identified by time-consuming techniques based on in vitro infection of cell lines to determine their ability to adhere to and invade intestinal epithelial cells as well as to survive and replicate within macrophages. Our aim was to find signature sequences that can be used to identify the AIEC pathotype. Comparative genomics was performed between three E. coli strain pairs, each pair comprised one AIEC and one non-AIEC with identical pulsotype, sequence type and virulence gene carriage. Genetic differences were further analysed in 22 AIEC and 28 non-AIEC isolated from CD patients and controls. The strain pairs showed similar genome structures, and no gene was specific to AIEC. Three single nucleotide polymorphisms displayed different nucleotide distributions between AIEC and non-AIEC, and four correlated with increased adhesion and/or invasion indices. Here, we present a classification algorithm based on the identification of three allelic variants that can predict the AIEC phenotype with 84% accuracy. Our study corroborates the absence of an AIEC-specific genetic marker distributed across all AIEC strains. Nonetheless, point mutations putatively involved in the AIEC phenotype can be used for the molecular identification of the AIEC pathotype.


September 22, 2019

Complete genome sequence of Bacillus velezensis 157 isolated from Eucommia ulmoides with pathogenic bacteria inhibiting and lignocellulolytic enzymes production by SSF.

Bacillus velezensis 157 was isolated from the bark of Eucommia ulmoides, and exhibited antagonistic activity against a broad spectrum of pathogenic bacteria and fungi. Moreover, B. velezensis 157 also showed various lignocellulolytic activities including cellulase, xylanase, a-amylase, and pectinase, which had the ability of using the agro-industrial waste (soybean meal, wheat bran, sugarcane bagasse, wheat straw, rice husk, maize flour and maize straw) under solid-state fermentation and obtained several industrially valuable enzymes. Soybean meal appeared to be the most efficient substrate for the single fermentation of B. velezensis 157. Highest yield of pectinase (19.15 ± 2.66 U g-1), cellulase (46.69 ± 1.19 U g-1) and amylase (2097.18 ± 15.28 U g-1) was achieved on untreated soybean meal. Highest yield of xylanase (22.35 ± 2.24 U g-1) was obtained on untreated wheat bran. Here, we report the complete genome sequence of the B. velezensis 157, composed of a circular 4,013,317 bp chromosome with 3789 coding genes and a G + C content of 46.41%, one circular 8439 bp plasmid and a G + C content of 40.32%. The genome contained a total of 8 candidate gene clusters (bacillaene, difficidin, macrolactin, butirosin, bacillibactin, bacilysin, fengycin and surfactin), and dedicates over 15.8% of the whole genome to synthesize secondary metabolite biosynthesis. In addition, the genes encoding enzymes involved in degradation of cellulose, xylan, lignin, starch, mannan, galactoside and arabinan were found in the B. velezensis 157 genome. Thus, the study of B. velezensis 157 broadened that B. velezensis can not only be used as biocontrol agents, but also has potentially a wide range of applications in lignocellulosic biomass conversion.


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.