Menu
April 21, 2020

Systematic analysis of dark and camouflaged genes reveals disease-relevant genes hiding in plain sight.

The human genome contains “dark” gene regions that cannot be adequately assembled or aligned using standard short-read sequencing technologies, preventing researchers from identifying mutations within these gene regions that may be relevant to human disease. Here, we identify regions with few mappable reads that we call dark by depth, and others that have ambiguous alignment, called camouflaged. We assess how well long-read or linked-read technologies resolve these regions.Based on standard whole-genome Illumina sequencing data, we identify 36,794 dark regions in 6054 gene bodies from pathways important to human health, development, and reproduction. Of these gene bodies, 8.7% are completely dark and 35.2% are =?5% dark. We identify dark regions that are present in protein-coding exons across 748 genes. Linked-read or long-read sequencing technologies from 10x Genomics, PacBio, and Oxford Nanopore Technologies reduce dark protein-coding regions to approximately 50.5%, 35.6%, and 9.6%, respectively. We present an algorithm to resolve most camouflaged regions and apply it to the Alzheimer’s Disease Sequencing Project. We rescue a rare ten-nucleotide frameshift deletion in CR1, a top Alzheimer’s disease gene, found in disease cases but not in controls.While we could not formally assess the association of the CR1 frameshift mutation with Alzheimer’s disease due to insufficient sample-size, we believe it merits investigating in a larger cohort. There remain thousands of potentially important genomic regions overlooked by short-read sequencing that are largely resolved by long-read technologies.


April 21, 2020

A high-quality de novo genome assembly from a single mosquito using PacBio sequencing

A high-quality reference genome is a fundamental resource for functional genetics, comparative genomics, and population genomics, and is increasingly important for conservation biology. PacBio Single Molecule, Real-Time (SMRT) sequencing generates long reads with uniform coverage and high consensus accuracy, making it a powerful technology for de novo genome assembly. Improvements in throughput and concomitant reductions in cost have made PacBio an attractive core technology for many large genome initiatives, however, relatively high DNA input requirements (~5 µg for standard library protocol) have placed PacBio out of reach for many projects on small organisms that have lower DNA content, or on projects with limited input DNA for other reasons. Here we present a high-quality de novo genome assembly from a single Anopheles coluzzii mosquito. A modified SMRTbell library construction protocol without DNA shearing and size selection was used to generate a SMRTbell library from just 100 ng of starting genomic DNA. The sample was run on the Sequel System with chemistry 3.0 and software v6.0, generating, on average, 25 Gb of sequence per SMRT Cell with 20 h movies, followed by diploid de novo genome assembly with FALCON-Unzip. The resulting curated assembly had high contiguity (contig N50 3.5 Mb) and completeness (more than 98% of conserved genes were present and full-length). In addition, this single-insect assembly now places 667 (>90%) of formerly unplaced genes into their appropriate chromosomal contexts in the AgamP4 PEST reference. We were also able to resolve maternal and paternal haplotypes for over 1/3 of the genome. By sequencing and assembling material from a single diploid individual, only two haplotypes were present, simplifying the assembly process compared to samples from multiple pooled individuals. The method presented here can be applied to samples with starting DNA amounts as low as 100 ng per 1 Gb genome size. This new low-input approach puts PacBio-based assemblies in reach for small highly heterozygous organisms that comprise much of the diversity of life.


April 21, 2020

Tandem-genotypes: robust detection of tandem repeat expansions from long DNA reads.

Tandemly repeated DNA is highly mutable and causes at least 31 diseases, but it is hard to detect pathogenic repeat expansions genome-wide. Here, we report robust detection of human repeat expansions from careful alignments of long but error-prone (PacBio and nanopore) reads to a reference genome. Our method is robust to systematic sequencing errors, inexact repeats with fuzzy boundaries, and low sequencing coverage. By comparing to healthy controls, we prioritize pathogenic expansions within the top 10 out of 700,000 tandem repeats in whole genome sequencing data. This may help to elucidate the many genetic diseases whose causes remain unknown.


April 21, 2020

Genome plasticity favours double chromosomal Tn4401b-blaKPC-2 transposon insertion in the Pseudomonas aeruginosa ST235 clone.

Pseudomonas aeruginosa Sequence Type 235 is a clone that possesses an extraordinary ability to acquire mobile genetic elements and has been associated with the spread of resistance genes, including genes that encode for carbapenemases. Here, we aim to characterize the genetic platforms involved in resistance dissemination in blaKPC-2-positive P. aeruginosa ST235 in Colombia.In a prospective surveillance study of infections in adult patients attended in five ICUs in five distant cities in Colombia, 58 isolates of P. aeruginosa were recovered, of which, 27 (46.6%) were resistant to carbapenems. The molecular analysis showed that 6 (22.2%) and 4 (14.8%) isolates harboured the blaVIM and blaKPC-2 genes, respectively. The four blaKPC-2-positive isolates showed a similar PFGE pulsotype and belonged to ST235. Complete genome sequencing of a representative ST235 isolate shows a unique chromosomal contig of 7097.241?bp with eight different resistance genes identified and five transposons: a Tn6162-like with ant(2?)-Ia, two Tn402-like with ant(3?)-Ia and blaOXA-2 and two Tn4401b with blaKPC-2. All transposons were inserted into the genomic islands. Interestingly, the two Tn4401b copies harbouring blaKPC-2 were adjacently inserted into a new genomic island (PAGI-17) with traces of a replicative transposition process. This double insertion was probably driven by several structural changes within the chromosomal region containing PAGI-17 in the ST235 background.This is the first report of a double Tn4401b chromosomal insertion in P. aeruginosa, just within a new genomic island (PAGI-17). This finding indicates once again the great genomic plasticity of this microorganism.


April 21, 2020

Full-length transcript sequencing and comparative transcriptomic analysis to evaluate the contribution of osmotic and ionic stress components towards salinity tolerance in the roots of cultivated alfalfa (Medicago sativa L.).

Alfalfa is the most extensively cultivated forage legume. Salinity is a major environmental factor that impacts on alfalfa’s productivity. However, little is known about the molecular mechanisms underlying alfalfa responses to salinity, especially the relative contribution of the two important components of osmotic and ionic stress.In this study, we constructed the first full-length transcriptome database for alfalfa root tips under continuous NaCl and mannitol treatments for 1, 3, 6, 12, and 24?h (three biological replicates for each time points, including the control group) via PacBio Iso-Seq. This resulted in the identification of 52,787 full-length transcripts, with an average length of 2551?bp. Global transcriptional changes in the same 33 stressed samples were then analyzed via BGISEQ-500 RNA-Seq. Totals of 8861 NaCl-regulated and 8016 mannitol-regulated differentially expressed genes (DEGs) were identified. Metabolic analyses revealed that these DEGs overlapped or diverged in the cascades of molecular networks involved in signal perception, signal transduction, transcriptional regulation, and antioxidative defense. Notably, several well characterized signalling pathways, such as CDPK, MAPK, CIPK, and PYL-PP2C-SnRK2, were shown to be involved in osmotic stress, while the SOS core pathway was activated by ionic stress. Moreover, the physiological shifts of catalase and peroxidase activity, glutathione and proline content were in accordance with dynamic transcript profiles of the relevant genes, indicating that antioxidative defense system plays critical roles in response to salinity stress.Overall, our study provides evidence that the response to salinity stress in alfalfa includes both osmotic and ionic components. The key osmotic and ionic stress-related genes are candidates for future studies as potential targets to improve resistance to salinity stress via genetic engineering.


April 21, 2020

Whole-genome sequencing of Klebsiella pneumoniae isolates to track strain progression in a single patient with recurrent urinary tract infection.

Klebsiella pneumoniae is an important uropathogen that increasingly harbors broad-spectrum antibiotic resistance determinants. Evidence suggests that some same-strain recurrences in women with frequent urinary tract infections (UTIs) may emanate from a persistent intravesicular reservoir. Our objective was to analyze K. pneumoniae isolates collected over weeks from multiple body sites of a single patient with recurrent UTI in order to track ordered strain progression across body sites, as has been employed across patients in outbreak settings. Whole-genome sequencing of 26 K. pneumoniae isolates was performed utilizing the Illumina platform. PacBio sequencing was used to create a refined reference genome of the original urinary isolate (TOP52). Sequence variation was evaluated by comparing the 26 isolate sequences to the reference genome sequence. Whole-genome sequencing of the K. pneumoniae isolates from six different body sites of this patient with recurrent UTI demonstrated 100% chromosomal sequence identity of the isolates, with only a small P2 plasmid deletion in a minority of isolates. No single nucleotide variants were detected. The complete absence of single-nucleotide variants from 26 K. pneumoniae isolates from multiple body sites collected over weeks from a patient with recurrent UTI suggests that, unlike in an outbreak situation with strains collected from numerous patients, other methods are necessary to discern strain progression within a single host over a relatively short time frame.


April 21, 2020

Genome sequence and transcriptomic profiles of a marine bacterium, Pseudoalteromonas agarivorans Hao 2018.

Members of the marine genus Pseudoalteromonas have attracted great interest because of their ability to produce a large number of biologically active substances. Here, we report the complete genome sequence of Pseudoalteromonas agarivorans Hao 2018, a strain isolated from an abalone breeding environment, using second-generation Illumina and third-generation PacBio sequencing technologies. Illumina sequencing offers high quality and short reads, while PacBio technology generates long reads. The scaffolds of the two platforms were assembled to yield a complete genome sequence that included two circular chromosomes and one circular plasmid. Transcriptomic data for Pseudoalteromonas were not available. We therefore collected comprehensive RNA-seq data using Illumina sequencing technology from a fermentation culture of P. agarivorans Hao 2018. Researchers studying the evolution, environmental adaptations and biotechnological applications of Pseudoalteromonas may benefit from our genomic and transcriptomic data to analyze the function and expression of genes of interest.


April 21, 2020

Origin and recent expansion of an endogenous gammaretroviral lineage in domestic and wild canids.

Vertebrate genomes contain a record of retroviruses that invaded the germlines of ancestral hosts and are passed to offspring as endogenous retroviruses (ERVs). ERVs can impact host function since they contain the necessary sequences for expression within the host. Dogs are an important system for the study of disease and evolution, yet no substantiated reports of infectious retroviruses in dogs exist. Here, we utilized Illumina whole genome sequence data to assess the origin and evolution of a recently active gammaretroviral lineage in domestic and wild canids.We identified numerous recently integrated loci of a canid-specific ERV-Fc sublineage within Canis, including 58 insertions that were absent from the reference assembly. Insertions were found throughout the dog genome including within and near gene models. By comparison of orthologous occupied sites, we characterized element prevalence across 332 genomes including all nine extant canid species, revealing evolutionary patterns of ERV-Fc segregation among species as well as subpopulations.Sequence analysis revealed common disruptive mutations, suggesting a predominant form of ERV-Fc spread by trans complementation of defective proviruses. ERV-Fc activity included multiple circulating variants that infected canid ancestors from the last 20 million to within 1.6 million years, with recent bursts of germline invasion in the sublineage leading to wolves and dogs.


April 21, 2020

Comprehensive analysis of full genome sequence and Bd-milRNA/target mRNAs to discover the mechanism of hypovirulence in Botryosphaeria dothidea strains on pear infection with BdCV1 and BdPV1

Pear ring rot disease, mainly caused by Botryosphaeria dothidea, is widespread in most pear and apple-growing regions. Mycoviruses are used for biocontrol, especially in fruit tree disease. BdCV1 (Botryosphaeria dothidea chrysovirus 1) and BdPV1 (Botryosphaeria dothidea partitivirus 1) influence the biological characteristics of B. dothidea strains. BdCV1 is a potential candidate for the control of fungal disease. Therefore, it is vital to explore interactions between B. dothidea and mycovirus to clarify the pathogenic mechanisms of B. dothidea and hypovirulence of B. dothidea in pear. A high-quality full-length genome sequence of the B. dothidea LW-Hubei isolate was obtained using Single Molecule Real-Time sequencing. It has high repeat sequence with 9.3% and DNA methylation existence in the genome. The 46.34?Mb genomes contained 14,091 predicted genes, which of 13,135 were annotated. B. dothidea was predicted to express 3833 secreted proteins. In bioinformatics analysis, 351 CAZy members, 552 transporters, 128 kinases, and 1096 proteins associated with plant-host interaction (PHI) were identified. RNA-silencing components including two endoribonuclease Dicer, four argonaute (Ago) and three RNA-dependent RNA polymerase (RdRp) molecules were identified and expressed in response to mycovirus infection. Horizontal transfer of the LW-C and LW-P strains indicated that BdCV1 induced host gene silencing in LW-C to suppress BdPV1 transmission. To investigate the role of RNA-silencing in B. dothidea defense, we constructed four small RNA libraries and sequenced B. dothidea micro-like RNAs (Bd-milRNAs) produced in response to BdCV1 and BdPV1 infection. Among these, 167 conserved and 68 candidate novel Bd-milRNAs were identified, of which 161 conserved and 20 novel Bd-milRNA were differentially expressed. WEGO analysis revealed involvement of the differentially expressed Bd-milRNA-targeted genes in metabolic process, catalytic activity, cell process and response to stress or stimulus. BdCV1 had a greater effect on the phenotype, virulence, conidiomata, vertical and horizontal transmission ability, and mycelia cellular structure biological characteristics of B. dothidea strains than BdPV1 and virus-free strains. The results obtained in this study indicate that mycovirus regulates biological processes in B. dothidea through the combined interaction of antiviral defense mediated by RNA-silencing and milRNA-mediated regulation of target gene mRNA expression.


January 23, 2017

Tutorial: Long Amplicon Analysis application

This tutorial provides an overview of the Long Amplicon Analysis (LAA) application. The LAA algorithm generates highly accurate, phased and full-length consensus sequences from long amplicons. Applications of LAA include…


January 23, 2017

Tutorial: Iso-Seq analysis application

This tutorial provides an overview of the Isoform Sequencing (Iso-Seq) analysis application. The Iso-Seq application provides reads that span entire transcript isoforms, from the 5′ end to the 3′ polyA-tail….


January 23, 2017

Tutorial: HGAP4 de novo assembly application

This tutorial provides an overview of the Hierarchical Genome Assembly Process (HGAP4) de novo assembly analysis application. HGAP4 generates accurate de novo assemblies using only PacBio data. HGAP4 is suitable…


Talk with an expert

If you have a question, need to check the status of an order, or are interested in purchasing an instrument, we're here to help.