Sequencing platform Archives

June 1, 2021

Genome sequencing of endosymbiotic bacterial Streptomyces sp. from Antartic lichen using Single Molecule Real-time Sequencing (SMRT) technology.

Along with the advent of next-generation sequencing (NGS) techniques, it has become possible to sequence a microbial genome very quickly with high coverage. Recently, PacificBioscience developed single molecule real-time sequencing (SMRT) technology, 3rd generation sequencing platform, which provide much longer (average read length: 1.5Kb) reads without PCR amplification. We did de novo sequencing of Streptomyces sp. using Illumina GAIIx, Roche 454 and PacBio RS system and compared the data. The endosymbiotic bacteria Streptomyces sp. PAMC 26508 was isolated from Antarctic lichen Psoroma sp. that grows attached rocks on Barton Peninsula, King George Island, Antarctica (62, 13’S, 58, 47’W). With 4 SMRT cells, we could get more than 15x coverage of corrected sequence data for de novo assembly. Comparing the performance of other sequencing platforms, PacBio platform could generate data on similar manner with general mid-level GC content organism. In conclusion, PacBio RS system, SMRT technology, shows better performance with high GC content organisms and is expected to be the new tool to improve the de novo sequencing and assembly.

June 1, 2021

Advances in sequence consensus and clustering algorithms for effective de novo assembly and haplotyping applications.

One of the major applications of DNA sequencing technology is to bring together information that is distant in sequence space so that understanding genome structure and function becomes easier on a large scale. The Single Molecule Real Time (SMRT) Sequencing platform provides direct sequencing data that can span several thousand bases to tens of thousands of bases in a high-throughput fashion. In contrast to solving genomic puzzles by patching together smaller piece of information, long sequence reads can decrease potential computation complexity by reducing combinatorial factors significantly. We demonstrate algorithmic approaches to construct accurate consensus when the differences between reads are dominated by insertions and deletions. High-performance implementations of such algorithms allow more efficient de novo assembly with a pre-assembly step that generates highly accurate, consensus-based reads which can be used as input for existing genome assemblers. In contrast to recent hybrid assembly approach, only a single ~10 kb or longer SMRTbell library is necessary for the hierarchical genome assembly process (HGAP). Meanwhile, with a sensitive read-clustering algorithm with the consensus algorithms, one is able to discern haplotypes that differ by less than 1% different from each other over a large region. One of the related applications is to generate accurate haplotype sequences for HLA loci. Long sequence reads that can cover the whole 3 kb to 4 kb diploid genomic regions will simplify the haplotyping process. These algorithms can also be applied to resolve individual populations within mixed pools of DNA molecules that are similar to each, e.g., by sequencing viral quasi-species samples.

June 1, 2021

Complex alternative splicing patterns in hematopoietic cell subpopulations revealed by third-generation long reads.

Background: Alternative splicing expands the repertoire of gene functions and is a signature for different cell populations. Here we characterize the transcriptome of human bone marrow subpopulations including progenitor cells to understand their contribution to homeostasis and pathological conditions such as atherosclerosis and tumor metastasis. To obtain full-length transcript structures, we utilized long reads in addition to RNA-seq for estimating isoform diversity and abundance. Method: Freshly harvested, viable human bone marrow tissues were extracted from discarded harvesting equipment and separated into total bone marrow (total), lineage-negative (lin-) progenitor cells and differentiated cells (lin+) by magnetic bead sorting with antibodies to surface markers of hematopoietic cell lineages. Sequencing was done with SOLiD, Illumina HiSeq (100bp paired-end reads), and PacBio RS II (full-length cDNA library protocol for 1 – 6 kb libraries). Short reads were assembled using both Trinity for de novo assembly and Cufflinks for genome-guided assembly. Full-length transcript consensus sequences were obtained for the PacBio data using the RS_IsoSeq protocol from PacBios SMRTAnalysis software. Quantitation for each sample was done independently for each sequencing platform using Sailfish to obtain the TPM (transcripts per million) using k-mer matching. Results: PacBios long read sequencing technology is capable of sequencing full-length transcripts up to 10 kb and reveals heretofore-unseen isoform diversity and complexity within the hematopoietic cell populations. A comparison of sequencing depth and de novo transcript assembly with short read, second-generation sequencing reveals that, while short reads provide precision in determining portions of isoform structure and supporting larger 5 and 3 UTR regions, it fails in providing a complete structure especially when multiple isoforms are present at the same locus. Increased breadth of isoform complexity is revealed by long reads that permits further elaboration of full isoform diversity and specific isoform abundance within each separate cell population. Sorting the distribution of major and minor isoforms reveals a cell population-specific balance focused on distinct genome loci and shows how tissue specificity and diversity are modulated by alternative splicing.

June 1, 2021

Targeted SMRT Sequencing and phasing using Roche NimbleGen’s SeqCap EZ enrichment

As a cost-effective alternative to whole genome human sequencing, targeted sequencing of specific regions, such as exomes or panels of relevant genes, has become increasingly common. These methods typically include direct PCR amplification of the genomic DNA of interest, or the capture of these targets via probe-based hybridization. Commonly, these approaches are designed to amplify or capture exonic regions and thereby result in amplicons or fragments that are a few hundred base pairs in length, a length that is well-addressed with short-read sequencing technologies. These approaches typically provide very good coverage and can identify SNPs in the targeted region, but are unable to haplotype these variants. Here we describe a targeted sequencing workflow that combines Roche NimbleGen’s SeqCap EZ enrichment technology with Pacific Biosciences’ SMRT Sequencing to provide a more comprehensive view of variants and haplotype information over multi-kilobase regions. While the SeqCap EZ technology is typically used to capture 200 bp fragments, we demonstrate that 6 kb fragments can also be utilized to enrich for long fragments that extend beyond the targeted capture site and well into (and often across) the flanking intronic regions. When combined with the long reads of SMRT Sequencing, multi-kilobase regions of the human genome can be phased and variants detected in exons, introns and intergenic regions.

June 1, 2021

Structural variation with Pacific Biosciences long reads

2015 SMRT Informatics Developers Conference Presentation Slides: Ali Bashir of Mount Sinai School of Medicine discussed methods for characterizing structural variation in human genomes across a variety of coverage levels.

June 1, 2021

Targeted sequencing of genes from soybean using NimbleGen SeqCap EZ and PacBio SMRT Sequencing

Full-length gene capture solutions offer opportunities to screen and characterize structural variations and genetic diversity to understand key traits in plants and animals. Through a combined Roche NimbleGen probe capture and SMRT Sequencing strategy, we demonstrate the capability to resolve complex gene structures often observed in plant defense and developmental genes spanning multiple kilobases. The custom panel includes members of the WRKY plant-defense-signaling family, members of the NB-LRR disease-resistance family, and developmental genes important for flowering. The presence of repetitive structures and low-complexity regions makes short-read sequencing of these genes difficult, yet this approach allows researchers to obtain complete sequences for unambiguous resolution of gene models. This strategy has been applied to genomic DNA samples from soybean coupled with barcoding for multiplexing.

June 1, 2021

Target enrichment using a neurology panel for 12 barcoded genomic DNA samples on the PacBio SMRT Sequencing platform

Target enrichment is a powerful tool for studies involved in understanding polymorphic SNPs with phasing, tandem repeats, and structural variations. With increasing availability of reference genomes, researchers can easily design a cost-effective targeted investigation with custom probes specific to regions of interest. Using PacBio long-read technology in conjunction with probe capture, we were able to sequence multi-kilobase enriched regions to fully investigate intronic and exonic regions, distinguish haplotypes, and characterize structural variations. Furthermore, we demonstrate this approach is advantageous for studying complex genomic regions previously inaccessible through other sequencing platforms. In the present work, 12 barcoded genomic DNA (gDNA) samples were sheared to 6 kb for target enrichment analysis using the Neurology panel provided by Roche NimbleGen. Probe-captured DNA was used to make SMRTbell libraries for SMRT Sequencing on the PacBio RS II. Our results demonstrate the ability to multiplex 12 samples and achieve 1300x enrichment of targeted regions. In addition, we achieved an even representation of on-target rate of 70% across the 12 barcoded genomic DNA samples.

June 1, 2021

“SMRTer Confirmation”: Scalable clinical read-through variant confirmation using the Pacific Biosciences SMRT Sequencing platform

Next-generation sequencing (NGS) has significantly improved the cost and turnaround time for diagnostic genetic tests. ACMG recommends variant confirmation by an orthogonal method, unless sufficiently high sensitivity and specificity can be demonstrated using NGS alone. Most NGS laboratories make extensive use of Sanger sequencing for secondary confirmation of single nucleotide variants (SNVs) and indels, representing a large fraction of the cost and time required to deliver high quality genetic testing data to clinicians and patients. Despite its established data quality, Sanger is not a high-throughput method by today’s standards from either an assay or analysis standpoint as it can involve manual review of Sanger traces and is not amenable to multiplexing. Toward a scalable solution for confirmation, Invitae has developed a fully automated and LIMS-tracked assay and informatics pipeline that utilizes the Pacific Biosciences SMRT sequencing platform. Invitae’s pipeline generates PCR amplicons that encompass the variant(s) of interest, which are converted to closed DNA structures (SMRTbells) and sequenced in pools of 96 per SMRTcell. Each amplicon is appended with a 16nt barcode that encodes the patient and variant IDs. Per-sample de-multiplexing, alignment, variant calling, and confirmation resolution are handled via an automated pipeline. The confirmation process was validated by analyzing 243 clinical SNVs and indels in parallel with the gold standard Sanger sequencing method. Amplicons were sequenced and analyzed in technical replicates to demonstrate reproducibility. In this study, the PacBio-based confirmation pipeline demonstrated high reproducibility (97.5%), and outperformed Sanger in the fraction of primary NGS variants confirmed (PacBio = 93.4% and 94.7% confirmed across two replicates, Sanger = 84.8%) while having 100% concordance of confirmation status among overlapping confirmation calls.

February 5, 2021

Podcast: Going beyond the $1,000 genome with Mark Gerstein

Mark Gerstein is the co-director of the Yale Computational Biology and Bioinformatics program where he focuses on better annotation of the human genome and better ways to mine big genomics…

April 21, 2020

Long-read sequencing for rare human genetic diseases.

During the past decade, the search for pathogenic mutations in rare human genetic diseases has involved huge efforts to sequence coding regions, or the entire genome, using massively parallel short-read sequencers. However, the approximate current diagnostic rate is <50% using these approaches, and there remain many rare genetic diseases with unknown cause. There may be many reasons for this, but one plausible explanation is that the responsible mutations are in regions of the genome that are difficult to sequence using conventional technologies (e.g., tandem-repeat expansion or complex chromosomal structural aberrations). Despite the drawbacks of high cost and a shortage of standard analytical methods, several studies have analyzed pathogenic changes in the genome using long-read sequencers. The results of these studies provide hope that further application of long-read sequencers to identify the causative mutations in unsolved genetic diseases may expand our understanding of the human genome and diseases. Such approaches may also be applied to molecular diagnosis and therapeutic strategies for patients with genetic diseases in the future.

April 21, 2020

Whole-genome sequence of Arthrinium phaeospermum, a globally distributed pathogenic fungus.

Arthrinium phaeospermum (Corda) M.B. Ellis is a globally distributed pathogenic fungus with a wide host range; its hosts include not only plants, but also humans and animals. This study aimed to develop genomic resources for A. phaeospermum to provide solid data and a theoretical basis for further studies of its pathogenesis, transcriptomics, proteomics, metabolomics and RNA genomics. The genome was obtained from the mycelia of the strain AP-Z13 using a combination of analyses with the high-throughput Illumina HiSeq 4000 system and PacBio RSII LongRead sequencing platform. Functional annotation was performed by BLASTing protein sequences against those in different publicly available databases to obtain their corresponding annotations. The genome is 48.45?Mb in size, with an N90 scaffold size of 1,931,147?bp, and encodes 19,836 putative predicted genes. This is the first report of the genome-scale assembly and annotation for A. phaeospermum, the first species in the genus Arthrinium to be subjected to whole genome sequencing. Copyright © 2019 Elsevier Inc. All rights reserved.

April 21, 2020

Complete genome sequence of Paracoccus sp. Arc7-R13, a silver nanoparticles synthesizing bacterium isolated from Arctic Ocean sediments

Paracoccus sp. Arc7-R13, a silver nanoparticles (AgNPs) synthesizing bacterium, was isolated from Arctic Ocean sediment. Here we describe the complete genome of Paracoccus sp. Arc7-R13. The complete genome contains 4,040,012?bp with 66.66?mol%?G?+?C content, including one circular chromosome of 3,231,929?bp (67.45?mol%?G?+?C content), and eight plasmids with length ranging from 24,536?bp to 199,685?bp. The genome contains 3835 protein-coding genes (CDSs), 49 tRNA genes, as well as 3 rRNA operons as 16S-23S-5S rRNA. Based on the gene annotation and Swiss-Prot analysis, a total of 15 genes belonging to 11 kinds, including silver exporting P-type ATPase (SilP), alkaline phosphatase, nitroreductase, thioredoxin reductase, NADPH dehydrogenase and glutathione peroxidase, might be related to the synthesis of AgNPs. Meanwhile, many additional genes associated with synthesis of AgNPs such as protein-disulfide isomerase, c-type cytochrome, glutathione synthase and dehydrogenase reductase were also identified.

April 21, 2020

Comparative Genomic Analysis of Virulence, Antimicrobial Resistance, and Plasmid Profiles of Salmonella Dublin Isolated from Sick Cattle, Retail Beef, and Humans in the United States.

Salmonella enterica serovar Dublin is a host-adapted serotype associated with typhoidal disease in cattle. While rare in humans, it usually causes severe illness, including bacteremia. In the United States, Salmonella Dublin has become one of the most multidrug-resistant (MDR) serotypes. To understand the genetic elements that are associated with virulence and resistance, we sequenced 61 isolates of Salmonella Dublin (49 from sick cattle and 12 from retail beef) using the Illumina MiSeq and closed 5 genomes using the PacBio sequencing platform. Genomic data of eight human isolates were also downloaded from NCBI (National Center for Biotechnology Information) for comparative analysis. Fifteen Salmonella pathogenicity islands (SPIs) and a spv operon (spvRABCD), which encodes important virulence factors, were identified in all 69 (100%) isolates. The 15 SPIs were located on the chromosome of the 5 closed genomes, with each of these isolates also carrying 1 or 2 plasmids with sizes between 36 and 329?kb. Multiple antimicrobial resistance genes (ARGs), including blaCMY-2, blaTEM-1B, aadA12, aph(3′)-Ia, aph(3′)-Ic, strA, strB, floR, sul1, sul2, and tet(A), along with spv operons were identified on these plasmids. Comprehensive antimicrobial resistance genotypes were determined, including 17 genes encoding resistance to 5 different classes of antimicrobials, and mutations in the housekeeping gene (gyrA) associated with resistance or decreased susceptibility to fluoroquinolones. Together these data revealed that this panel of Salmonella Dublin commonly carried 15 SPIs, MDR/virulence plasmids, and ARGs against several classes of antimicrobials. Such genomic elements may make important contributions to the severity of disease and treatment failures in Salmonella Dublin infections in both humans and cattle.

April 21, 2020

Antibiotic susceptibility of plant-derived lactic acid bacteria conferring health benefits to human.

Lactic acid bacteria (LAB) confer health benefits to human when administered orally. We have recently isolated several species of LAB strains from plant sources, such as fruits, vegetables, flowers, and medicinal plants. Since antibiotics used to treat bacterial infection diseases induce the emergence of drug-resistant bacteria in intestinal microflora, it is important to evaluate the susceptibility of LAB strains to antibiotics to ensure the safety and security of processed foods. The aim of the present study is to determine the minimum inhibitory concentration (MIC) of antibiotics against several plant-derived LAB strains. When aminoglycoside antibiotics, such as streptomycin (SM), kanamycin (KM), and gentamicin (GM), were evaluated using LAB susceptibility test medium (LSM), the MIC was higher than when using Mueller-Hinton (MH) medium. Etest, which is an antibiotic susceptibility assay method consisting of a predefined gradient of antibiotic concentrations on a plastic strip, is used to determine the MIC of antibiotics world-wide. In the present study, we demonstrated that Etest was particularly valuable while testing LAB strains. We also show that the low susceptibility of the plant-derived LAB strains against each antibiotic tested is due to intrinsic resistance and not acquired resistance. This finding is based on the whole-genome sequence information reflecting the horizontal spread of the drug-resistance genes in the LAB strains.

April 21, 2020

RNA sequencing: the teenage years.

Over the past decade, RNA sequencing (RNA-seq) has become an indispensable tool for transcriptome-wide analysis of differential gene expression and differential splicing of mRNAs. However, as next-generation sequencing technologies have developed, so too has RNA-seq. Now, RNA-seq methods are available for studying many different aspects of RNA biology, including single-cell gene expression, translation (the translatome) and RNA structure (the structurome). Exciting new applications are being explored, such as spatial transcriptomics (spatialomics). Together with new long-read and direct RNA-seq technologies and better computational tools for data analysis, innovations in RNA-seq are contributing to a fuller understanding of RNA biology, from questions such as when and where transcription occurs to the folding and intermolecular interactions that govern RNA function.

Auto Tag: Sequencing platform

Genome sequencing of endosymbiotic bacterial Streptomyces sp. from Antartic lichen using Single Molecule Real-time Sequencing (SMRT) technology.

Advances in sequence consensus and clustering algorithms for effective de novo assembly and haplotyping applications.

Complex alternative splicing patterns in hematopoietic cell subpopulations revealed by third-generation long reads.

Targeted SMRT Sequencing and phasing using Roche NimbleGen’s SeqCap EZ enrichment

Structural variation with Pacific Biosciences long reads

Targeted sequencing of genes from soybean using NimbleGen SeqCap EZ and PacBio SMRT Sequencing

Target enrichment using a neurology panel for 12 barcoded genomic DNA samples on the PacBio SMRT Sequencing platform

“SMRTer Confirmation”: Scalable clinical read-through variant confirmation using the Pacific Biosciences SMRT Sequencing platform

Podcast: Going beyond the $1,000 genome with Mark Gerstein

Long-read sequencing for rare human genetic diseases.

Whole-genome sequence of Arthrinium phaeospermum, a globally distributed pathogenic fungus.

Complete genome sequence of Paracoccus sp. Arc7-R13, a silver nanoparticles synthesizing bacterium isolated from Arctic Ocean sediments

Comparative Genomic Analysis of Virulence, Antimicrobial Resistance, and Plasmid Profiles of Salmonella Dublin Isolated from Sick Cattle, Retail Beef, and Humans in the United States.

Antibiotic susceptibility of plant-derived lactic acid bacteria conferring health benefits to human.

RNA sequencing: the teenage years.

Subscribe for blog updates:

Filter by topic

Talk with an expert

ALS case study

Subscribe for blog updates:

Filter by topic

Talk with an expert