protein-coding Archives

April 21, 2020 |

Improved assembly and variant detection of a haploid human genome using single-molecule, high-fidelity long reads.

The sequence and assembly of human genomes using long-read sequencing technologies has revolutionized our understanding of structural variation and genome organization. We compared the accuracy, continuity, and gene annotation of genome assemblies generated from either high-fidelity (HiFi) or continuous long-read (CLR) datasets from the same complete hydatidiform mole human genome. We find that the HiFi sequence data assemble an additional 10% of duplicated regions and more accurately represent the structure of tandem repeats, as validated with orthogonal analyses. As a result, an additional 5 Mbp of pericentromeric sequences are recovered in the HiFi assembly, resulting in a 2.5-fold increase in the NG50 within 1 Mbp of the centromere (HiFi 480.6 kbp, CLR 191.5 kbp). Additionally, the HiFi genome assembly was generated in significantly less time with fewer computational resources than the CLR assembly. Although the HiFi assembly has significantly improved continuity and accuracy in many complex regions of the genome, it still falls short of the assembly of centromeric DNA and the largest regions of segmental duplication using existing assemblers. Despite these shortcomings, our results suggest that HiFi may be the most effective standalone technology for de novo assembly of human genomes. © 2019 John Wiley & Sons Ltd/University College London.

April 21, 2020 |

Tandem repeats lead to sequence assembly errors and impose multi-level challenges for genome and protein databases.

The widespread occurrence of repetitive stretches of DNA in genomes of organisms across the tree of life imposes fundamental challenges for sequencing, genome assembly, and automated annotation of genes and proteins. This multi-level problem can lead to errors in genome and protein databases that are often not recognized or acknowledged. As a consequence, end users working with sequences with repetitive regions are faced with ‘ready-to-use’ deposited data whose trustworthiness is difficult to determine, let alone to quantify. Here, we provide a review of the problems associated with tandem repeat sequences that originate from different stages during the sequencing-assembly-annotation-deposition workflow, and that may proliferate in public database repositories affecting all downstream analyses. As a case study, we provide examples of the Atlantic cod genome, whose sequencing and assembly were hindered by a particularly high prevalence of tandem repeats. We complement this case study with examples from other species, where mis-annotations and sequencing errors have propagated into protein databases. With this review, we aim to raise the awareness level within the community of database users, and alert scientists working in the underlying workflow of database creation that the data they omit or improperly assemble may well contain important biological information valuable to others. © The Author(s) 2019. Published by Oxford University Press on behalf of Nucleic Acids Research.

April 21, 2020 |

The bracteatus pineapple genome and domestication of clonally propagated crops.

Domestication of clonally propagated crops such as pineapple from South America was hypothesized to be a ‘one-step operation’. We sequenced the genome of Ananas comosus var. bracteatus CB5 and assembled 513?Mb into 25 chromosomes with 29,412 genes. Comparison of the genomes of CB5, F153 and MD2 elucidated the genomic basis of fiber production, color formation, sugar accumulation and fruit maturation. We also resequenced 89 Ananas genomes. Cultivars ‘Smooth Cayenne’ and ‘Queen’ exhibited ancient and recent admixture, while ‘Singapore Spanish’ supported a one-step operation of domestication. We identified 25 selective sweeps, including a strong sweep containing a pair of tandemly duplicated bromelain inhibitors. Four candidate genes for self-incompatibility were linked in F153, but were not functional in self-compatible CB5. Our findings support the coexistence of sexual recombination and a one-step operation in the domestication of clonally propagated crops. This work guides the exploration of sexual and asexual domestication trajectories in other clonally propagated crops.

April 21, 2020 |

Analyses of the Complete Genome Sequence of the Strain Bacillus pumilus ZB201701 Isolated from Rhizosphere Soil of Maize under Drought and Salt Stress.

Bacillus pumilus ZB201701 is a rhizobacterium with the potential to promote plant growth and tolerance to drought and salinity stress. We herein present the complete genome sequence of the Gram-positive bacterium B. pumilus ZB201701, which consists of a linear chromosome with 3,640,542 base pairs, 3,608 protein-coding sequences, 24 ribosomal RNAs, and 80 transfer RNAs. Genome analyses using bioinformatics revealed some of the putative gene clusters involved in defense mechanisms. In addition, activity analyses of the strain under salt and simulated drought stress suggested its potential tolerance to abiotic stress. Plant growth-promoting bacteria-based experiments indicated that the strain promotes the salt tolerance of maize. The complete genome of B. pumilus ZB201701 provides valuable insights into rhizobacteria-mediated salt and drought tolerance and rhizobacteria-based solutions for abiotic stress in agriculture.

April 21, 2020 |

De novo assembly of a wild pear (Pyrus betuleafolia) genome.

China is the origin and evolutionary centre of Oriental pears. Pyrus betuleafolia is a wild species native to China and distributed in the northern region, and it is widely used as rootstock. Here, we report the de novo assembly of the genome of P. betuleafolia-Shanxi Duli using an integrated strategy that combines PacBio sequencing, BioNano mapping and chromosome conformation capture (Hi-C) sequencing. The genome assembly size was 532.7 Mb, with a contig N50 of 1.57 Mb. A total of 59 552 protein-coding genes and 247.4 Mb of repetitive sequences were annotated for this genome. The expansion genes in P. betuleafolia were significantly enriched in secondary metabolism, which may account for the organism’s considerable environmental adaptability. An alignment analysis of orthologous genes showed that fruit size, sugar metabolism and transport, and photosynthetic efficiency were positively selected in Oriental pear during domestication. A total of 573 nucleotide-binding site (NBS)-type resistance gene analogues (RGAs) were identified in the P. betuleafolia genome, 150 of which are TIR-NBS-LRR (TNL)-type genes, which represented the greatest number of TNL-type genes among the published Rosaceae genomes and explained the strong disease resistance of this wild species. The study of flavour metabolism-related genes showed that the anthocyanidin reductase (ANR) metabolic pathway affected the astringency of pear fruit and that sorbitol transporter (SOT) transmembrane transport may be the main factor affecting the accumulation of soluble organic matter. This high-quality P. betuleafolia genome provides a valuable resource for the utilization of wild pear in fundamental pear studies and breeding. © 2019 The Authors. Plant Biotechnology Journal published by Society for Experimental Biology and The Association of Applied Biologists and John Wiley & Sons Ltd.

April 21, 2020 |

Chlorella vulgaris genome assembly and annotation reveals the molecular basis for metabolic acclimation to high light conditions.

Chlorella vulgaris is a fast-growing fresh-water microalga cultivated at the industrial scale for applications ranging from food to biofuel production. To advance our understanding of its biology and to establish genetics tools for biotechnological manipulation, we sequenced the nuclear and organelle genomes of Chlorella vulgaris 211/11P by combining next generation sequencing and optical mapping of isolated DNA molecules. This hybrid approach allowed to assemble the nuclear genome in 14 pseudo-molecules with an N50 of 2.8 Mb and 98.9% of scaffolded genome. The integration of RNA-seq data obtained at two different irradiances of growth (high light-HL versus low light -LL) enabled to identify 10,724 nuclear genes, coding for 11,082 transcripts. Moreover 121 and 48 genes were respectively found in the chloroplast and mitochondrial genome. Functional annotation and expression analysis of nuclear, chloroplast and mitochondrial genome sequences revealed peculiar features of Chlorella vulgaris. Evidence of horizontal gene transfers from chloroplast to mitochondrial genome was observed. Furthermore, comparative transcriptomic analyses of LL vs HL provide insights into the molecular basis for metabolic rearrangement in HL vs. LL conditions leading to enhanced de novo fatty acid biosynthesis and triacylglycerol accumulation. The occurrence of a cytosolic fatty acid biosynthetic pathway can be predicted and its upregulation upon HL exposure is observed, consistent with increased lipid amount under HL. These data provide a rich genetic resource for future genome editing studies, and potential targets for biotechnological manipulation of Chlorella vulgaris or other microalgae species to improve biomass and lipid productivity.This article is protected by copyright. All rights reserved.

April 21, 2020 |

De novo assembly and annotation of the Ganoderma australe genome.

The Ganoderma genus represents clear biotechnological potential, due to the large quantity of molecules with biological activity that could be explored. However, available information regarding the biotechnological importance of species within Ganoderma, other than G. lucidum, is quite limited. Genomic studies of little-known species can contribute to the knowledge thereof, as well as the search for metabolic pathways and the identification of genes which code for proteins that may be of biotechnological relevance. Therefore, the objective of the present study was to obtain the G. australe genome, through the use of new sequencing technologies. Genomic DNA from G. australe was sequenced with the PacBio Sequel system, to a depth of 100×. The genome was assembled de novo with the Canu assembly tool, and gene prediction and annotation were performed with a funannotate pipeline. An assembled 84?Mb genome was obtained, and 22,756 putative protein-coding sequences were predicted in the G. australe genome. Ganoderic acid pathways were annotated and listed in the funannotate pipeline, and were recognized using Pfam and Antismash signals. Thus, the G. australe genome shows great potential, mainly, due to the annotation of putative sequences that could be employed in biotechnological approaches. Copyright © 2019 Elsevier Inc. All rights reserved.

April 21, 2020 |

Complete genome sequence of Marinobacter sp. LQ44, a haloalkaliphilic phenol-degrading bacterium isolated from a deep-sea hydrothermal vent

Marinobacter sp. strain LQ44, an alkaliphile and moderate halophile from a deep-sea hydrothermal vent on the East Pacific Rise, is a novel phenol-degrading bacterium that is capable of utilizing phenol as sole carbon and energy sources. Here, we present the complete genome sequence of strain LQ44, which consists of 4,435,564?bp with a circular chromosome, 4164 protein-coding genes, 3 rRNA operons and 50 tRNAs. Genome analysis revealed that strain LQ44 may degrade phenol via meta-cleavage pathway. The LQ44 genome contains multiple genes involved in pH adaptation and osmotic adjustment. Genes related to hydrocarbon degradation, aerobic denitrification and potential industrial important enzymes were also identified from the genome. To our knowledge, this is the first report of a genome sequence of a haloalkaliphilic phenol-degrading bacterium, which will provide insights into the survival of this bacterium under salt-alkali conditions and the potential for biotechnological applications.

April 21, 2020 |

Genome assembly provides insights into the genome evolution and flowering regulation of orchardgrass.

Orchardgrass (Dactylis glomerata L.) is an important forage grass for cultivating livestock worldwide. Here, we report an ~1.84-Gb chromosome-scale diploid genome assembly of orchardgrass, with a contig N50 of 0.93 Mb, a scaffold N50 of 6.08 Mb and a super-scaffold N50 of 252.52 Mb, which is the first chromosome-scale assembled genome of a cool-season forage grass. The genome includes 40 088 protein-coding genes, and 69% of the assembled sequences are transposable elements, with long terminal repeats (LTRs) being the most abundant. The LTRretrotransposons may have been activated and expanded in the grass genome in response to environmental changes during the Pleistocene between 0 and 1 million years ago. Phylogenetic analysis reveals that orchardgrass diverged after rice but before three Triticeae species, and evolutionarily conserved chromosomes were detected by analysing ancient chromosome rearrangements in these grass species. We also resequenced the whole genome of 76 orchardgrass accessions and found that germplasm from Northern Europe and East Asia clustered together, likely due to the exchange of plants along the ‘Silk Road’ or other ancient trade routes connecting the East and West. Last, a combined transcriptome, quantitative genetic and bulk segregant analysis provided insights into the genetic network regulating flowering time in orchardgrass and revealed four main candidate genes controlling this trait. This chromosome-scale genome and the online database of orchardgrass developed here will facilitate the discovery of genes controlling agronomically important traits, stimulate genetic improvement of and functional genetic research on orchardgrass and provide comparative genetic resources for other forage grasses. © 2019 The Authors. Plant Biotechnology Journal published by Society for Experimental Biology and The Association of Applied Biologists and John Wiley & Sons Ltd.

April 21, 2020 |

Complete genome sequence of Paracoccus sp. Arc7-R13, a silver nanoparticles synthesizing bacterium isolated from Arctic Ocean sediments

Paracoccus sp. Arc7-R13, a silver nanoparticles (AgNPs) synthesizing bacterium, was isolated from Arctic Ocean sediment. Here we describe the complete genome of Paracoccus sp. Arc7-R13. The complete genome contains 4,040,012?bp with 66.66?mol%?G?+?C content, including one circular chromosome of 3,231,929?bp (67.45?mol%?G?+?C content), and eight plasmids with length ranging from 24,536?bp to 199,685?bp. The genome contains 3835 protein-coding genes (CDSs), 49 tRNA genes, as well as 3 rRNA operons as 16S-23S-5S rRNA. Based on the gene annotation and Swiss-Prot analysis, a total of 15 genes belonging to 11 kinds, including silver exporting P-type ATPase (SilP), alkaline phosphatase, nitroreductase, thioredoxin reductase, NADPH dehydrogenase and glutathione peroxidase, might be related to the synthesis of AgNPs. Meanwhile, many additional genes associated with synthesis of AgNPs such as protein-disulfide isomerase, c-type cytochrome, glutathione synthase and dehydrogenase reductase were also identified.

April 21, 2020 |

Pseudo-chromosome length genome assembly of a double haploid ‘Bartlett’ pear (Pyrus communis L.)

We report an improved assembly and scaffolding of the European pear (Pyrus communis L.) genome (referred to as BartlettDHv2.0), obtained using a combination of Pacific Biosciences RSII Long read sequencing (PacBio), Bionano optical mapping, chromatin interaction capture (Hi-C), and genetic mapping. A total of 496.9 million bases (Mb) corresponding to 97% of the estimated genome size were assembled into 494 scaffolds. Hi-C data and a high-density genetic map allowed us to anchor and orient 87% of the sequence on the 17 chromosomes of the pear genome. About 50% (247 Mb) of the genome consists of repetitive sequences. Comparison with previous assemblies of Pyrus communis. and Pyrus x bretschneideri confirmed the presence of 37,445 protein-coding genes, which is 13% fewer than previously predicted.

April 21, 2020 |

RNA sequencing: the teenage years.

Over the past decade, RNA sequencing (RNA-seq) has become an indispensable tool for transcriptome-wide analysis of differential gene expression and differential splicing of mRNAs. However, as next-generation sequencing technologies have developed, so too has RNA-seq. Now, RNA-seq methods are available for studying many different aspects of RNA biology, including single-cell gene expression, translation (the translatome) and RNA structure (the structurome). Exciting new applications are being explored, such as spatial transcriptomics (spatialomics). Together with new long-read and direct RNA-seq technologies and better computational tools for data analysis, innovations in RNA-seq are contributing to a fuller understanding of RNA biology, from questions such as when and where transcription occurs to the folding and intermolecular interactions that govern RNA function.

April 21, 2020 |

Benchmarking Transposable Element Annotation Methods for Creation of a Streamlined, Comprehensive Pipeline

Sequencing technology and assembly algorithms have matured to the point that high-quality de novo assembly is possible for large, repetitive genomes. Current assemblies traverse transposable elements (TEs) and allow for annotation of TEs. There are numerous methods for each class of elements with unknown relative performance metrics. We benchmarked existing programs based on a curated library of rice TEs. Using the most robust programs, we created a comprehensive pipeline called Extensive de-novo TE Annotator (EDTA) that produces a condensed TE library for annotations of structurally intact and fragmented elements. EDTA is open-source and freely available: https://github.com/oushujun/EDTA.List of abbreviationsTETransposable ElementsLTRLong Terminal RepeatLINELong Interspersed Nuclear ElementSINEShort Interspersed Nuclear ElementMITEMiniature Inverted Transposable ElementTIRTerminal Inverted RepeatTSDTarget Site DuplicationTPTrue PositivesFPFalse PositivesTNTrue NegativeFNFalse NegativesGRFGeneric Repeat FinderEDTAExtensive de-novo TE Annotator

April 21, 2020 |

Genomic analysis of Marinobacter sp. NP-4 and NP-6 isolated from the deep-sea oceanic crust on the western flank of the Mid-Atlantic Ridge

Two Marinobacter sp. NP-4 and NP-6 were isolated from a deep oceanic basaltic crust at North Pond, located at the western flank of the Mid-Atlantic Ridge. These two strains are capable of using multiple carbon sources such as acetate, succinate, glucose and sucrose while take oxygen as a primary electron acceptor. The strain NP-4 is also able to grow anaerobically under 20?MPa, with nitrate as the electron acceptor, thus represents a piezotolerant. To explore the metabolic potentials of Marinobacter sp. NP-4 and NP-6, the complete genome of NP-4 and close-to-complete genome of NP-6 were sequenced. The genome of NP-4 contains one chromosome and two plasmids with the size of 4.6?Mb in total, and with average GC content of 57.0%. The genome of NP-6 is 4.5?Mb and consists of 6 scaffolds, with an average GC content of 57.1%. Complete glycolysis, citrate cycle and aromatics compounds degradation pathways are identified in genomes of these two strains, suggesting that they possess a heterotrophic life style. Additionally, one plasmid of NP-4 contains genes for alkane degradation, phosphonate ABC transporter and cation efflux system, enabling NP-4 extra surviving abilities. In total, genomic information of these two strains provide insights into the physiological features and adaptation strategies of Marinobacter spp. in the deep oceanic crust biosphere.

April 21, 2020 |

Complete genome sequence of Paenisporosarcina antarctica CGMCC 1.6503 T, a marine psychrophilic bacterium isolated from Antarctica

A marine psychrophilic bacterium _Paenisporosarcina antarctica_ CGMCC 1.6503T (= JCM 14646T) was isolated off King George Island, Antarctica (62°13’31? S 58°57’08? W). In this study, we report the complete genome sequence of _Paenisporosarcina antarctica_, which is comprised of 3,972,524?bp with a mean G?+?C content of 37.0%. By gene function and metabolic pathway analyses, studies showed that strain CGMCC 1.6503T encodes a series of genes related to cold adaptation, including encoding fatty acid desaturases, dioxygenases, antifreeze proteins and cold shock proteins, and possesses several two-component regulatory systems, which could assist this strain in responding to the cold stress, the oxygen stress and the osmotic stress in Antarctica. The complete genome sequence of _P. antarctica_ may provide further insights into the genetic mechanism of cold adaptation for Antarctic marine bacteria.

Auto Tag: protein-coding

Improved assembly and variant detection of a haploid human genome using single-molecule, high-fidelity long reads.

Tandem repeats lead to sequence assembly errors and impose multi-level challenges for genome and protein databases.

The bracteatus pineapple genome and domestication of clonally propagated crops.

Analyses of the Complete Genome Sequence of the Strain Bacillus pumilus ZB201701 Isolated from Rhizosphere Soil of Maize under Drought and Salt Stress.

De novo assembly of a wild pear (Pyrus betuleafolia) genome.

Chlorella vulgaris genome assembly and annotation reveals the molecular basis for metabolic acclimation to high light conditions.

De novo assembly and annotation of the Ganoderma australe genome.

Complete genome sequence of Marinobacter sp. LQ44, a haloalkaliphilic phenol-degrading bacterium isolated from a deep-sea hydrothermal vent

Genome assembly provides insights into the genome evolution and flowering regulation of orchardgrass.

Complete genome sequence of Paracoccus sp. Arc7-R13, a silver nanoparticles synthesizing bacterium isolated from Arctic Ocean sediments

Pseudo-chromosome length genome assembly of a double haploid ‘Bartlett’ pear (Pyrus communis L.)

RNA sequencing: the teenage years.

Benchmarking Transposable Element Annotation Methods for Creation of a Streamlined, Comprehensive Pipeline

Genomic analysis of Marinobacter sp. NP-4 and NP-6 isolated from the deep-sea oceanic crust on the western flank of the Mid-Atlantic Ridge

Complete genome sequence of Paenisporosarcina antarctica CGMCC 1.6503 T, a marine psychrophilic bacterium isolated from Antarctica

Subscribe for blog updates:

Filter by topic

Talk with an expert

ALS case study

Subscribe for blog updates:

Filter by topic

Talk with an expert