ENCODE Archives - Page 5 of 45

April 21, 2020

Complete Genome Sequence of the Wolbachia wAlbB Endosymbiont of Aedes albopictus.

Wolbachia, an alpha-proteobacterium closely related to Rickettsia, is a maternally transmitted, intracellular symbiont of arthropods and nematodes. Aedes albopictus mosquitoes are naturally infected with Wolbachia strains wAlbA and wAlbB. Cell line Aa23 established from Ae. albopictus embryos retains only wAlbB and is a key model to study host-endosymbiont interactions. We have assembled the complete circular genome of wAlbB from the Aa23 cell line using long-read PacBio sequencing at 500× median coverage. The assembled circular chromosome is 1.48 megabases in size, an increase of more than 300 kb over the published draft wAlbB genome. The annotation of the genome identified 1,205 protein coding genes, 34 tRNA, 3 rRNA, 1 tmRNA, and 3 other ncRNA loci. The long reads enabled sequencing over complex repeat regions which are difficult to resolve with short-read sequencing. Thirteen percent of the genome comprised insertion sequence elements distributed throughout the genome, some of which cause pseudogenization. Prophage WO genes encoding some essential components of phage particle assembly are missing, while the remainder are found in five prophage regions/WO-like islands or scattered around the genome. Orthology analysis identified a core proteome of 535 orthogroups across all completed Wolbachia genomes. The majority of proteins could be annotated using Pfam and eggNOG analyses, including ankyrins and components of the Type IV secretion system. KEGG analysis revealed the absence of five genes in wAlbB which are present in other Wolbachia. The availability of a complete circular chromosome from wAlbB will enable further biochemical, molecular, and genetic analyses on this strain and related Wolbachia. © The Author(s) 2019. Published by Oxford University Press on behalf of the Society for Molecular Biology and Evolution.

April 21, 2020

Characterizing the major structural variant alleles of the human genome.

In order to provide a comprehensive resource for human structural variants (SVs), we generated long-read sequence data and analyzed SVs for fifteen human genomes. We sequence resolved 99,604 insertions, deletions, and inversions including 2,238 (1.6 Mbp) that are shared among all discovery genomes with an additional 13,053 (6.9 Mbp) present in the majority, indicating minor alleles or errors in the reference. Genotyping in 440 additional genomes confirms the most common SVs in unique euchromatin are now sequence resolved. We report a ninefold SV bias toward the last 5 Mbp of human chromosomes with nearly 55% of all VNTRs (variable number of tandem repeats) mapping to this portion of the genome. We identify SVs affecting coding and noncoding regulatory loci improving annotation and interpretation of functional variation. These data provide the framework to construct a canonical human reference and a resource for developing advanced representations capable of capturing allelic diversity. Copyright © 2018 Elsevier Inc. All rights reserved.

April 21, 2020

Genetic Variation, Comparative Genomics, and the Diagnosis of Disease.

The discovery of mutations associated with human genetic dis- ease is an exercise in comparative genomics (see Glossary). Although there are many different strategies and approaches, the central premise is that affected persons harbor a significant excess of pathogenic DNA variants as com- pared with a group of unaffected persons (controls) that is either clinically defined1 or established by surveying large swaths of the general population.2 The more exclu- sive the variant is to the disease, the greater its penetrance, the larger its effect size, and the more relevant it becomes to both disease diagnosis and future therapeutic investigation. The most popular approach used by researchers in human genetics is the case–control design, but there are others that can be used to track variants and disease in a family context or that consider the probability of different classes of mutations based on evolutionary patterns of divergence or de novo mutational change.3,4 Although the approaches may be straightforward, the discovery of patho- genic variation and its mechanism of action often is less trivial, and decades of research can be required in order to identify the variants underlying both mendelian and complex genetic traits.

April 21, 2020

Stout camphor tree genome fills gaps in understanding of flowering plant genome evolution.

We present reference-quality genome assembly and annotation for the stout camphor tree (Cinnamomum kanehirae (Laurales, Lauraceae)), the first sequenced member of the Magnoliidae comprising four orders (Laurales, Magnoliales, Canellales and Piperales) and over 9,000 species. Phylogenomic analysis of 13 representative seed plant genomes indicates that magnoliid and eudicot lineages share more recent common ancestry than monocots. Two whole-genome duplication events were inferred within the magnoliid lineage: one before divergence of Laurales and Magnoliales and the other within the Lauraceae. Small-scale segmental duplications and tandem duplications also contributed to innovation in the evolutionary history of Cinnamomum. For example, expansion of the terpenoid synthase gene subfamilies within the Laurales spawned the diversity of Cinnamomum monoterpenes and sesquiterpenes.

April 21, 2020

Finding Nemo’s Genes: A chromosome-scale reference assembly of the genome of the orange clownfish Amphiprion percula.

The iconic orange clownfish, Amphiprion percula, is a model organism for studying the ecology and evolution of reef fishes, including patterns of population connectivity, sex change, social organization, habitat selection and adaptation to climate change. Notably, the orange clownfish is the only reef fish for which a complete larval dispersal kernel has been established and was the first fish species for which it was demonstrated that antipredator responses of reef fishes could be impaired by ocean acidification. Despite its importance, molecular resources for this species remain scarce and until now it lacked a reference genome assembly. Here, we present a de novo chromosome-scale assembly of the genome of the orange clownfish Amphiprion percula. We utilized single-molecule real-time sequencing technology from Pacific Biosciences to produce an initial polished assembly comprised of 1,414 contigs, with a contig N50 length of 1.86 Mb. Using Hi-C-based chromatin contact maps, 98% of the genome assembly were placed into 24 chromosomes, resulting in a final assembly of 908.8 Mb in length with contig and scaffold N50s of 3.12 and 38.4 Mb, respectively. This makes it one of the most contiguous and complete fish genome assemblies currently available. The genome was annotated with 26,597 protein-coding genes and contains 96% of the core set of conserved actinopterygian orthologs. The availability of this reference genome assembly as a community resource will further strengthen the role of the orange clownfish as a model species for research on the ecology and evolution of reef fishes. © 2018 The Authors. Molecular Ecology Resources Published by John Wiley & Sons Ltd.

April 21, 2020

Transmission of ESBL-producing Escherichia coli between broilers and humans on broiler farms.

ESBL and AmpC ß-lactamases are an increasing concern for public health. Studies suggest that ESBL/pAmpC-producing Escherichia coli and their plasmids carrying antibiotic resistance genes can spread from broilers to humans working or living on broiler farms. These studies used traditional typing methods, which may not have provided sufficient resolution to reliably assess the relatedness of these isolates.Eleven suspected transmission events among broilers and humans living/working on eight broiler farms were investigated using whole-genome short-read (Illumina) and long-read sequencing (PacBio). Core genome MLST (cgMLST) was performed to investigate the occurrence of strain transmission. Horizontal plasmid and gene transfer were analysed using BLAST.Of eight suspected strain transmission events, six were confirmed. The isolate pairs had identical ESBL/AmpC genes and fewer than eight allelic differences according to the cgMLST, and five had an almost identical plasmid composition. On one of the farms, cgMLST revealed that the isolate pairs belonging to ST10 from a broiler and a household member of the farmer had 475 different alleles, but that the plasmids were identical, indicating horizontal transfer of mobile elements rather than strain transfer. Of three suspected horizontal plasmid transmission events, one was confirmed. In addition, gene transfer between plasmids was found.The present study confirms transmission of strains as well as horizontal plasmid and gene transfer between broilers and farmers and household members on the same farm. WGS is an important tool to confirm suspected zoonotic strain and resistance gene transmission. © The Author(s) 2019. Published by Oxford University Press on behalf of the British Society for Antimicrobial Chemotherapy. All rights reserved. For permissions, please email: journals.permissions@oup.com.

April 21, 2020

Immunogenetic factors driving formation of ultralong VH CDR3 in Bos taurus antibodies.

The antibody repertoire of Bos taurus is characterized by a subset of variable heavy (VH) chain regions with ultralong third complementarity determining regions (CDR3) which, compared to other species, can provide a potent response to challenging antigens like HIV env. These unusual CDR3 can range to over seventy highly diverse amino acids in length and form unique ß-ribbon ‘stalk’ and disulfide bonded ‘knob’ structures, far from the typical antigen binding site. The genetic components and processes for forming these unusual cattle antibody VH CDR3 are not well understood. Here we analyze sequences of Bos taurus antibody VH domains and find that the subset with ultralong CDR3 exclusively uses a single variable gene, IGHV1-7 (VHBUL) rearranged to the longest diversity gene, IGHD8-2. An eight nucleotide duplication at the 3′ end of IGHV1-7 encodes a longer V-region producing an extended F ß-strand that contributes to the stalk in a rearranged CDR3. A low amino acid variability was observed in CDR1 and CDR2, suggesting that antigen binding for this subset most likely only depends on the CDR3. Importantly a novel, potentially AID mediated, deletional diversification mechanism of the B. taurus VH ultralong CDR3 knob was discovered, in which interior codons of the IGHD8-2 region are removed while maintaining integral structural components of the knob and descending strand of the stalk in place. These deletions serve to further diversify cysteine positions, and thus disulfide bonded loops. Hence, both germline and somatic genetic factors and processes appear to be involved in diversification of this structurally unusual cattle VH ultralong CDR3 repertoire.

April 21, 2020

Complete Genome Sequence of Lactic Acid Bacterium Pediococcus acidilactici Strain ATCC 8042, an Autolytic Anti-bacterial Peptidoglycan Hydrolase Producer

Pediococcus acidilactici is a probiotic bacterium that is industrially utilized in the food industry and antibiotics development. Here, we determine the complete nucleotide sequence of the genome of Pediococcus acidilactici ATCC 8042. The genome was sequenced by the PacBio RSII to generate a single contig consisting of circular chromosome sequence. Illumina MiniSeq sequencing platform and Sanger sequencing method were additionally utilized to correct errors resulting from the long-read sequencing platform. The sequence consists of 2,009,598 bp with a G + C content of 42.1% and contains 1,865 protein-coding sequences. Based on the sequence information, we could confirm and predict the presence of four peptidoglycan hydrolases by HyPe software. This work, therefore, provides the complete genomic information of P. acidilactici ATCC 8042 with a profitable potential of genome-scale comprehension of anti-pathogenic activity, which can be applied in nutraceutical and pharmaceutical biotechnology field.

April 21, 2020

Genomic variation and strain-specific functional adaptation in the human gut microbiome during early life.

The human gut microbiome matures towards the adult composition during the first years of life and is implicated in early immune development. Here, we investigate the effects of microbial genomic diversity on gut microbiome development using integrated early childhood data sets collected in the DIABIMMUNE study in Finland, Estonia and Russian Karelia. We show that gut microbial diversity is associated with household location and linear growth of children. Single nucleotide polymorphism- and metagenomic assembly-based strain tracking revealed large and highly dynamic microbial pangenomes, especially in the genus Bacteroides, in which we identified evidence of variability deriving from Bacteroides-targeting bacteriophages. Our analyses revealed functional consequences of strain diversity; only 10% of Finnish infants harboured Bifidobacterium longum subsp. infantis, a subspecies specialized in human milk metabolism, whereas Russian infants commonly maintained a probiotic Bifidobacterium bifidum strain in infancy. Groups of bacteria contributing to diverse, characterized metabolic pathways converged to highly subject-specific configurations over the first two years of life. This longitudinal study extends the current view of early gut microbial community assembly based on strain-level genomic variation.

April 21, 2020

Antarctic blackfin icefish genome reveals adaptations to extreme environments.

Icefishes (suborder Notothenioidei; family Channichthyidae) are the only vertebrates that lack functional haemoglobin genes and red blood cells. Here, we report a high-quality genome assembly and linkage map for the Antarctic blackfin icefish Chaenocephalus aceratus, highlighting evolved genomic features for its unique physiology. Phylogenomic analysis revealed that Antarctic fish of the teleost suborder Notothenioidei, including icefishes, diverged from the stickleback lineage about 77 million years ago and subsequently evolved cold-adapted phenotypes as the Southern Ocean cooled to sub-zero temperatures. Our results show that genes involved in protection from ice damage, including genes encoding antifreeze glycoprotein and zona pellucida proteins, are highly expanded in the icefish genome. Furthermore, genes that encode enzymes that help to control cellular redox state, including members of the sod3 and nqo1 gene families, are expanded, probably as evolutionary adaptations to the relatively high concentration of oxygen dissolved in cold Antarctic waters. In contrast, some crucial regulators of circadian homeostasis (cry and per genes) are absent from the icefish genome, suggesting compromised control of biological rhythms in the polar light environment. The availability of the icefish genome sequence will accelerate our understanding of adaptation to extreme Antarctic environments.

April 21, 2020

Genetic map-guided genome assembly reveals a virulence-governing minichromosome in the lentil anthracnose pathogen Colletotrichum lentis.

Colletotrichum lentis causes anthracnose, which is a serious disease on lentil and can account for up to 70% crop loss. Two pathogenic races, 0 and 1, have been described in the C. lentis population from lentil. To unravel the genetic control of virulence, an isolate of the virulent race 0 was sequenced at 1481-fold genomic coverage. The 56.10-Mb genome assembly consists of 50 scaffolds with N50 scaffold length of 4.89 Mb. A total of 11 436 protein-coding gene models was predicted in the genome with 237 coding candidate effectors, 43 secondary metabolite biosynthetic enzymes and 229 carbohydrate-active enzymes (CAZymes), suggesting a contraction of the virulence gene repertoire in C. lentis. Scaffolds were assigned to 10 core and two minichromosomes using a population (race 0 × race 1, n = 94 progeny isolates) sequencing-based, high-density (14 312 single nucleotide polymorphisms) genetic map. Composite interval mapping revealed a single quantitative trait locus (QTL), qClVIR-11, located on minichromosome 11, explaining 85% of the variability in virulence of the C. lentis population. The QTL covers a physical distance of 0.84 Mb with 98 genes, including seven candidate effector and two secondary metabolite genes. Taken together, the study provides genetic and physical evidence for the existence of a minichromosome controlling the C. lentis virulence on lentil. © 2018 The Authors. New Phytologist © 2018 New Phytologist Trust.

April 21, 2020

Nephromyces encodes a urate metabolism pathway and predicted peroxisomes, demonstrating that these are not ancient losses of apicomplexans.

The phylum Apicomplexa is a quintessentially parasitic lineage, whose members infect a broad range of animals. One exception to this may be the apicomplexan genus Nephromyces, which has been described as having a mutualistic relationship with its host. Here we analyze transcriptome data from Nephromyces and its parasitic sister taxon, Cardiosporidium, revealing an ancestral purine degradation pathway thought to have been lost early in apicomplexan evolution. The predicted localization of many of the purine degradation enzymes to peroxisomes, and the in silico identification of a full set of peroxisome proteins, indicates that loss of both features in other apicomplexans occurred multiple times. The degradation of purines is thought to play a key role in the unusual relationship between Nephromyces and its host. Transcriptome data confirm previous biochemical results of a functional pathway for the utilization of uric acid as a primary nitrogen source for this unusual apicomplexan.

April 21, 2020

In-depth analysis of the genome of Trypanosoma evansi, an etiologic agent of surra.

Trypanosoma evansi is the causative agent of the animal trypanosomiasis surra, a disease with serious economic burden worldwide. The availability of the genome of its closely related parasite Trypanosoma brucei allows us to compare their genetic and evolutionarily shared and distinct biological features. The complete genomic sequence of the T. evansi YNB strain was obtained using a combination of genomic and transcriptomic sequencing, de novo assembly, and bioinformatic analysis. The genome size of the T. evansi YNB strain was 35.2 Mb, showing 96.59% similarity in sequence and 88.97% in scaffold alignment with T. brucei. A total of 8,617 protein-coding genes, accounting for 31% of the genome, were predicted. Approximately 1,641 alternative splicing events of 820 genes were identified, with a majority mediated by intron retention, which represented a major difference in post-transcriptional regulation between T. evansi and T. brucei. Disparities in gene copy number of the variant surface glycoprotein, expression site-associated genes, microRNAs, and RNA-binding protein were clearly observed between the two parasites. The results revealed the genomic determinants of T. evansi, which encoded specific biological characteristics that distinguished them from other related trypanosome species.

April 21, 2020

Bioinformatic analysis of the complete genome sequence of Pectobacterium carotovorum subsp. brasiliense BZA12 and candidate effector screening

AbstractPectobacterium carotovorum subsp. brasiliense (Pcb) is a gram-negative, plant pathogenic bacterium of the soft rot Enterobacteriaceae (SRE) family. We present the complete genome sequence of Pcb strain BZA12, which reveals that Pcb strain BZA12 carries a single 4,924,809 bp chromosome with 51.97% GC content and comprises 4508 predicted protein-coding genes.Geneannotationofthese genes utilizedGO, KEGG,and COG databases.Incomparison withthree closely related soft-rot pathogens, strain BZA12 has 3797 gene families, among which 3107 gene families are identified as orthologous with those of both P. carotovorum subsp. carotovorum PCC21 and P. carotovorum subsp. odoriferum BCS7, as well as 36 putative Unique Gene Families. We selected five putative effectors from the BZA12 genome and transiently expressed them in Nicotiana benthamiana. Candidate effector A12GL002483 was localized in the cell nucleus and induced cell death. This study provides a foundation for a better understanding of the genomic structure and function of Pcb, particularly in the discovery of potential pathogenic factors and for the development of more effective strategies against this pathogen.

April 21, 2020

Physiological properties and genetic analysis related to exopolysaccharide (EPS) production in the fresh-water unicellular cyanobacterium Aphanothece sacrum (Suizenji Nori).

The clonal strains, phycoerythrin(PE)-rich- and PE-poor strains, of the unicellular, fresh water cyanobacterium Aphanothece sacrum (Suringar) Okada (Suizenji Nori, in Japanese) were isolated from traditional open-air aquafarms in Japan. A. sacrum appeared to be oligotrophic on the basis of its growth characteristics. The optimum temperature for growth was around 20°C. Maximum growth and biomass increase at 20°C was obtained under light intensities between 40 to 80 µmol m-2 s-1 (fluorescent lamps, 12 h light/12 h dark cycles) and between 40 to 120 µmol m-2 s-1 for PE-rich and PE-poor strains, respectively, of A. sacrum . Purified exopolysaccharide (EPS) of A. sacrum has a molecular weight of ca. 104 kDa with five major monosaccharides (glucose, xylose, rhamnose, galactose and mannose; =85 mol%). We also deciphered the whole genome sequence of the two strains of A. sacrum. The putative genes involved in the polymerization, chain length control, and export of EPS would contribute to understand the biosynthetic process of their extremely high molecular weight EPS. The putative genes encoding Wzx-Wzy-Wzz- and Wza-Wzb-Wzc were conserved in the A. sacrum strains FPU1 and FPU3. This result suggests that the Wzy-dependent pathway participates in the EPS production of A. sacrum.

Auto Tag: ENCODE

Complete Genome Sequence of the Wolbachia wAlbB Endosymbiont of Aedes albopictus.

Characterizing the major structural variant alleles of the human genome.

Genetic Variation, Comparative Genomics, and the Diagnosis of Disease.

Stout camphor tree genome fills gaps in understanding of flowering plant genome evolution.

Finding Nemo’s Genes: A chromosome-scale reference assembly of the genome of the orange clownfish Amphiprion percula.

Transmission of ESBL-producing Escherichia coli between broilers and humans on broiler farms.

Immunogenetic factors driving formation of ultralong VH CDR3 in Bos taurus antibodies.

Complete Genome Sequence of Lactic Acid Bacterium Pediococcus acidilactici Strain ATCC 8042, an Autolytic Anti-bacterial Peptidoglycan Hydrolase Producer

Genomic variation and strain-specific functional adaptation in the human gut microbiome during early life.

Antarctic blackfin icefish genome reveals adaptations to extreme environments.

Genetic map-guided genome assembly reveals a virulence-governing minichromosome in the lentil anthracnose pathogen Colletotrichum lentis.

Nephromyces encodes a urate metabolism pathway and predicted peroxisomes, demonstrating that these are not ancient losses of apicomplexans.

In-depth analysis of the genome of Trypanosoma evansi, an etiologic agent of surra.

Bioinformatic analysis of the complete genome sequence of Pectobacterium carotovorum subsp. brasiliense BZA12 and candidate effector screening

Physiological properties and genetic analysis related to exopolysaccharide (EPS) production in the fresh-water unicellular cyanobacterium Aphanothece sacrum (Suizenji Nori).

Subscribe for blog updates:

Filter by topic

Talk with an expert

Antimicrobial resistance research

Subscribe for blog updates:

Filter by topic

Talk with an expert