---
author:
- contributor_roles: []
  family: Edmunds
  given: Scott
  url: https://orcid.org/0000-0001-6444-1436
blog:
  authors: null
  community_id: 52db0518-e228-4260-8c54-c4e323b2569d
  created: 1675555200
  current_feed_url: null
  description: Data driven blogging from the GigaScience editors
  doi: https://doi.org/10.59350/gigablog
  favicon: https://rogue-scholar.org/api/communities/52db0518-e228-4260-8c54-c4e323b2569d/logo
  feed_format: application/atom+xml
  feed_url: http://gigasciencejournal.com/blog/feed/atom/
  filter: null
  generator: Other
  home_page_url: https://gigasciencejournal.com/blog
  issn: null
  language: eng
  license: https://creativecommons.org/licenses/by/4.0/legalcode
  prefix: '10.59350'
  relative_url: null
  secure: false
  slug: gigablog
  status: archived
  subfield: '1311'
  title: GigaBlog
  updated: null
  use_api: null
container: GigaBlog
date: '2018-12-13T00:00:00+00:00'
date_updated: '2025-12-06T10:24:52+00:00'
guid: http://gigasciencejournal.com/blog/?p=2612
identifier: https://doi.org/10.59350/22rps-x6446
image: http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/old-goldmine-and-water-tanks-300x186.jpg
images:
- alt: microbial genomics goldmine
  height: '186'
  sizes: '(max-width: 300px) 100vw, 300px'
  src: http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/old-goldmine-and-water-tanks-300x186.jpg
  srcset: http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/old-goldmine-and-water-tanks-300x186.jpg,
    http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/old-goldmine-and-water-tanks-768x477.jpg,
    http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/old-goldmine-and-water-tanks-1024x636.jpg,
    http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/old-goldmine-and-water-tanks.jpg
  width: '300'
- alt: Microbial goldmine
  height: '832'
  sizes: '(max-width: 1024px) 100vw, 1024px'
  src: http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/Lisa_ICG-1024x832.jpg
  srcset: http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/Lisa_ICG-1024x832.jpg,
    http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/Lisa_ICG-300x244.jpg,
    http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/Lisa_ICG-768x624.jpg
  width: '1024'
issn: null
keywords:
- Biology
- Genomics
- ICG Prize
- ICG13
- Metagenomics
lang: en
license: https://creativecommons.org/licenses/by/4.0/legalcode
rid: fy4tf-gg360
rights: https://creativecommons.org/licenses/by/4.0/legalcode
summary: '*Out today is the winner of our ICG13 Prize, presenting work that can aid
  in revealing new biologically relevant findings and missed genes from previously
  generated transcriptome assemblies. Teaching old data new tricks, and maximising
  every last nugget of information from previously funded research.'
title: 'Reprocessing the Microbial Genomic Goldmine: Winner of the ICG13 Prize'
url: https://wayback.archive-it.org/22098/2025-05-01T17:13:42Z/http://gigasciencejournal.com/blog/microbial-genomics-goldmine
version: v1
---

\*[Out today](https://doi.org/10.1093/gigascience/giy158)[![microbial
genomics
goldmine](http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/old-goldmine-and-water-tanks-300x186.jpg){.alignright
.wp-image-2614 .size-medium loading="lazy" decoding="async"
srcset="http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/old-goldmine-and-water-tanks-300x186.jpg 300w, http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/old-goldmine-and-water-tanks-768x477.jpg 768w, http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/old-goldmine-and-water-tanks-1024x636.jpg 1024w, http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/old-goldmine-and-water-tanks.jpg 1885w"
sizes="(max-width: 300px) 100vw, 300px" width="300"
height="186"}](https://doi.org/10.1093/gigascience/giy158) is the winner
of our ICG13 Prize, presenting work that can aid in revealing new
biologically relevant findings and missed genes from previously
generated transcriptome assemblies. Teaching old data new tricks, and
maximising every last nugget of information from previously funded
research. Here we present some insight into why the reviewers and judges
felt this work was so novel and these efforts to reprocess the microbial
genomics goldmine has such promise in the field of reproducibility and
data reuse.\
\*

Analyses of genomic data often miss a large amount of information that
is present in the data due to the presence of so-called genomics \"dark
data\". This is information that is actually contained within the data,
but due to limitations in the analysis methods and variation in the
analyses tools used, it is usually missed. Titus Brown and his lab have
now created an automated pipeline to assemble and annotate
previously-analyzed raw data to dig out this hidden information. This
study mines a huge marine microbial dataset from the [Microbial
Transcriptome Sequencing Project
(MMETSP)](https://doi.org/10.1371/journal.pbio.1001889), demonstrating
that re-analysing old data with new tools can yield new results.

Previous [work on the
MMETSP](https://doi.org/10.1371/journal.pbio.1001889) sequenced 678
transcriptomes and assembled genes that spanned 396 different strains of
marine eukaryotes. This dataset has been an invaluable resource within
the oceanographic community, exponentially expanding the accessible
genetic information base of marine protistan life. In the 5 years since
the original analysis was completed, tools, techniques and databases
have been improved on. While analysis of this historical data could
potentially be carried out again to produce new and more accurate
findings, re-analysis of previously generated data with new tools is not
commonplace, and it is unclear what the best practice would be. Running
analyses again produces different results, and the effects of using
different pipelines are poorly understood, making it difficult to
determine the usefulness of the new results relative to the previous
findings.

The authors of this study tackled this challenge in a systematic manner.
They created an automated pipeline to assemble and annotate the original
raw data from the MMETSP data. The resulting new transcriptome
assemblies were then automatically evaluated in the pipeline and
compared against previously-generated assemblies from the original
assembly pipeline developed by the National Center for Genome Research.
As there is no one-size-fits-all protocol for transcriptome assembly,
and as software tools are constantly improving, this pipeline enabled
improvements to be tested and quantified. The new assemblies generated
containing the majority of the previous data as well as new content. On
average, 7.8% of the annotated sequence in the new assemblies had novel
gene names not found in the historical assemblies, demonstrating that
new findings can be gleaned from old data. Asked why having the most
accurate and up-to-date assembly is important, author Lisa Johnson
stated: \"Having the best possible quality reference is necessary to be
able to accurately characterize new RNAseq data. This is especially true
if significant investments will be made downstream based on differential
expression results, as is sometimes the case with biomarker and drug
discovery in the agriculture, food and pharmaceutical fields.\"

::: {#attachment_2615 .wp-caption .aligncenter style="width: 1034px"}
![Microbial
goldmine](http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/Lisa_ICG-1024x832.jpg){.wp-image-2615
.size-large loading="lazy" decoding="async"
aria-describedby="caption-attachment-2615"
srcset="http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/Lisa_ICG-1024x832.jpg 1024w, http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/Lisa_ICG-300x244.jpg 300w, http://gigasciencejournal.com/blog/wp-content/uploads/2018/12/Lisa_ICG-768x624.jpg 768w"
sizes="(max-width: 1024px) 100vw, 1024px" width="1024" height="832"}

Lisa Johnson presenting the work at ICG13 in Shenzhen.
:::

While raw sequencing data is commonly shared and well cared for by
government funded public archives, the resulting assembled genomes,
annotations and results generally are not. For the microbial genomics
work carried out in this article, the processed data and results of each
re-analysis would ideally be deposited in a discoverable location, with
versions and automated redirection to newer versions. The authors have
attempted to do just that, with the resulting outputs archived [in the
public Zenodo repository](https://doi.org/10.5281/zenodo.746048) hosted
by CERN, as well as snapshots from the study archived [in our GigaDB
database](http://dx.doi.org/10.5524/100522). If we are to take full
advantage of public data, this work demonstrates that researchers need
to make these products \"forward discoverable\", automatically notifying
users when a dataset is updated or changed. For researchers low on
resources, the benefit would make it possible to improve downstream work
without significant additional funding, experimentation or sequencing.

This microbial genomics work was selected by our international panel of
judges as the winner of our [second *GigaScience*
prize](http://gigasciencejournal.com/blog/icg13-gigascience-prize/), and
the first author Lisa Johnson came out to the International Conference
on Genomics in Shenzhen to present the work. In a follow up posting
we\'ll present a Q&A with Lisa going into more detail on this work, but
we also have [accompanying
commentary](https://doi.org/10.1093/gigascience/giy159) by the authors
expanding on the important lessons for reproducibility. The [video of
her talk](https://www.youtube.com/watch?v=WGmMOw0Jsqk&t=14s) is also
available to view below, and the slides are also [available to
view](https://www.slideshare.net/GigaScience/lisa-johnson-at-icg13-reassembly-quality-evaluation-and-annotation-of-678-microbial-eukaryotic-reference-transcriptomes).

\*\*Further Reading\
\*\*Johnson, LK et al. (2018): Re-assembly, quality evaluation, and
annotation of 678 microbial eukaryotic reference transcriptomes.
*GigaScience*. doi:
[10.1093/gigascience/giy158](https://doi.org/10.1093/gigascience/giy158)

Alexander, H et al. (2018): Keeping it light: (Re)analyzing
community-wide datasets without major infrastructure. *GigaScience*.
doi:
[10.1093/gigascience/giy159](https://doi.org/10.1093/gigascience/giy159)

The post [Reprocessing the Microbial Genomic Goldmine: Winner of the
ICG13
Prize](http://gigasciencejournal.com/blog/microbial-genomics-goldmine/){rel="nofollow"}
appeared first on
[GigaBlog](http://gigasciencejournal.com/blog){rel="nofollow"}.