---
author:
- contributor_roles: []
  family: Edmunds
  given: Scott
  url: https://orcid.org/0000-0001-6444-1436
blog:
  authors: null
  community_id: 52db0518-e228-4260-8c54-c4e323b2569d
  created: 1675555200
  current_feed_url: null
  description: Data driven blogging from the GigaScience editors
  doi: https://doi.org/10.59350/gigablog
  favicon: https://rogue-scholar.org/api/communities/52db0518-e228-4260-8c54-c4e323b2569d/logo
  feed_format: application/atom+xml
  feed_url: http://gigasciencejournal.com/blog/feed/atom/
  filter: null
  generator: Other
  home_page_url: https://gigasciencejournal.com/blog
  issn: null
  language: eng
  license: https://creativecommons.org/licenses/by/4.0/legalcode
  prefix: '10.59350'
  relative_url: null
  secure: false
  slug: gigablog
  status: archived
  subfield: '1311'
  title: GigaBlog
  updated: null
  use_api: null
container: GigaBlog
date: '2012-07-12T00:00:00+00:00'
date_updated: '2025-12-06T10:48:32+00:00'
guid: http://finaloriginalblogs.dev/gigablog/2012/07/12/gigascience-launches-overseeing-the-transition-from-papers-to-executable-research-objects/
identifier: https://doi.org/10.59350/9p2mj-m6k12
issn: null
keywords:
- Open Access
- GigaDB
- GigaScience
- Launch
lang: en
license: https://creativecommons.org/licenses/by/4.0/legalcode
rid: 3eqrw-c9325
rights: https://creativecommons.org/licenses/by/4.0/legalcode
summary: Research papers have been the predominant form of scholarly communication
  for the past few centuries, and despite moves towards online publication and open
  access, the process and structure of publication has not fundamentally changed in
  that time.
title: 'GigaScience launches: overseeing the transition from papers to "executable
  research objects"'
url: https://wayback.archive-it.org/22098/2025-05-01T17:13:42Z/http://gigasciencejournal.com/blog/gigascience-launches-overseeing-the-transition-from-papers-to-executable-research-objects
version: v1
---

Research papers have been the predominant form of scholarly
communication for the past few centuries, and despite moves towards
online publication and open access, the process and structure of
publication has not fundamentally changed in that time. With biological
and biomedical research becoming increasingly data-driven, and the
amount of information, computational tools, and code supporting a
publication in areas such as
[genomics](http://www.genome.gov/sequencingcosts/) and
[imaging](http://iht2blog.com/tag/imaging-data/ "IHT imaging data page"){target="_blank"
rel="noopener"} growing at exponential rates, the lack of access to the
resources that the paper is built upon is leading to a growing
\"reproducibility gap\".
[Recent](http://www.nature.com/news/2011/110111/full/469139a.html "Naturenews: Cancer trial errors revealed"){target="_blank"
rel="noopener"}
[scandals](http://www.nature.com/news/2011/111101/full/479015a.html "Naturenews: Dutch university fraud story"){target="_blank"
rel="noopener"} relating to falsified data that went long undetected
(including
[this](http://arstechnica.com/science/2012/07/new-record-for-faking-data-set-by-japanese-researcher/ "172 falsified publication story"){target="_blank"
rel="noopener"} particularly egregious recent example of an author
fabricating the data supporting at least 172 of their publications)
further highlight the need to make data easily accessible for purposes
of validation and to maintain public trust in science.

As research has shifted to work within, and handle, this data-rich
environment and to utilize advances such as cloud computing and
automated workflow systems,  publishing needs to be able to follow in a
similar direction. For a number of years
[much](http://www.executablepapers.com/ "Executable paper challenge"){target="_blank"
rel="noopener"}
[talk](http://www.wf4ever-project.org/about "Wf4Ever about page"){target="_blank"
rel="noopener"} has been made about the potential for executable papers:
aiding review and re-use of data by having all of the tools and data
associated with a publication accessible in a reproducible and
standardized environment. An important first step in the path has been
reached today with the first articles published in
[*GigaScience*](http://www.gigasciencejournal.com/ "GigaScience homepage"){target="_blank"
rel="noopener"}.

Aiming to become a home to research from the growing number of
biological and biomedical fields handling \"big-data\",
[*GigaScience*](http://www.gigasciencejournal.com/ "GigaScience homepage"){target="_blank"
rel="noopener"} is a new type of open access, open data journal that
provides standard scientific publishing linked directly to a database
that hosts its relevant data. Our associated
[*Giga*DB](http://gigadb.org/ "GigaDB homepage"){target="_blank"
rel="noopener"} database,
[launched](http://blogs.biomedcentral.com/gigablog/2012/07/gigascience_at_icg6_announcing_the "Launch blog"){target="_blank"
rel="noopener"} last year, provides a home to all of the supporting data
and tools associated with research, thus overcoming one of the biggest
challenges holding back reproducible research. Through
[*Giga*DB](http://gigadb.org/ "GigaDB homepage"){target="_blank"
rel="noopener"}, we assign
[DataCite](http://datacite.org/ "DataCite homepage"){target="_blank"
rel="noopener"}
[DOIs](http://en.wikipedia.org/wiki/Digital_object_identifier "What are DOIs: wikipedia"){target="_blank"
rel="noopener"} to these accompanying datasets to provide additional
credit to the authors who make their data publicly available and to
boost data discoverability and tracking of data. Data citation is
important in incentivizing the effort needed to present data to the
public in a usable form, and recent successes by others and us in
promoting its use were highlighted in
[this](http://www.biomedcentral.com/1756-0500/5/223 "Adventures in Data Citation paper"){target="_blank"
rel="noopener"} recent commentary. Our commitment to data DOIs also
means that we are well placed to take advantage of the upcoming data
citation index
[announced](http://www.reuters.com/article/2012/06/22/idUS109861+22-Jun-2012+HUG20120622 "TR data citiation index announcement"){target="_blank"
rel="noopener"} recently by Thomson Reuters.

Using the data hosting capabilities and expertise in data handling and
[cloud
computing](https://www.easygenomics.com/index "BGI Easy Genomics homepage"){target="_blank"
rel="noopener"} of our partner,
[BGI](http://en.genomics.cn/navigation/index.action "BGI homepage"){target="_blank"
rel="noopener"}, we are able to host a much larger and broader range of
datasets than journal supplementary files are usually able to handle,
outside the capacities of most other journals and repositories. We have
been testing and building up our new informatics platform by releasing a
number of datasets from BGI to demonstrate new mechanisms of
pre-publication data-release. For example, the deadly 2011 outbreak *E.
coli* dataset was the [first we
released](http://blogs.biomedcentral.com/gigablog/2011/08/03/notes-from-an-e-coli-tweenome-lessons-learned-from-our-first-data-doi/ "Notes from a tweenome"){target="_blank"
rel="noopener"}, and this led to the crowdsourcing of its analysis,
which was
[cited](http://royalsociety.org/policy/projects/science-public-enterprise/report/ "RS report page"){target="_blank"
rel="noopener"} in the recent Royal Society \"Science as an Open
Enterprise\" report as an example of \"The power of intelligently open
data\".

Now hosting datasets up to 14TB in size (such as
[this](http://dx.doi.org/10.5524/100034 "88 HCC cancer genomes"){target="_blank"
rel="noopener"} enormous resource of 88 tumor-normal paired genomes) and
containing several datasets from articles currently in press (including
[this](http://dx.doi.org/10.5524/100037 "Single cell cancer data"){target="_blank"
rel="noopener"} cancer single-cell genome data that will be published
shortly in *GigaScience*), we are providing examples of how data and
papers can be combined in our launch issue today. Exemplifying
*GigaScience* and *Giga*DB\'s innovative approach is a research
[article](http://www.gigasciencejournal.com/content/1/1/3 "Methylome paper"){target="_blank"
rel="noopener"} from Stephan Beck\'s
[group](http://www.ucl.ac.uk/cancer/medical-genomics/mg_staff "S Beck UCL group page"){target="_blank"
rel="noopener"} at UCL focusing on ways to conduct whole-genome analyses
of DNA methylation. In addition to having the raw data available in
[NCBI](http://trace.ncbi.nlm.nih.gov/Traces/sra/?study=SRP005934 "NCBI raw data: SRP005934"){target="_blank"
rel="noopener"}, all of the supporting data and software tools needed to
recreate the experiments --- a total of 84 GB --- are freely available
for
[download](http://dx.doi.org/10.5524/100035 "GigaDB methylome paper data"){target="_blank"
rel="noopener"} and reuse under the most open
[CC0](http://creativecommons.org/publicdomain/zero/1.0/ "CC0 licence"){target="_blank"
rel="noopener"} public domain waiver from *Giga*DB. We are also
maximizing data-interoperability and re-use by encouraging and rewarding
authors for providing much richer metadata. For this dataset, metadata
are available in ISA-tab format -- the first time a journal has taken
ISA-compliant submissions (for more, see the
[recent](http://www.nature.com/ng/journal/v44/n2/full/ng.1054.html "ISA-community Nature Genetics paper"){target="_blank"
rel="noopener"} ISA-commons community paper in *Nature Genetics* we
contributed to).

The *[Giga](http://gigadb.org/ "GigaDB homepage"){target="_blank"
rel="noopener"}*[DB](http://gigadb.org/ "GigaDB homepage"){target="_blank"
rel="noopener"} website is continuing to evolve and the next version
will be released in a few months with a more extensive search interface.
As the final step in attaining fully executable and reproducible papers,
we will be working with authors to make the computational tools and data
processing pipelines described in their papers available and, where
possible, executable on an informatics platform we are developing with
[collaborators](http://www.cuhk.edu.hk/cbiit/research.html "Tin-Lap Lee, CUHK-BGI trans-omic collaboration page"){target="_blank"
rel="noopener"} at the Chinese University of Hong Kong. We hope that by
making both the data and processes involved in their analysis freely
accessible, this novel form of publication will help articles published
in our journal to have a much higher impact in the scientific literature
and to maximize their reuse within the community. Check out our two
editorials for more on our goals with the
[database](http://www.gigasciencejournal.com/content/1/1/11 "GigaDB editorial"){target="_blank"
rel="noopener"} and
[journal](http://www.gigasciencejournal.com/content/1/1/1 "GigaScience editorial"){target="_blank"
rel="noopener"}.

As well as this innovative, big-data-driven publication format, the
journal also provides reviews and commentaries that address the many
hurdles that still need to be surmounted to improve future big-data
handling. Many of these are part of our first [thematic series (GSC and
beyond)](http://www.gigasciencejournal.com/series/GSC_and_beyond "GSC series page"){target="_blank"
rel="noopener"} covering the best practices in genomics research,
published [in
concert](http://blogs.biomedcentral.com/gigablog/2012/07/gigascience_at_the_genomic_standards "GSC series call-for-papers blog posting"){target="_blank"
rel="noopener"} with the [Genomic Standards
Consortium](http://gensc.org/gc_wiki/index.php/Main_Page "GSC homepage"){target="_blank"
rel="noopener"}. In addition to commentaries discussing data-sharing
issues in
[neuroimaging](http://www.gigasciencejournal.com/content/1/1/9 "Neuroimaging paper"){target="_blank"
rel="noopener"}and
[genomics](http://www.gigasciencejournal.com/content/1/1/10 "Sansone GSC paper"){target="_blank"
rel="noopener"}, our launch issue has a more detailed
[review](http://www.gigasciencejournal.com/content/1/1/2 "Data compression/archiving review"){target="_blank"
rel="noopener"} tackling data-compression from Guy Cochrane and Ewan
Birney at the EBI, a white-paper on DNA collection for large vertebrate
genome projects, and two papers on using genomics data to characterize
ecosystem and disease outbreaks. We also have a software/technical note
paper detailing a novel data format that facilitates the
interoperability of bioinformatics tools and an associated commentary
from Jonathan Eisen on \'badomics\' terminology and the explosion of
\"omes\" (good and
[bad](http://www.ark-genomics.org/badomics-generator "Badomics generator"){target="_blank"
rel="noopener"}), a topic that readers of his
[blog](http://phylogenomics.blogspot.com/search/label/bad%20omics%20word%20of%20the%20day "#badomics"){target="_blank"
rel="noopener"} will well know.

We hope you enjoy our first set of articles, and please follow our
progress here in our blog, our various social media outlets, and on the
journal
[homepage](http://www.gigasciencejournal.com/ "GigaScience homepage"){target="_blank"
rel="noopener"}. We would like to thank the BGI and our collaborators at
the Chinese University of Hong Kong, British Library, DataCite, and
ISA-Tab for their help and support building the data platform. We\'d
also like to thank our authors, reviewers and fantastic Editorial Board,
as well as [BioMed
Central](http://www.biomedcentral.com/ "BMC homepage"){target="_blank"
rel="noopener"} for their part in getting this issue out. BGI is
generously covering the open access [article processing
charges](http://www.biomedcentral.com/about/apcfaq "APC FAQ"){target="_blank"
rel="noopener"} for the journal\'s first year, so please contact us at
<editorial@gigasciencejournal.com> if you have related work you would
like to submit to this series or journal; alternatively submit a
manuscript [here](http://www.gigasciencejournal.com/manuscript). This
week we will be at the
[ISMB](http://www.iscb.org/ismb2012 "ISMB 2012"){target="_blank"
rel="noopener"} meeting in Long Beach, so come and get hold of us at the
BMC booth (#36) if you\'d like to meet the editors.

The post [GigaScience launches: overseeing the transition from papers to
\"executable research
objects\"](http://gigasciencejournal.com/blog/gigascience-launches-overseeing-the-transition-from-papers-to-executable-research-objects/){rel="nofollow"}
appeared first on
[GigaBlog](http://gigasciencejournal.com/blog){rel="nofollow"}.