---
author:
- contributor_roles: []
  family: Edmunds
  given: Scott
  url: https://orcid.org/0000-0001-6444-1436
blog:
  authors: null
  community_id: 52db0518-e228-4260-8c54-c4e323b2569d
  created: 1675555200
  current_feed_url: null
  description: Data driven blogging from the GigaScience editors
  doi: https://doi.org/10.59350/gigablog
  favicon: https://rogue-scholar.org/api/communities/52db0518-e228-4260-8c54-c4e323b2569d/logo
  feed_format: application/atom+xml
  feed_url: http://gigasciencejournal.com/blog/feed/atom/
  filter: null
  generator: Other
  home_page_url: https://gigasciencejournal.com/blog
  issn: null
  language: eng
  license: https://creativecommons.org/licenses/by/4.0/legalcode
  prefix: '10.59350'
  relative_url: null
  secure: false
  slug: gigablog
  status: archived
  subfield: '1311'
  title: GigaBlog
  updated: null
  use_api: null
container: GigaBlog
date: '2012-12-27T00:00:00+00:00'
date_updated: '2025-12-06T10:47:52+00:00'
guid: http://finaloriginalblogs.dev/gigablog/?p=355
identifier: https://doi.org/10.59350/jrf70-d5t78
image: http://gigasciencejournal.com/blog/wp-content/uploads/2012/12/ScientificReview-300x168.jpg
images:
- height: '168'
  src: http://gigasciencejournal.com/blog/wp-content/uploads/2012/12/ScientificReview-300x168.jpg
  width: '300'
- src: http://gigasciencejournal.com/blog/wp-content/uploads/2012/12/ScientificReview.jpg
issn: null
keywords:
- Open Access
- Data Citation
- GigaDB
- GigaScience
- Open Peer Review
lang: en
license: https://creativecommons.org/licenses/by/4.0/legalcode
rid: yd18a-ze045
rights: https://creativecommons.org/licenses/by/4.0/legalcode
summary: With everyone in a reflective mood as the year comes to a close, one of the
  big scientific trends of 2012 has obviously been the high profile that open-access
  and more open methods of carrying out science has received.
title: 'Opening peer-review: our new paper on SOAPdenovo2 shows how it works'
url: https://wayback.archive-it.org/22098/2025-05-01T17:13:42Z/http://gigasciencejournal.com/blog/opening-peer-review-our-new-paper-on-soapdenovo2-shows-how-it-works
version: v1
---

[![](http://gigasciencejournal.com/blog/wp-content/uploads/2012/12/ScientificReview-300x168.jpg){.alignleft
.size-medium .wp-image-357 loading="lazy" decoding="async" width="300"
height="168"}](http://gigasciencejournal.com/blog/wp-content/uploads/2012/12/ScientificReview.jpg)With
everyone in a reflective mood as the year comes to a close, one of the
big scientific trends of 2012 has obviously been the high profile that
open-access and more open methods of carrying out science has received.
With the [Elsevier
boycott](http://thecostofknowledge.com/ "Boycott Elseview"){target="_blank"
rel="noopener"}, UK [Finch
report](http://www.guardian.co.uk/science/2012/jun/19/open-access-academic-publishing-finch-report "Finch report in Guardian"){target="_blank"
rel="noopener"}, and launch of a number of innovative new schemes in
publishing open-access research and data (including [*F1000
Research*](http://f1000research.com/ "F1000 Research"){target="_blank"
rel="noopener"},
[*eLife*](www.elifesciences.org/ "eLife"){target="_blank"
rel="noopener"}, [*PeerJ*](https://peerj.com/ "PeerJ"){target="_blank"
rel="noopener"} and of course
*[GigaScience](http://www.gigasciencejournal.com "GigaHomepage"){target="_blank"
rel="noopener"}*), 2012 has been talked of as the year of an [\"academic
spring\"](http://en.wikipedia.org/wiki/Academic_Spring "Academic Spring"){target="_blank"
rel="noopener"} that has started to shake up the centuries old, stuffy
and closed system of scientific discourse.

On top of changes to the way scientists and readers are demanding they
can access and mine the literature and data, new incentives and
mechanisms to release and publish data (of which we have [written
extensively](http://blogs.biomedcentral.com/gigablog/tag/data-citation/ "GigaBlog on data citation"){target="_blank"
rel="noopener"}), the process of peer-review has also come under the
spotlight, and there has been a lot of talk on the
[deficiencies](http://www.michaeleisen.org/blog/?p=694 "Eisen blog"){target="_blank"
rel="noopener"} [of this
system](http://www.guardian.co.uk/science/lost-worlds/2012/dec/01/dinosaurs-fossils "Guardian blog on peer review"){target="_blank"
rel="noopener"}. Many newly launched journals have tried to make the
system more transparent, using systems such as post-publication
peer-review (e.g. [F1000
Research](http://f1000research.com/about/ "F1000 Research about page"){target="_blank"
rel="noopener"}), pre-print servers (e.g. the increasing acceptance of
[arXiv](http://arxiv.org/ "arXiv"){target="_blank" rel="noopener"} in
biology), providing access to anonymized (e.g. [EMBO
journals](http://www.nature.com/emboj/about/process.html#Transparent_Process "EMBO J peer review"){target="_blank"
rel="noopener"}) or partial
([elife](http://www.elifesciences.org/the-journal/review-process/ "eLife review process"){target="_blank"
rel="noopener"}) parts of the peer-review history, or encouraging
reviewers to opt-into open peer-review
([*PeerJ*](https://peerj.com/about/policies-and-procedures/#open-peer-review "PeerJ open review policies"){target="_blank"
rel="noopener"}, and [experimented
with](http://www.plosone.org/static/reviewerGuidelines;jsessionid=A48FAA073462674555394D0178417DB9#anonymity){target="_blank"
rel="noopener"} a little at *PLOS*). At *GigaScience* we have decided to
take this process one step further and ask for [open peer-review as
default](http://www.gigasciencejournal.com/about/reviewers "GigaReviewers page"){target="_blank"
rel="noopener"}, and as our aims are to promote more open, reproducible
and transparent-science we feel it promotes accountability, fairness,
and importantly gives credit to reviewers for their hard efforts. A [new
publication](http://www.gigasciencejournal.com/content/1/1/18/abstract "SOAPdenovo2 paper"){target="_blank"
rel="noopener"} in the journal today provides a particularly useful
example of how this process has worked, so we have decided to highlight
it here in
[GigaBlog](http://blogs.biomedcentral.com/gigablog/ "GigaBlog"){target="_blank"
rel="noopener"}, and would welcome feedback and comments on our
approach.

**What is SOAPdenovo2?**\
Today we
[publish](http://www.gigasciencejournal.com/content/1/1/18/abstract "SOAPdenovo2 paper"){target="_blank"
rel="noopener"} an updated version of BGI\'s popular SOAPdenovo software
application (the [original
version](http://genome.cshlp.org/content/20/2/265.short "SOAPdenovo1 paper"){target="_blank"
rel="noopener"} having 460 citations according to
[googlescholar](http://scholar.google.com.hk/scholar?cites=11447276992969970821&as_sdt=2005&sciodt=0,5&hl=en "google scholar citations"){target="_blank"
rel="noopener"}), a start of the art tool for *de novo* genome assembly.
*De novo* assembly -- piecing together genomes from sequencing data
without the aid of a previously assembled reference, is a particularly
computationally intensive and technically challenging task. BGI and
their SOAPdenovo tool has been particularly adept in this area, using it
to assemble hundreds of new plant and animal species genomes, as well as
finding huge amounts of [previously undetected structural
changes](http://www.nature.com/nbt/journal/v29/n8/full/nbt.1904.html "Nature Biotech SV paper"){target="_blank"
rel="noopener"} when applied to individual human genomes. *De novo*
genome assembly is an important and competitive area in bioinformatics,
and there have been a number of assembly competitions and genome
assembler \"bake offs\" to compare and benchmark the various
applications and methods available for carrying this out, the
[Assemblathon](http://assemblathon.org/ "Assemblathon"){target="_blank"
rel="noopener"} and
[GAGE](http://gage.cbcb.umd.edu/ "GAGE homepage"){target="_blank"
rel="noopener"} assembly competitions and evaluations being notable
examples of this.

New developments in version 2 of SOAPdenovo have focused on using more
efficient algorithms and data structures to reduce the memory
requirements, better optimizing and handling of errors and low coverage
or heterozygous regions, as well as improved closing of gaps. To
demonstrate the improvements and that the application truly is the
state-of-the-art for de novo assembly of large vertebrate genomes, the
authors reassembled BGI\'s [YH Asian reference
genome](http://yh.genomics.org.cn/ "YH homepage"){target="_blank"
rel="noopener"} with the new and original versions of SOAPdenovo,
version 2 producing contig sizes 3 times larger, and with nearly two
thirds of the maximum memory consumption. Doing comparisons against
other state of the-art assemblers such as ALLPATHS-LG, SOAPdenovo2
outperformed them for many metrics on the Assemblathon and GAGE
benchmark datasets tested, really showcasing and demonstrating the
potential power and utility of this new application for the
bioinformatics community.

**Open peer-review, *GigaScience* style**\
Stating that SOAPdenevo2 can perform better than other state-of-the-art
assembly tools is one thing, but to justify and prove this review and
testing by independent peers is needed, and the larger and more
complicated an application is (particularly an issue for us being a
journal that focuses on data heavy research studies), the more
challenging this can be. In order to ease, throw light and credit the
reviewers in this process *GigaScience* uses a much more transparent,
accountable and open peer-review process. Tailoring the process for such
data heavy studies our [criteria for
publication](http://www.gigasciencejournal.com/about/reviewers "GigaReviewers page"){target="_blank"
rel="noopener"} is based more on relative amount of data created or
used, and transparency and availability more than subjective and
unpredictable measures such as supposed \"impact\". For software and
methods papers what is being presented obviously has to be an
improvement on what is currently available, but for scenarios such as
genome assembly where there is obviously no \"one-size-fits-all\"
solution, assessment has to be based on the new method being an
improvement in at least one potential application.

During peer review we host all of the supporting information and data
(totaling 78GB in this case) and our curators work and make all of it
available to the peer-reviewers from our ftp servers. In this case we
worked with three groups of expert reviewers (8 independent experts in
total) who thoroughly tested the software against various tools and
datasets provided to ensure the claims made by the authors were correct.
On top of providing all of the test data and scripts and tools that
support the paper, to aid the process the authors also provide detailed
pipelines with the tools and configured packages including commands and
necessary utilities to reproduce the different tests carried out in the
paper.

Whilst used in a number of medical journals, almost unprecedentedly in
biology we ask as default all of the reviewers to carry out open
peer-review, and in this case all 8 of them consented and signed their
names to the reports that are now [available to
view](http://www.gigasciencejournal.com/content/1/1/18/prepub "SOAPdenovo prepub history"){target="_blank"
rel="noopener"} from the pre-publication history section associated with
our published articles. To see how this looks you can follow the history
of the SOAPdenovo2 paper
[here](http://www.gigasciencejournal.com/content/1/1/18/prepub "SOAPdenovo prepub history"){target="_blank"
rel="noopener"}.

While we do have the option for reviewers to opt-out and anonymize their
reports if they have concerns about this process, it is encouraging that
for all of the papers we have reviewed so far none have asked to do
this. A number of new journals are starting to encourage reviewers to
sign their reports, but this is the default option for *GigaScience*,
with the option of opting out if the referees have reasons to remain
anonymous. We also give the reviewers the option of making confidential
comments to the editors (particularly on ethical and policy issues), but
so far the quality of the reports has generally been very constructive,
and
[previous](http://bjp.rcpsych.org/content/176/1/47.long "Open peer review RCT paper"){target="_blank"
rel="noopener"}
[studies](http://www.bmj.com/content/318/7175/23?view=long&pmid=9872878 "BMJ open peer review quality paper"){target="_blank"
rel="noopener"} on open peer-review have also found that quality and
courteousness of reviews were increased, with little if any negative
effects. By making the process more open and transparent competing
interests and biases are reduced, and reviewers are able to take credit
for the hard efforts they have put into the review process, and even
declare and include it in their CV if they wish as we would like to put
the content of accepted papers reviews under a [CC-BY
license](http://creativecommons.org/licenses/by/3.0/ "CC-BY"){target="_blank"
rel="noopener"}. The benefits of this increased transparency to readers
are also useful, as they do not have to take it on trust that published
manuscripts were reviewed by qualified reviewers, and for educational
purposes they can see good examples of how peer review operates.

**Promoting reproducibility, *GigaScience* style**\
On top of boosting transparency and reproducibility of peer-review of
data-heavy studies, *GigaScience* also carries this over to the
publication process, and this paper is also an excellent example of this
goal. On top of SOAPdenovo2 meeting our requirements of being open
source and having its code in a repository
([sourceforge](http://soapdenovo2.sourceforge.net/ "SOAPdenovo2 sourceforge page"){target="_blank"
rel="noopener"}), the authors also provide detailed pipelines with the
tools and configured packages including commands and necessary utilities
to reproduce the different tests carried out in the paper. With 78GB of
test data and 30MB of tools and scripts being much larger than any other
journal is able to handle, we have made all of these available from our
[GigaDB database](http://gigadb.org/ "GigaDB"){target="_blank"
rel="noopener"} as separate citable DOIs. Taking this a step further, on
top of being able to be downloaded by ftp and our
[Aspera](http://asperasoft.com/ "Aspera"){target="_blank"
rel="noopener"} license (allowing up to 10-100X faster access), the
software and analyses are also currently being integrated into our
Galaxy-workflow system based data platform.

While we have previously published software articles and pipeline
studies combining reference datasets and tools before (see our
[methylome pipeline
paper](http://www.gigasciencejournal.com/content/1/1/3 "Mouse methylome paper"){target="_blank"
rel="noopener"} with 84GB of [supporting
information](http://dx.doi.org/10.5524/100035 "Mouse methylome GigaDB data"){target="_blank"
rel="noopener"}), this is the first paper that we have given separate
DOIs to the
[tools](http://dx.doi.org/10.5524/100044 "software DOI"){target="_blank"
rel="noopener"} and
[data](http://dx.doi.org/10.5524/100038 "V2 YH genome DOI"){target="_blank"
rel="noopener"}. The logic for doing this is that both can now be
credited to potentially different groups of authors, the data and
methods/analyses may be used and cited independently of each other, and
each can be tracked and credited to each author via DOIs listed in their
[ORCID](http://about.orcid.org/ "ORCID about page"){target="_blank"
rel="noopener"} account. We feel that it is important to credit method
as well as data production, and while
[DataCite](http://datacite.org/ "DataCite"){target="_blank"
rel="noopener"} currently recognizes \"Software\" as a resource type, we
are encouraging and working with them to add \"Worklow\" to their list
of handled objects.

Work is ongoing on the workflow and data platform side, but we are
currently in the process of reviewing a number of other software
articles (called [Technical
Note](http://www.gigasciencejournal.com/authors/instructions/technicalnote "Technical Note I4A"){target="_blank"
rel="noopener"} in *GigaScience*), and if you have similar studies you
are interested in having reviewed in a more transparent and constructive
manner please contact us at editorial@gigasciencejournal or submit it
through our [online submission
system](http://www.gigasciencejournal.com/manuscript "Submit to GigaScience"){target="_blank"
rel="noopener"}. As process is still currently evolving and being
fine-tuned we would welcome any feedback via this blog,
[twitter](https://twitter.com/gigascience "@gigascience"){target="_blank"
rel="noopener"} or email. We would like to thank the team of reviewers
of
[this](http://www.gigasciencejournal.com/content/1/1/18/prepub "SOAPdenovo prepub history"){target="_blank"
rel="noopener"} and our other manuscripts so far for their hard efforts,
as well as the authors for being so helpful in making their work and
data available in such a reproducible manner. Many journals are
tentatively starting to experiment going down a partially more open
route, but from our positive experiences so far we would encourage them
and others to be more bold and go all of the way.

**Further Reading**

1.  Luo R et al., SOAPdenovo2: an empirically improved memory-efficient
    short-read de novo assembler *GigaScience* 2012, **1**:18\
    [2.](http://www.bmj.com/content/318/7175/23?view=long&pmid=9872878 "BMJ open peer review quality paper"){target="_blank"
    rel="noopener"} van Rooyen S et al., Effect of open peer review on
    quality of reviews and on reviewers\' recommendations: a randomised
    trial. *BMJ* 1999, **318**:23-7\
    [3.](http://bjp.rcpsych.org/content/176/1/47.long "Open peer review RCT paper"){target="_blank"
    rel="noopener"} Walsh E et al., Open peer review: a randomised
    controlled trial. *Br J Psychiatry* 2000, **176**:47-51.\
    [4.](http://dx.doi.org/10.5524/100038 "V2 YH genome DOI"){target="_blank"
    rel="noopener"} Wang, J; et al., (2012): Updated genome assembly of
    YH: the first diploid genome sequence of a Han Chinese individual
    (version 2, 07/2012). GigaScience Database.
    [http://dx.doi.org/10.5524/100038](http://dx.doi.org/10.5524/100038 "V2 YH genome DOI"){target="_blank"
    rel="noopener"}\
    [5.](http://dx.doi.org/10.5524/100044 "software DOI"){target="_blank"
    rel="noopener"} Luo, R; et al., (2012): Software and supporting
    material for \"SOAPdenovo2: An empirically improved memory-efficient
    short read de novo assembly\". GigaScience Database.
    [http://dx.doi.org/10.5524/100044](http://dx.doi.org/10.5524/100044 "software DOI"){target="_blank"
    rel="noopener"}

UPDATE 24th Jan 2013: we have [produced an
editorial](http://www.gigasciencejournal.com/content/2/1/1 "GigaScience peer-review editorial"){target="_blank"
rel="noopener"} on our peer-review policies based on this blog and the
feedback we received on it. Also check out the great work the Homolog_us
blog has done [testing and studying
SOAPdenovo2](http://www.homolog.us/blogs/category/soapdenovo/ "Homolog_us blog SOAPdenovo2 tag"){target="_blank"
rel="noopener"} making the source-code even more transparent via this
wiki:
[http://homolog.us/wiki1/index.php?title=SOAPdenovo2](http://homolog.us/wiki1/index.php?title=SOAPdenovo2 "SOAPdenovo2 wiki"){target="_blank"
rel="noopener"}

The post [Opening peer-review: our new paper on SOAPdenovo2 shows how it
works](http://gigasciencejournal.com/blog/opening-peer-review-our-new-paper-on-soapdenovo2-shows-how-it-works/){rel="nofollow"}
appeared first on
[GigaBlog](http://gigasciencejournal.com/blog){rel="nofollow"}.