---
author:
- contributor_roles: []
  family: Edmunds
  given: Scott
  url: https://orcid.org/0000-0001-6444-1436
blog:
  authors: null
  community_id: 52db0518-e228-4260-8c54-c4e323b2569d
  created: 1675555200
  current_feed_url: null
  description: Data driven blogging from the GigaScience editors
  doi: https://doi.org/10.59350/gigablog
  favicon: https://rogue-scholar.org/api/communities/52db0518-e228-4260-8c54-c4e323b2569d/logo
  feed_format: application/atom+xml
  feed_url: http://gigasciencejournal.com/blog/feed/atom/
  filter: null
  generator: Other
  home_page_url: https://gigasciencejournal.com/blog
  issn: null
  language: eng
  license: https://creativecommons.org/licenses/by/4.0/legalcode
  prefix: '10.59350'
  relative_url: null
  secure: false
  slug: gigablog
  status: archived
  subfield: '1311'
  title: GigaBlog
  updated: null
  use_api: null
container: GigaBlog
date: '2013-08-09T00:00:00+00:00'
date_updated: '2025-12-06T10:42:59+00:00'
guid: http://finaloriginalblogs.dev/gigablog/?p=721
identifier: https://doi.org/10.59350/4sc41-9fz56
image: http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/big_data_colour1.png
images:
- height: '842'
  sizes: '(max-width: 595px) 100vw, 595px'
  src: http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/big_data_colour1.png
  srcset: http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/big_data_colour1.png,
    http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/big_data_colour1-212x300.png
  width: '595'
- height: '225'
  sizes: '(max-width: 300px) 100vw, 300px'
  src: http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/Screen-shot-2013-08-10-at-12.04.22-AM-300x225.png
  srcset: http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/Screen-shot-2013-08-10-at-12.04.22-AM-300x225.png,
    http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/Screen-shot-2013-08-10-at-12.04.22-AM.png
  width: '300'
- src: http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/big_data_colour1.png
- src: http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/Screen-shot-2013-08-10-at-12.04.22-AM.png
issn: null
keywords:
- Open Access
- Publishing
- Berlin
- Conferences
- ECCB
lang: en
license: https://creativecommons.org/licenses/by/4.0/legalcode
rid: 0e5gm-ss456
rights: https://creativecommons.org/licenses/by/4.0/legalcode
summary: <strong> Big Data Publishing </strong> (credit Jenny Cham, CC-BY) As mentioned
  in our previous posting, on top of the many great talks and sessions we attended
  at ISMB in Berlin last month, we were kept even busy helping to organize and present
  in a special Beyond-the-PDF inspired "What Bioinformaticians need to know about
  digital publishing beyond the PDF" workshop.
title: 'More on our ISMB workshop: What Bioinformaticians need to know about digital
  publishing beyond the PDF'
url: https://wayback.archive-it.org/22098/2025-05-01T17:13:42Z/http://gigasciencejournal.com/blog/more-on-our-ismb-workshop-what-bioinformaticians-need-to-know-about-digital-publishing-beyond-the-pdf
version: v1
---

**Big Data Publishing** (credit [Jenny
Cham](http://www.flickr.com/photos/97823772@N02/9390887593/in/set-72157634564701268 "Jenny Cham Flickr picture"){target="_blank"
rel="noopener"}, CC-BY)\
[![](http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/big_data_colour1.png){.alignleft
.size-full .wp-image-734 loading="lazy" decoding="async"
srcset="http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/big_data_colour1.png 595w, http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/big_data_colour1-212x300.png 212w"
sizes="(max-width: 595px) 100vw, 595px" width="595"
height="842"}](http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/big_data_colour1.png)As
mentioned in our [previous
posting](http://blogs.biomedcentral.com/gigablog/2013/08/01/bioinformaticians-breaking-down-barriers-in-berlin/ "ISMB 2013 blog"){target="_blank"
rel="noopener"}, on top of the many great talks and sessions we attended
at
[ISMB](http://www.iscb.org/ismbeccb2013 "ISMB/ECCB homepage"){target="_blank"
rel="noopener"} in Berlin last month, we were kept even busy helping to
organize and present in a special
[Beyond-the-PDF](http://www.force11.org/beyondthepdf2 "#BtPDF2 homepage"){target="_blank"
rel="noopener"} inspired \"[What Bioinformaticians need to know about
digital publishing beyond the
PDF](http://www.iscb.org/cms_addon/conferences/ismbeccb2013/workshops.php#WK03 "ISMB workshops program"){target="_blank"
rel="noopener"}\" workshop. Most of the heavy lifting has to be credited
to the hard work of Marco Roos, but the other organisers included Oscar
Corcho, Carole Goble, Barend Mons, Jun Zhao and Erik Schultes, and we
had a ridiculously overqualified list of speakers, panelists and
supporters that help make it a success.

You can see the blurb and line-up on the [ISMB
website](http://www.iscb.org/cms_addon/conferences/ismbeccb2013/workshops.php#WK03 "ISMB workshop program"){target="_blank"
rel="noopener"}, but the main aim of the session was to inform
participants of changes and new opportunities in scientific
communication. While there have been a lot of recent developments
spurred by related future of scholarly communication conferences and
projects (see
[Force11](http://www.force11.org/ "Force 11 website"){target="_blank"
rel="noopener"}), we wanted to take these sorts of discussions out of
the usual publishing crowd, and present them to researchers who may
potential use them. As mentioned in our [previous
posting](http://blogs.biomedcentral.com/gigablog/2013/08/01/bioinformaticians-breaking-down-barriers-in-berlin/ "ISMB 2013 blog"){target="_blank"
rel="noopener"}, ISMB and the computational biology community are a
particularly receptive and positive audience for open science as the
whole field has been built on data sharing and open-source tools, so we
wanted to provide visual guidelines and present tools that take this
open approach even further.

Phil Bourne brilliantly set the scene, giving an overview and his
thoughts on what the issues are, and how people can participate to make
a change (slides
[here](http://www.slideshare.net/pebourne/ismb2013 "Phil Bourne workshop talk sldies"){target="_blank"
rel="noopener"}). Giving some sense of the urgency that researchers need
to address the issue of making their data available, as few are prepared
for the implications of [new NIH open access
policy](http://guides.library.upenn.edu/content.php?pid=388232&sid=3181434 "PMCID policy"){target="_blank"
rel="noopener"} that starts being enforced next month (if their work
doesn\'t end up in PMC, grants will not get renewed), and similar
mandates for data will be here sooner than most are prepared for too.
Second on the program was Rebecca Lawrence from
[F1000](http://f1000.com/ "F1000"){target="_blank" rel="noopener"},
talking about data publishing and data peer-review initiatives. Showing
some of the great examples
[F1000Research](http://f1000research.com/ "F1000 Research website"){target="_blank"
rel="noopener"} have done in this area, their post publication
peer-review pipeline has managed to review papers and take them from
submission to publication in as little as 34 hours. If you have seen our
[recent
postings](http://blogs.biomedcentral.com/gigablog/2013/07/25/genome-assembly-in-the-spotlight/ "Assemblathon2 blog"){target="_blank"
rel="noopener"} on the unusual peer-review of our [Assemblathon2
paper](http://www.gigasciencejournal.com/content/2/1/10/abstract "Assemblathon2 paper"){target="_blank"
rel="noopener"}, like us, they have also had very positive experiences
with open peer review. Also touching on data review, Rebecca has been
very closely involved with the [JISC
PREPARDE](http://proj.badc.rl.ac.uk/preparde "Preparde homepage"){target="_blank"
rel="noopener"} (Peer REview for Publication & Accreditation of Research
data in the Earth sciences) and some of this work was presented.

**GigaScience meets ISA, RO and Nanopublications**\
Scott\'s
[talk](http://www.slideshare.net/GigaScience/scott-edmunds-ismb-talk-on-big-data-publishing "Scott Edmunds Slides"){target="_blank"
rel="noopener"} on \"Big Data Publishing\" in the session covered work
on deconstructing the paper to reward reproducibility, deposition and
transparency of data, methods and analyses. Using our [SOAPdenovo2
paper](http://www.gigasciencejournal.com/content/1/1/18 "SOAPdenovo2 paper"){target="_blank"
rel="noopener"} as an example to we show how we can issue separate DOIs
resolving to all of the associated
[data](http://dx.doi.org/10.5524/100038 "Data DOI"){target="_blank"
rel="noopener"} as well as the [scripts, pipelines and
workflows](http://dx.doi.org/10.5524/100044 "Software DOI"){target="_blank"
rel="noopener"} we are hosting.

[![](http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/Screen-shot-2013-08-10-at-12.04.22-AM-300x225.png){.alignleft
.size-medium .wp-image-731 loading="lazy" decoding="async"
srcset="http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/Screen-shot-2013-08-10-at-12.04.22-AM-300x225.png 300w, http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/Screen-shot-2013-08-10-at-12.04.22-AM.png 636w"
sizes="(max-width: 300px) 100vw, 300px" width="300"
height="225"}](http://gigasciencejournal.com/blog/wp-content/uploads/2013/08/Screen-shot-2013-08-10-at-12.04.22-AM.png)

::: {style="margin-bottom:5px"}
**[Scott Edmunds ISMB talk on Big Data
Publishing](http://www.slideshare.net/GigaScience/scott-edmunds-ismb-talk-on-big-data-publishing "Scott Edmunds ISMB talk on Big Data Publishing"){target="_blank"
rel="noopener"}** from **[GigaScience, BGI Hong
Kong](http://www.slideshare.net/GigaScience){target="_blank"
rel="noopener"}**
:::

We have already presented in the past on our
[GigaDB](gigadb.org/ "GigaDB homepage"){target="_blank" rel="noopener"}
database and
[GigaGalaxy](http://galaxy.cbiit.cuhk.edu.hk/ "GigaGalaxy"){target="_blank"
rel="noopener"} data platform, but tied with this workshop and related
presentations at the meeting this was the first time we presenting some
initial findings from a case study we are carrying out with the
[ISA-TAB](http://isa-tools.org/ "ISA-tab homepage"){target="_blank"
rel="noopener"}, [Research Object
(RO)](www.researchobject.org/ "Research Objects"){target="_blank"
rel="noopener"} and
[Nanopublication](http://nanopub.org/wordpress/ "Nanopublications"){target="_blank"
rel="noopener"} communities to see how these three data models can
support representation of scholarly artifacts. Studying how
complimentary these models can be, how much value we can add to the
publishing process, as well as encourage their use by demoing them as
\"digital instruction for authors\", this work is very much still in
progress, but we were keen to show the results so far to the workshop
audience. We have been working with the ISA team for a while, using
their interoperable metadata format in a number of our datasets (see the
[methylated nematode
genome](http://dx.doi.org/10.5524/100043 "methylated nematode data"){target="_blank"
rel="noopener"} -- of which there will be an announcement next month),
and as we have been working with [Galaxy
workflows](https://main.g2.bx.psu.edu/){target="_blank" rel="noopener"}
the workflow-centric focus of the [Research
Object](www.researchobject.org/ "Research Objects"){target="_blank"
rel="noopener"} model made it an obvious system to trial in this case
study too.
[Nanopublications](http://nanopub.org "Nanopublications"){target="_blank"
rel="noopener"} are a new area to explore for us, but following our
\"deconstructed paper\" approach, being smallest the unit of publishable
information (an assertion), the ability to attribute and cite these
makes them attractive through their potential to provide incentives for
researchers to make their data available.

While the work on the case study so far has shown we can represent the
parts of the paper in these models, and that we can recreate results
from the paper using our
[GigaGalaxy](http://galaxy.cbiit.cuhk.edu.hk/ "GigaGalaxy"){target="_blank"
rel="noopener"} workflow, we are still struggling to recreate other
parts of the results and represent them as nanopublications. It was
great seeing our board member Carole Goble further elaborate on the
costliness of reproducibility in her ISMB keynote the following day (see
her [great
slides](http://www.slideshare.net/carolegoble/ismb2013-keynotecleangoble "Carole Goble Keynote"){target="_blank"
rel="noopener"}) and include some of this case study as examples to the
wider ISMB audience. The RO and Nanopublication work was elaborated upon
by Marco Roos in his Tech Track talk (see
[abstract](http://www.iscb.org/uploaded/css/148/28327.pdf "TT34 abstract"){target="_blank"
rel="noopener"}), and Mark Thompson in his
[poster](http://www.iscb.org/cms_addon/conferences/ismbeccb2013/posterlist.php?cat=B#B28 "B28 poster"){target="_blank"
rel="noopener"}, and for completeness sake we will try to post these
when they are made available online.

The session ended with a packed panel chaired by Barend Mons, and
including Niklas Blomberg (Director of
[Elixir](http://www.elixir-europe.org/ "Elixir homepage"){target="_blank"
rel="noopener"}), Carole Goble, Larry Hunter, Winston Hide, as well as
the speakers and organizers summarising the topic together and fielding
questions from the audience. Polling the audience at the end of the
session, Marco got a very positive response when asking whether to do
the workshop again [next year in
Boston](http://www.iscb.org/ismb2014 "ISMB 2014, Boston"){target="_blank"
rel="noopener"}, so watch this space for updates on the ongoing case
study, and if there will be a follow up workshop. We\'d like to thank
all of the case study participants, particularly Jun Zhao, Philippe
Rocca-Sera and Alejandra Gonzalez-Beltran at Oxford, and Mark Thompson
at LUMC, and Marco Roos in particular for doing most of the hard work
setting up the workshop, as well as the ISMB organisers for selecting it
and helping us make it happen. We\'d also like to thank Jennifer Cham at
the EBI for the [great
sketch](http://www.flickr.com/photos/97823772@N02/9390887593/in/set-72157634564701268 "Jenny Cham Flickr picture"){target="_blank"
rel="noopener"} featured above, and for making it CC-BY. You can see the
other fantastic sketches she made at the conference from the her [Flickr
page](http://www.flickr.com/photos/97823772@N02/sets/72157634564701268/ "Jenny Cham Flickr"){target="_blank"
rel="noopener"}.

The post [More on our ISMB workshop: What Bioinformaticians need to know
about digital publishing beyond the
PDF](http://gigasciencejournal.com/blog/more-on-our-ismb-workshop-what-bioinformaticians-need-to-know-about-digital-publishing-beyond-the-pdf/){rel="nofollow"}
appeared first on
[GigaBlog](http://gigasciencejournal.com/blog){rel="nofollow"}.