---
author:
- contributor_roles: []
  family: Edmunds
  given: Scott
  url: https://orcid.org/0000-0001-6444-1436
blog:
  authors: null
  community_id: 52db0518-e228-4260-8c54-c4e323b2569d
  created: 1675555200
  current_feed_url: null
  description: Data driven blogging from the GigaScience editors
  doi: https://doi.org/10.59350/gigablog
  favicon: https://rogue-scholar.org/api/communities/52db0518-e228-4260-8c54-c4e323b2569d/logo
  feed_format: application/atom+xml
  feed_url: http://gigasciencejournal.com/blog/feed/atom/
  filter: null
  generator: Other
  home_page_url: https://gigasciencejournal.com/blog
  issn: null
  language: eng
  license: https://creativecommons.org/licenses/by/4.0/legalcode
  prefix: '10.59350'
  relative_url: null
  secure: false
  slug: gigablog
  status: archived
  subfield: '1311'
  title: GigaBlog
  updated: null
  use_api: null
container: GigaBlog
date: '2020-04-08T00:00:00+00:00'
date_updated: '2025-12-06T10:19:48+00:00'
guid: http://gigasciencejournal.com/blog/?p=3274
identifier: https://doi.org/10.59350/zxcdd-ns360
image: http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/giaa026fig7-300x206.jpeg
images:
- alt: ShinyLearner
  height: '206'
  sizes: '(max-width: 300px) 100vw, 300px'
  src: http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/giaa026fig7-300x206.jpeg
  srcset: http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/giaa026fig7-300x206.jpeg,
    http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/giaa026fig7.jpeg
  width: '300'
- alt: ShinyLearner author
  height: '135'
  sizes: '(max-width: 135px) 100vw, 135px'
  src: http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/brightspotcdn.byu_.edu_-150x150.jpg
  srcset: http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/brightspotcdn.byu_.edu_-150x150.jpg,
    http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/brightspotcdn.byu_.edu_.jpg
  width: '135'
- alt: CODECHECK certificate
  height: '208'
  sizes: '(max-width: 222px) 100vw, 222px'
  src: http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/Screenshot-2020-04-06-at-4.27.42-PM-300x281.png
  srcset: http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/Screenshot-2020-04-06-at-4.27.42-PM-300x281.png,
    http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/Screenshot-2020-04-06-at-4.27.42-PM-768x721.png,
    http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/Screenshot-2020-04-06-at-4.27.42-PM-1024x961.png,
    http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/Screenshot-2020-04-06-at-4.27.42-PM.png
  width: '222'
issn: null
keywords:
- Technology
- Code Ocean
- Containers
- Docker
- Open Science
lang: en
license: https://creativecommons.org/licenses/by/4.0/legalcode
reference:
- id: https://doi.org/10.1093/gigascience/giaa026
  unstructured: 'Piccolo, S. R., Lee, T. J., Suh, E., &amp; Hill, K. (2020). ShinyLearner:
    A containerized benchmarking tool for machine-learning classification of tabular
    data. <i>GigaScience</i>, <i>9</i>(4).'
- id: https://doi.org/10.5281/zenodo.3674056
  unstructured: Eglen, S. J. (2020). <i>CODECHECK Certificate 2020-001</i>. Zenodo.
- id: https://doi.org/10.1186/s13742-016-0135-4
  unstructured: Piccolo, S. R., &amp; Frampton, M. B. (2016). Tools and techniques
    for computational reproducibility. <i>GigaScience</i>, <i>5</i>(1).
- id: http://gigasciencejournal.com/blog/shinylearner-codecheck/
  unstructured: Unknown title
- id: http://gigasciencejournal.com/blog
  unstructured: Unknown title
rid: r6hbh-9hm05
rights: https://creativecommons.org/licenses/by/4.0/legalcode
summary: This week we showcased a new way of peer reviewing software,&nbsp;testing
  code in an independent manner and providing a CODECHECK <em> "certificate of reproducible
  computation" </em> when the results in the paper can be reproduced.
title: Reproducible Classification. Q&A on ShinyLearner & the CODECHECK certificate,
  pt. 2
url: https://wayback.archive-it.org/22098/2025-05-01T17:13:42Z/http://gigasciencejournal.com/blog/shinylearner-codecheck
version: v1
---

![ShinyLearner](http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/giaa026fig7-300x206.jpeg){.alignright
.size-medium .wp-image-3275 loading="lazy" decoding="async"
srcset="http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/giaa026fig7-300x206.jpeg 300w, http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/giaa026fig7.jpeg 766w"
sizes="(max-width: 300px) 100vw, 300px" width="300" height="206"}This
week [we
showcased](http://gigasciencejournal.com/blog/codecheck-certificate/) a
new way of peer reviewing software, testing code in an independent
manner and providing a CODECHECK *\"certificate of reproducible
computation\"* when the results in the paper can be reproduced. We\'ve
[written a post on the CODECHECK
process](http://gigasciencejournal.com/blog/codecheck-certificate/)
featuring a Q&A with [CODECHECK](http://codecheck.org.uk/) founder
Stephen Eglan, and here we\'ll provide a follow up with a Q&A with the
author of the software that was subject to the
[CODECHECK](http://codecheck.org.uk/) certification process. Another
Stephen, author Stephen Piccolo talks here about his new paper on
ShinyLearner, a benchmarking tool for machine-learning classification
algorithms.

![ShinyLearner
author](http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/brightspotcdn.byu_.edu_-150x150.jpg){.alignleft
.wp-image-3276 loading="lazy" decoding="async"
srcset="http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/brightspotcdn.byu_.edu_-150x150.jpg 150w, http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/brightspotcdn.byu_.edu_.jpg 200w"
sizes="(max-width: 135px) 100vw, 135px" width="135" height="135"}The
paper stands out by making this process very systematic and
reproducible, and we ask Stephen about why [classification
algorithms](https://dzone.com/articles/introduction-to-classification-algorithms)
are so important, why he uses reproducible and open research in his
work, and how the CODECHECK review process was from an authors
perspective. Stephen is an [Assistant Professor in the College of Life
Science](http://piccolo.byu.edu/) at Brigham Young University, working
in an open and reproducible way to integrate knowledge and techniques
across biology, computer science, medicine, and statistics.\*\*\
\*\*

**Why are there so many classification algorithms and why do they need
to be benchmarked?**

Computer scientists and statisticians have been creating and refining
classification algorithms [for
decades](https://doi.org/10.1111/j.1469-1809.1936.tb02137.x). Research
papers have been published on hundreds of these algorithms, and
implementations are available in the public domain for many of them. In
many research areas (biological or otherwise), classification algorithms
can help to identify patterns that distinguish two or more groups and
then be used to predict the group to which new observations belong.
However, it is difficult for scientists to know which classification
algorithm is best for a particular research application. In addition,
most algorithms support
[hyperparameters](https://en.wikipedia.org/wiki/Hyperparameter_(machine_learning)),
which allow the scientist to modify the algorithm\'s behavior, but it is
inefficient and bias-prone to tune these algorithms in an *ad
hoc* manner. Benchmark analyses address this problem by making the
algorithm/parameter comparisons systematic in nature. We can apply tens
or even hundreds of algorithm variations to benchmark datasets and
identify which tend to perform best.

**Why did you build ShinyLearner and what technical approaches did you
take to design it?**

We created [ShinyLearner](https://github.com/srp33/ShinyLearner) to make
it easier to perform benchmark comparisons. Several open-source software
libraries are available for performing classification analyses. But most
require the scientist to write computer code to perform analysis.
Generally, these libraries perform parameter tuning, but they provide
little insight into this process. We sought to provide a tool that
requires no coding to perform the analysis and that generates simply
formatted output files so users can more easily gain insight into
benchmark results. It encapsulates 4 open-source, machine-learning
libraries into a \"software container\" and provides a consistent
interface for working with any of them. The use of software containers
makes the installation process easier because it already includes all
the software you need to execute the analysis. All you need is to
install the [Docker](https://www.docker.com/) software (or a related
tool for working with containers) and download the container image. The
user then executes the software at the command line (terminal). To ease
this process, we created a web application that helps the user construct
the commands they need to execute. One more tidbit: within the software
container, we use a combination of 4 programming languages to piece
everything together: Python, R, Java, and bash scripting. Most projects
would limit themselves to 1-2 programming language, but we needed to
write code in these various languages to interface with the different
machine-learning libraries but also because in some cases, it seemed
more efficient to use one language over another to implement the desired
functionality. By using software containers, the end user doesn\'t need
to worry about this complexity or install (specific versions of )
runtimes for all of these languages.

**This was an open source tool that used many techniques for
computational reproducibility, so what drove you to follow open science
practices?**

Nowadays, most scientists will not take seriously a journal article
describing a piece of research software unless it is open source. They
want to be able to see the code and verify its design and functionality
directly. This is a great thing! I am still baffled when I see a
research article that describes software or an algorithm but doesn\'t
make the code available.

In [this paper](https://doi.org/10.1093/gigascience/giaa026), we also
followed open-science practices by sharing our analysis scripts (in
addition to the actual [ShinyLearner
software](https://github.com/srp33/ShinyLearner)). We wanted others to
be able to see the exact code we used to generate the figures in our
paper. [We used CodeOcean](https://doi.org/10.24433/CO.5449763.v1), a
cloud-computing service that provides a free tier for open science
projects \[see [previous
blog](http://gigasciencejournal.com/blog/data-intensive-software-publishing-sailing-the-code-ocean-qa-with-ruibang-luo/)
on the platform\]. We did this for the sake of transparency and because
we knew that when we submitted the paper to a journal, at least one
reviewer would want us to tweak our analysis. Because we had already
packaged it up nicely and made it available for others, it was also easy
for us to remember each step of the analysis and tweak what we needed to
tweak. Finally, we used open-science practices because that\'s what we
like to see in others.

**How was the process of undergoing a CODECHECK review as an author?**

For me, the [CODECHECK](http://codecheck.org.uk/) process was easy. The
reviewer set everything up. I just needed to verify that his review was
reasonable and had been configured properly \[see the [certificate
here](http://doi.org/10.5281/zenodo.3674056)\].

**From an author\'s perspective does the CODECHECK certificate seem a
good incentive and reward for making your work easy to use? Do you think
it may drive authors to improve the way they release and write up
software papers?**

![CODECHECK
certificate](http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/Screenshot-2020-04-06-at-4.27.42-PM-300x281.png){.alignright
.wp-image-3272 loading="lazy" decoding="async"
srcset="http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/Screenshot-2020-04-06-at-4.27.42-PM-300x281.png 300w, http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/Screenshot-2020-04-06-at-4.27.42-PM-768x721.png 768w, http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/Screenshot-2020-04-06-at-4.27.42-PM-1024x961.png 1024w, http://gigasciencejournal.com/blog/wp-content/uploads/2020/04/Screenshot-2020-04-06-at-4.27.42-PM.png 1068w"
sizes="(max-width: 222px) 100vw, 222px" width="222" height="208"}It\'s
nice to have that stamp of approval and peace of mind knowing that at
least one outside person was able to reproduce our work. I wasn\'t aware
of this certificate before submitting to this journal. So it wasn\'t a
big motivation for me at the time. But going forward, I would be
motivated by the certificate because it would tell me the journal is
serious about following good scientific practices. Few other journals
make the extra effort to really verify your code when it is open source.
I would love to see this become standard practice. It\'s more work for
everyone in the short term, but in the long run it\'s very beneficial.
It won\'t always prevent research misconduct or ensure that a scientific
analysis is impactful, but it\'s a simple way to encourage open science.

If authors knew that they would need to pass a
[CODECHECK](http://codecheck.org.uk/) from the beginning, it would
motivate them to follow practices early in the process that would
support open, reproducible science. This would benefit themselves but
also the broader community. I touch on this in more detail in an earlier
*Gigascience* paper ([Tools and techniques for computational
reproducibility](https://doi.org/10.1186/s13742-016-0135-4)).

*Read more in our [previous post on
CODECHECK](http://gigasciencejournal.com/blog/codecheck-certificate/).
Stephen Eglen presented CODECHECK at The 14th Munin Conference on
Scholarly Publishing 2019 and you can watch a [video recording
here](https://mediasite.uit.no/Mediasite/Play/8027873496dc465ebc4b9b3ab0338ad01d?playFrom=1772000).*

### **References**

Piccolo SR. et al.,  ShinyLearner: A containerized benchmarking tool for
machine-learning classification of tabular data, *GigaScience*, Volume
9, Issue 4, April 2020,
doi:[10.1093/gigascience/giaa026](https://doi.org/10.1093/gigascience/giaa026)

:::: {.col-md-8}
::: {.blog-post}
Eglen SJ. CODECHECK Certificate 2020-001. Zenodo. 2020[
<http://doi.org/10.5281/zenodo.3674056>]{.ng-binding}
:::
::::

:::: {.col-md-3 .col-md-offset-1}
::: {.panel-body}
Piccolo SR, Frampton MB. Tools and techniques for computational
reproducibility. *Gigascience*. 2016;5(1):30. Published 2016 Jul 11.
doi:[10.1186/s13742-016-0135-4](https://doi.org/10.1186/s13742-016-0135-4)
:::
::::

The post [Reproducible Classification. Q&A on ShinyLearner & the
CODECHECK certificate, pt.
2](http://gigasciencejournal.com/blog/shinylearner-codecheck/){rel="nofollow"}
appeared first on
[GigaBlog](http://gigasciencejournal.com/blog){rel="nofollow"}.