---
author:
- contributor_roles: []
  family: Edmunds
  given: Scott
  url: https://orcid.org/0000-0001-6444-1436
blog:
  authors: null
  community_id: 52db0518-e228-4260-8c54-c4e323b2569d
  created: 1675555200
  current_feed_url: null
  description: Data driven blogging from the GigaScience editors
  doi: https://doi.org/10.59350/gigablog
  favicon: https://rogue-scholar.org/api/communities/52db0518-e228-4260-8c54-c4e323b2569d/logo
  feed_format: application/atom+xml
  feed_url: http://gigasciencejournal.com/blog/feed/atom/
  filter: null
  generator: Other
  home_page_url: https://gigasciencejournal.com/blog
  issn: null
  language: eng
  license: https://creativecommons.org/licenses/by/4.0/legalcode
  prefix: '10.59350'
  relative_url: null
  secure: false
  slug: gigablog
  status: archived
  subfield: '1311'
  title: GigaBlog
  updated: null
  use_api: null
container: GigaBlog
date: '2016-07-19T00:00:00+00:00'
date_updated: '2025-12-06T10:33:31+00:00'
guid: http://blogs.biomedcentral.com/gigablog/?p=1809
identifier: https://doi.org/10.59350/51w44-grp73
image: http://gigasciencejournal.com/blog/wp-content/uploads/2016/07/url.png
images:
- alt: 'Strategy for knowledge generation enabled by PhenoMeNal e-infrastructure.
    Source: PhenoMeNal consortium.'
  height: '565'
  sizes: '(max-width: 800px) 100vw, 800px'
  src: http://gigasciencejournal.com/blog/wp-content/uploads/2016/07/url.png
  srcset: http://gigasciencejournal.com/blog/wp-content/uploads/2016/07/url.png, http://gigasciencejournal.com/blog/wp-content/uploads/2016/07/url-300x212.png,
    http://gigasciencejournal.com/blog/wp-content/uploads/2016/07/url-768x542.png
  width: '800'
issn: null
keywords:
- Medicine
- Technology
- Big Data
- Guest Post
- Medical Data
lang: en
license: https://creativecommons.org/licenses/by/4.0/legalcode
rid: mf6ph-64c88
rights: https://creativecommons.org/licenses/by/4.0/legalcode
summary: <strong> Building an open, community-supported, e-infrastructure for medical
  metabolomics data. </strong> The European Union's Horizon 2020 research and innovation
  programme is funding the PhenoMeNal (Phenome and Metabolome aNalysis) project that
  aims to support data processing and analysis pipelines for molecular phenotype data
  generated from metabolomics applications.
title: 'Guest posting: Building a PhenoMeNal metabolomics e-infrastructure'
url: https://wayback.archive-it.org/22098/2025-05-01T17:13:42Z/http://gigasciencejournal.com/blog/guest-posting-building-phenomenal-metabolomics-e-infrastructure
version: v1
---

**Building an open, community-supported, e-infrastructure for medical
metabolomics data.**\
The European Union\'s Horizon 2020 research and innovation programme is
funding the [PhenoMeNal](http://phenomenal-h2020.eu/) (Phenome and
Metabolome aNalysis) project that aims to support data processing and
analysis pipelines for molecular phenotype data generated from
metabolomics applications. We aim to build an open, community-led and
community-supported e-infrastructure by leveraging existing cloud
infrastructures, tooling, and data repositories, under one umbrella of
services dedicated to the European biomedical community to begin with,
and eventually worldwide. However, the experience and tooling can be
applied to non clinical settings.

The project was formulated to address the complexity and high volumes of
data being generated, collected, and analysed that are quickly going
beyond current data management and computational capabilities. For
example, it is estimated that a single National Phenome Centre managing
only around 100,000 human samples per year might generate a velocity of
data amounting to more than 2PB annually. And this is a conservative
estimate.

When thinking of the near to middle-term future where phenotyping
efforts on a population-wide scale are on the horizon, such
e-infrastructures will need to deal with Exabyte-scales of biological
phenotyping data. It is essential to build such data infrastructures, in
particular for metabolomics phenotyping, as there is an urgent need to
improve the understanding of the causes and mechanisms underlying
health, healthy ageing and diseases beyond what can be done solely with
genomic approaches.

::: {#attachment_1811 .wp-caption .aligncenter style="width: 810px"}
![Strategy for knowledge generation enabled by PhenoMeNal
e-infrastructure. Source: PhenoMeNal
consortium.](http://gigasciencejournal.com/blog/wp-content/uploads/2016/07/url.png){.size-full
.wp-image-1811 data-fetchpriority="high" decoding="async"
aria-describedby="caption-attachment-1811"
srcset="http://gigasciencejournal.com/blog/wp-content/uploads/2016/07/url.png 800w, http://gigasciencejournal.com/blog/wp-content/uploads/2016/07/url-300x212.png 300w, http://gigasciencejournal.com/blog/wp-content/uploads/2016/07/url-768x542.png 768w"
sizes="(max-width: 800px) 100vw, 800px" width="800" height="565"}

Strategy for knowledge generation enabled by PhenoMeNal
e-infrastructure. Source: PhenoMeNal consortium.
:::

There are also additional challenges beyond simply addressing the
volumes, velocity and variety of data. Ethics and privacy are of central
concerns in the project since we aim to handle clinical metabolomics
data, and we are engaged with experts on ethics, legal, and social
implications of biological data. The project also aims to build
sustainable infrastructure that will live-on beyond the three years of
EU funding awarded to PhenoMeNal. This involves ensuring that all of the
infrastructure we develop is open source, well documented, and community
supported, and that we engage with existing research e-infrastructures
across Europe and worldwide.

\*\*Building part public-part private clouds\
\*\*So how are we building the PhenoMeNal e-infrastructure to address
these challenges?

In our first year, we have completed plenty of groundwork towards
setting up exemplar pipelines on public and private cloud computing
infrastructure. We are working on setting up the computing
infrastructure across European boundaries, with part hosting on the
European Bioinformatics Institute\'s EMBASSY Cloud in the UK, on an
installation at UPPMAX in Uppsala in Sweden (infrastructure previously
[covered by
*GigaScience*](https://gigascience.biomedcentral.com/articles/10.1186/2047-217X-2-9)),
and within CRS4\'s datacenter in Sardinia, Italy. Alongside these, we
are also working with Google on creating a hybrid public-private cloud
infrastructure that we can then demonstrate the migration of computing
tools to data that needs to be ring-fenced within institutions (e.g.
within a clinic\'s own infrastructure where data cannot be moved) while
at the same time still leveraging the power of the public cloud.

We have also been working on packaging metabolomics software to become
\'cloud-ready\'. We are utilizing the [Docker
system](http://phenomenal-h2020.eu/home/2016/04/15/phenomenal-public-docker-registry-is-live/)
for containerizing software, where we have already packaged tools such
as
[BATMAN](http://bioinformatics.oxfordjournals.org/content/early/2012/05/25/bioinformatics.bts308)
(Bayesian AuTomated Metabolite Analyser for NMR spectra),
[OpenMS](http://www.openms.de/),
[XCMS](https://metlin.scripps.edu/xcms/), as well as data management
tooling such as the [ISA metadata framework](http://www.isa-tools.org)
API.

All of this work is in collaboration with clinical and research sites,
to support some example use cases. We work closely with clinical
metabolomics centers such as [CARAMBA](http://www.medsci.uu.se/caramba/)
(Clinical Analysis and Research Applying Mass spectrometry and
Bioinformatics at Akademiska) at Uppsala University (Sweden), the
[MRC-NIHR National Phenome
Centre](http://www.imperial.ac.uk/phenome-centre) (UK), the [Phenome
Centre Birmingham](http://www.birmingham.ac.uk/phenome-centre) (UK),
[Hospital Clínic of Barcelona](http://www.hospitalclinic.org/en)
(Spain), and the [Netherlands Metabolomics
Centre](http://www.metabolomicscentre.nl/) (Netherlands).

\*\*Survey on Metabolomics Data: Call for participation\
\*\*To gain a better understanding of the requirements for data
infrastructures for metabolomics research, PhenoMeNal has commissioned a
survey to gain insight into current practices, future needs, and
opinions, on how metabolomics research data is managed, published, and
disseminated. The survey welcomes contributions from the community of
researchers, including clinicians, and also from those involved with
data management, engineering and software development involved with
research and application of metabolomics.

If you work in any role relating to metabolomics, be it in clinical
application, basic research, or in technical support, please help us
build PhenoMeNal as an open, community-led, metabolomics infrastructure
by letting us know your thoughts.

You can contribute to the survey (closes end August 2016) here:
<https://www.surveymonkey.co.uk/r/metabolomics-data>

**Acknowledgements**\
The PhenoMeNal project is led by a consortium of 14 partners at: the
European Molecular Biology Laboratory-European Bioinformatics Institute
(UK), Imperial College of Science, Technology and Medicine (UK),
Leibniz-Institute of Plant Biochemistry (Germany), Universitat de
Barcelona (Spain), The University of Birmingham (UK), Consorzio
Interuniversitario Risonanze Magnetiche di Metallo Proteine (Italy),
Universiteit Leiden (Netherlands), The University of Oxford (UK), Swiss
Institute of Bioinformatics (Switzerland), Uppsala Universitet (Sweden),
Biobanking and BioMolecular Resources Research Infrastructure-ERIC
(Austria), Commissariat à l\'énergie atomique et aux énergies
alternatives (France), Institut national de la recherche agronomique
(France), and Centro Di Ricerca, Sviluppo E Studi Superiori in Sardegna
(Italy).

The post [Guest posting: Building a PhenoMeNal metabolomics
e-infrastructure](http://gigasciencejournal.com/blog/guest-posting-building-phenomenal-metabolomics-e-infrastructure/){rel="nofollow"}
appeared first on
[GigaBlog](http://gigasciencejournal.com/blog){rel="nofollow"}.