{"found":53415,"hits":[{"document":{"authors":[{"affiliation":[{"id":"https://ror.org/02mb95055","name":"Birkbeck, University of London"}],"contributor_roles":[],"family":"Eve","given":"Martin Paul","url":"https://orcid.org/0000-0002-5589-8511"}],"blog":{"authors":[{"name":"Martin Paul Eve","url":"https://orcid.org/0000-0002-5589-8511"}],"community_id":"9224b0d7-fc03-497c-9c6f-85c9fd1e72da","created":1690329600,"current_feed_url":null,"description":null,"doi":"https://doi.org/10.59348/eve","favicon":"https://rogue-scholar.org/api/communities/9224b0d7-fc03-497c-9c6f-85c9fd1e72da/logo","feed_format":"application/atom+xml","feed_url":"https://eve.gd/feed_all.xml","filter":null,"generator":"Jekyll","home_page_url":"https://eve.gd","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59348","relative_url":null,"secure":true,"slug":"eve","status":"active","subfield":"1208","title":"Martin Paul Eve","updated":1785660546,"use_api":true},"blog_name":"Martin Paul Eve","blog_slug":"eve","content_html":"<p>This morning, I made a last-minute incorporation of Kornbrodt, Joseph, and David Zappone, 'Special Feature on <em>To the Journey: Looking Back at</em> Star Trek: Voyager', interview with Kate Mulgrew, 455 Films, 2026, Bluray into my <em>Voyager</em> book. I had watched the documentary itself in a digital copy earlier in the year. It's OK. Nothing very new. You do get a VERY enthusiastic Garrett Wang shepherding the program along with a slightly strange foray into a zero-gravity plane dive. I am unsure this has anything to do with <em>Voyager</em>, but it was mildly entertaining. And it's Harry Kim! And he clearly loves it and the <em>Voyager</em> fan community and all that goes with it. It's all quite heartwarming.</p>\n<p>But yesterday, the Bluray arrived and that has special features. And I really enjoyed the special feature interviews on this documentary! There was also some bad editing. Mulgrew's interview has a section that repeats itself and that is very frustrating as a viewer. However, I suppose that this is mostly b-reel footage. Also, the interviewer is not very good at moving her on beyond the \"women still struggle to have it all\" question (not that that isn't important).</p>\n<p>However, there is some great stuff in here!</p>\n<p>Just Mulgrew's work ethic. 3am starts, midnight finishes. Every part of her was invested in The Work. (In fact, so many of the interviewees talk with reverence about The Work. It's all there is, for them. A total dedication to it.) She talks of how she had to brace every part of her body's musculature in a specific way when playing Janeway. Her acting is a total body investment. She speaks of how playing Janeway for seven years meant that she was not there for her children, ever. She made a conscious choice to do this and, to this day, feels guilt. Her children have never seen her play Janeway. They have watched none of <em>Voyager</em> because, she says, they consider it the reason they had a motherless childhood. This seems to have had very deep, lasting family psychological problems. This is, then, an incredible commitment. A life given over to the role, really.</p>\n<p>She also states that she does not believe that introducing Janeway would work now. She believes that \"as a result of the digital age\" audiences \"are attracted to darkness\" and, for Mulgrew, \"Janeway was not dark. Janeway was light\". <em>Voyager</em> was, for her, an optimistic Trek, not a dark, cynical Trek. This is such a crucial point for my book. It is also very heartening to hear Mulgrew talk of \"the Janeway effect\" that has encouraged women in STEM and space. She talks of how female fans come up to her at conventions and spill their souls about how they went into science because of her. My heart did sink a little, as a man. Because I would love (but am too shy) to meet her at a convention and tell her how inspirational <em>I</em>, even though I am a man, found her. I never questioned her authority or thought it out of place that she was in command. It was clear. Obvious. Just how it should be, I thought (although I was in my teenage years when the later series of <em>Voyager</em> was airing for the first time.) So, Kate, if you ever read this, please know: men took inspiration from you, too. That's why <a href=\"https://janeway.systems/\">I named our software platform after you</a>.</p>\n<p>There's also a lovely, if not terribly informative, interview with Jeri Taylor (RIP). I suppose there IS lots of interest here; especially about the open story submission process that they ran. They simply didn't have enough ideas and so invited anyone to pitch to them. She does say that this means she sat through a lot of VERY bad pitches! But the main thing I took away from her interview was a sense of sadness. She talked of the intense workload and non-stop need for fresh material. But, as closing remarks, she said she would never, ever write for fun now. The fun has been taken from her by decades of writing to a schedule. This was somewhat terrible, even though she said it half-jokingly.</p>\n<p>There are interviews with Michael Piller's son and wife as well on the disc that give an interesting background to his life. Although I just thought more about how Shawn Piller's life had been extraordinary. Being introduced to Gene Roddenberry! Writing stories/scripts for <em>Star Trek: The Next Generation</em>! All opportunities opened up through his father (although he had to prove himself, even though he had this leg-up.) An amazing life. Few have that experience.</p>\n<p>There's also a fun section on directing (with a lot of Armin Shimerman featured, who is GREAT and hilarious!) but I was disappointed not to hear, in this section, from Roxann Dawson, who is doubtless the most successful crossover actor/director from the <em>Voyager</em> cast. But hey. I think the documentary had varying levels of commitment from the cast and was not high on their priority list, sadly. (Although they had lots of Mulgrew's time, which was generous of her.)</p>\n<p>Anyway, all in all, this was an informative and enjoyable set of interviews. It's probably only of great interest to dedicated fans. It can otherwise feel niche and repetitive. But for the hardcore fan, this is a worthwhile set of bonuses. I have still to watch the sections with the Paramount Executives. They're not generally liked! Mulgrew points out: the bottom line is always the money! But I suppose their executive position is worth knowing about. If there's anything significant, I will probably update this post.</p>\n<p><a href=\"https://eve.gd/2026/08/02/the-extended-interviews-on-ito-the-journey-looking-back-ati-star-trek-voyager/\">The extended interviews on <i>To the Journey: Looking Back at</i> Star Trek: Voyager</a> was originally published by Martin Paul Eve at <a href=\"https://eve.gd\">eve.gd: Martin Paul Eve</a> on August 02, 2026.</p>","doi":"https://doi.org/10.59348/qygeq-6n059","guid":"https://doi.org/10.59348/qygeq-6n059","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1785628800,"rid":"aszdc-x4j13","summary":"This morning, I made a last-minute incorporation of Kornbrodt, Joseph, and David Zappone, 'Special Feature on <em> To the Journey: Looking Back at </em> Star Trek: Voyager', interview with Kate Mulgrew, 455 Films, 2026, Bluray into my <em> Voyager </em> book. I had watched the documentary itself in a digital copy earlier in the year. It's OK. Nothing very new.","title":"The extended interviews on <i>To the Journey: Looking Back at</i> Star Trek: Voyager","updated_at":1785873800,"url":"https://eve.gd/2026/08/02/the-extended-interviews-on-ito-the-journey-looking-back-ati-star-trek-voyager/","version":"v1"}},{"document":{"authors":[{"affiliation":[{"id":"https://ror.org/048a87296","name":"Uppsala University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p><a href=\"http://www.nature.com/nchem/\">Nature Chemistry</a> just released the first issue with a few free papers,\nlike <em>Asymmetric total syntheses of (+)- and (-)-versicolamide B and biosynthetic implications</em> by Miller et al.\n(DOI:<a href=\"https://doi.org/10.1038/nchem.110\">10.1038/nchem.110</a>).</p>\n<p>Now, we've seen the Royal Society of Chemistry's <a href=\"http://chem-bla-ics.blogspot.com/search?q=project+prospect\">Project Prospect</a> <!-- keep link -->\n(see <a href=\"https://chem-bla-ics.linkedchemistry.info/2007/02/01/rsc-first-publisher-to-go-semantic.html\">RSC: the first publisher to go semantic! <i class=\"fa-solid fa-recycle fa-xs\"></i></a>)\nand ChemSpiders recent <a href=\"http://www.chemmantis.com/\">ChemMantis</a> system which enriches\nthe papers with machine readable representations of the molecules discussed in those\npapers. The new Nature publication has been in the works for a while, and they\n<a href=\"http://blogs.nature.com/thescepticalchymist/2008/05/jj_day_98_service_with_a_simpl.html\">asked</a>\nthe community before what a Nature Chemistry paper should like like, and I replied in\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2008/05/08/re-what-should-nature-chemistry-paper.html\">Re: What should a Nature Chemistry paper look like? <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.</p>\n<h2 id=\"the-verdict\">The verdict</h2>\n<p>So, have the been listening? Is the HTML they produce semantic? Is it data rich? Or is it\njust another hamburger? Well, I am very happy to see some of the suggestions I made picked\nup (though I do not fool myself in believing I am the only one that suggested those\nfeatures). A tour of good things, and points for improvement.</p>\n<p>The first impression is not shocking; it looks like any other interface, with molecules drawn as images in the paper:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem3.png\"/></p>\n<p>All structures that are numbered and linked (as in <em>C6-epi-stephacidin A (Compound <strong>13</strong>)</em>\nhave a hover-over function to popup a drawing of the structure:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem4.png\"/></p>\n<p>The popup image is a nice gimmick, but not really sematically useful. The link, however,\nis! It points to a separate supplementary page with further information which include\na image of the 2D structure and, following a link, the 3D structure in <a href=\"http://www.jmol.org/\">Jmol</a>.\nMoreover, it comes with the machine readable representations:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem5.png\"/></p>\n<p>This is indeed interesting, and a big step forward, though please do note my comments later.\nFor convenience, all molecules with such supplementary information is available from the\nspecial Chemical Compounds section of the paper:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem2.png\"/></p>\n<p>Excellent! This really is a step forward towards a data-rich paper! Indeed, I will shortly\nwrite up a <a href=\"http://www.bioclipse.net/\">Bioclipse</a> plugin for Nature Chemistry, which\nwill download all molecular structures based on the DOI! Anyway, more on that later\u2026\nFor this article, that table looks like:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem1.png\"/></p>\n<p>By now, you likely also noted the links to <a href=\"http://pubchem.ncbi.nlm.nih.gov/\">PubChem</a>, and\nindeed, upon publication of a paper, all structures are deposited in the public domain:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem6.png\"/></p>\n<p>At last but not least, each molecule is available in the <a href=\"http://en.wikipedia.org/wiki/Chemical_Markup_Language\">Chemical Markup Language</a>\n(with 2D coordinates)! And you know I am a very happy CML user for a long time (see e.g.\nPeter's recent blog <a href=\"https://doi.org/10.59350/jbq7c-szw40\">Egon Willighagen and CML <i class=\"fa-solid fa-recycle fa-xs\"></i></a>).\nBTW, one comment on the CML: the namespace used is the outdated namespace, <strong>not</strong>\nthe current one (see <a href=\"http://cmlexplained.blogspot.com/2007/06/there-can-be-only-one-namespace.html\">There can be only one (namespace)</a>).\n(But the <a href=\"http://cdk.sf.net/\">CDK</a> and Bioclipse will read it anyway.)</p>\n<h2 id=\"details-matter\">Details matter</h2>\n<p>So, while the first impression was not shocking, it was a bit deceptive. <em>Nature Chemistry</em>\nreally changes publishing of chemistry. But I have bad news too. They need to improve the\nHTML they produce.</p>\n<p>But before pointing out some missed chances, let me reply <em>inter alia</em> to Peter's recent\nwork on the Open Source plugin for including semantic chemistry in MS-Word documents\n(see <a href=\"https://doi.org/10.59350/wn2pv-gef13\">How can we publish semantic chemical documents? <i class=\"fa-solid fa-recycle fa-xs\"></i></a>):\nNature Chemistry seems to have done a great job with existing tools. Nevertheless, I fully\nback up Peters comment that while the plugin is useless without Word, the results produced\nwith the plugin are extremely Open Standard, and enormously reusable! Indeed, while the\nWord file format is only formally an true Open Standard, the file format is plain XML, and\nextracting content bearing the CML namespace is trivial.</p>\n<p>Which reminds me, if someone from the Nature Chemistry team is reading this, please point\nme to a blog what tools actually <em>are</em> involved in publishing a Nature Chemistry paper!\nI think we all like to know.</p>\n<p>Now, the <a href=\"http://en.wikipedia.org/wiki/HTML\">HTML</a> has room for improvement. First of all,\na look at the metadata defined for the web page of the article shows a <em>description</em>\nand <em>keywords</em> about the journal, not the article, and the same goes for the web pages for\nthe molecules:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem7.png\"/></p>\n<p>Additionally, the compound details web page has no special markup for the machine readable\ninformation:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem8.png\"/></p>\n<p>Or, if it does, it's still mixed with markup for visual pleasing output:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem9.png\"/></p>\n<p>Still, the HTML is clean enough to have some regular expressions extract a good deal of\ninformation, and there is also still the PubChem deposition.</p>\n<h2 id=\"beyond-connection-tables\">Beyond connection tables</h2>\n<p>Like many other chemistry journals, Nature Chemistry does not consider properties of\nthe molecule interesting, and NMR spectra are hidden in the Supplementary Information.\nThis paper in particular, disregards a lot of machine readable facts by putting all\nexperimental section bits in a PDF document. So, the next challenge for Nature Chemistry\nwill be to get the authors of papers contribute the original spectra (JCAMP-DX, CMLSpect,\netc) in the supplementary information section. Better, have the raw data or even the NMR\npeak-atom annotations deposited in public repositories such (see \n<a href=\"https://chem-bla-ics.linkedchemistry.info/2009/03/04/open-nmr-data-raw-curves-and-annotated.html\">Open NMR data: raw curves and annotated peak lists <i class=\"fa-solid fa-recycle fa-xs\"></i></a>).</p>\n<p>All in all, I am rather positive about the first Nature Chemistry issue, and like to\nthank the editors and paper authors for there efforts on improving publishing chemistry!</p>","doi":"https://doi.org/10.59350/40377-hz881","guid":"https://doi.org/10.59350/40377-hz881","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1237420800,"reference":[{"id":"https://doi.org/10.1038/nchem.110","unstructured":"Unknown title"},{"id":"https://doi.org/10.59350/jbq7c-szw40","unstructured":"Unknown title"},{"id":"https://doi.org/10.59350/wn2pv-gef13","unstructured":"Unknown title"}],"rid":"wasde-08n67","summary":"Nature Chemistry just released the first issue with a few free papers, like Asymmetric total syntheses of (+)- and (-)-versicolamide B and biosynthetic implications by Miller et al. (DOI:10.1038/nchem.110).","tags":["Inchi","Chemistry","Jmol"],"title":"Nature Chemistry improves publishing chemistry: a detailed analysis","updated_at":1785873774,"url":"https://chem-bla-ics.linkedchemistry.info/2009/03/19/nature-chemistry-improves-publishing.html","version":"v1"}},{"document":{"authors":[{"affiliation":[{"id":"https://ror.org/048a87296","name":"Uppsala University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>I have blogged about two Molecular Chemometrics principles so far:</p>\n<ul>\n<li><a href=\"https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html\">McPrinciple #1: access to data</a></li>\n<li><a href=\"https://chem-bla-ics.linkedchemistry.info/2010/08/12/molecular-chemometrics-principles-2-be.html\">McPrinciple #2: be clear in what you mean</a></li>\n</ul>\n<p>Peter's post <a href=\"https://doi.org/10.59350/hphjc-qgr72\">#solo10: Green Chain Reaction; where to store the data? DSR? IR? BioTorrent, OKF or ??? <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\ngives me enough basis to write up a third principle:</p>\n<p><strong>Molecular Chemometrics Principles #3</strong>: We make scientific progress if we build on past achievements.</p>\n<p>Sounds logical, right? Practically, the way we share our cheminformatics knowledge makes this standing on shoulders pretty difficult.\nBut there is one particular aspect I would like to ask your attention for: you can contribute by making clear what shoulders\nyou would like to stand on. That is, where do you prefer to put your effort, and what message would you like to give to your user community.</p>\n<p>In the aforelinked post, Peter asks where he should upload his data, and he suggest <a href=\"http://www.biotorrents.net/\">BioTorrent</a> (see my review\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2010/04/18/bittorrents-for-science.html\">BitTorrents for Science <i class=\"fa-solid fa-recycle fa-xs\"></i></a>), DSpace, and <a href=\"http://www.ckan.net/\">CKAN</a>.\nNow, his <a href=\"http://www.google.se/search?sourceid=chrome&amp;client=ubuntu&amp;channel=cs&amp;ie=UTF-8&amp;q=%22Green+Chain+Reaction%22\">Green Chain Reaction</a>\nis picked up (see <a href=\"http://researchremix.wordpress.com/2010/08/11/green-chain-reaction-project-putting-my-minutes-where-my-mouth-is/\">these</a>\n<a href=\"http://scienceonlinelondon.wikidot.com/topics:green-chain-reaction\">few</a> <a href=\"https://doi.org/10.59350/h2jq5-3np88\">blog <i class=\"fa-solid fa-recycle fa-xs\"></i></a> posts),\nand the resulting data should be distributed as much as possible. The exact location does not really matter\u2026</p>\n<p>But\u2026</p>\n<p>By picking where you upload, you make a statement to your community: \"<em>Look guys, we are distributing our data via Foo, because we believe those guys are doing good work! Perhaps you can support them too.</em>\".</p>\n<p>This principle does not only apply to data, it applies to things too. For example, when\n<a href=\"http://www.chemspider.com/blog/ichemlabs-and-rsc-chemspider-announce-partnership.html\">iChemLabs and RSC ChemSpider Announce Partnership</a>\nthey do not just improve the user experience of ChemSpider (which I certainly won't object against), but they also imply\n\"<em>Look dudes, your product is just not good enough and we do not want to help you improve it either</em>\".\nOf course, ChemSpider has every right, and for them to succeed it is crucial to make decisions like this. Fortunately,\n<a href=\"http://web.chemdoodle.com/installation.php\">ChemDoodle is GPL</a>.</p>\n<p>Every project with a user base has the opportunity to support shoulders, if they only visibly stand on them. By merely discussion the\n<em>Green Chain Reaction</em>, I show to support this social web experiment. You can too. Use these powers wisely. May the McPrinciples be with you.</p>","doi":"https://doi.org/10.59350/832vn-qwh10","guid":"https://doi.org/10.59350/832vn-qwh10","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1281744000,"reference":[{"id":"https://doi.org/10.59350/hphjc-qgr72","unstructured":"Unknown title"},{"id":"https://doi.org/10.59350/h2jq5-3np88","unstructured":"Unknown title"}],"rid":"55qe3-cbt36","summary":"I have blogged about two Molecular Chemometrics principles so far:","tags":["Mcprinciples","Solo10","Chemdoodle","Chemspider","Javascript"],"title":"The Molecular Chemometrics Principles #3: stand on shoulders","updated_at":1785873773,"url":"https://chem-bla-ics.linkedchemistry.info/2010/08/14/molecular-chemometrics-principles-3.html","version":"v1"}},{"document":{"authors":[{"affiliation":[{"id":"https://ror.org/048a87296","name":"Uppsala University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>About 6 months ago I <a href=\"http://chem-bla-ics.blogspot.com/2009/03/nmrshiftdb-enters-rdfopenmoleculesnet.html\">reported</a> about my efforts to RDF-ize the data from the\n<a href=\"http://www.nmrshiftdb.org/\">NMRShiftDB</a>. Since then, time was consumed by many other things, but now that <a href=\"http://www.bioclipse.net/\">Bioclipse</a> can query\n<a href=\"http://en.wikipedia.org/wiki/SPARQL\">SPARQL</a> end points, that I want to contribute the triple set (it is <a href=\"http://www.gnu.org/copyleft/fdl.html\">GNU FDL</a>-licensed)\nto <a href=\"http://www.bio2rdf.org/\">Bio2RDF</a>, that a student started working in my group (now larger than just me :) on reasoning on life sciences data, and that I\nrecently contributed my <a href=\"http://egonw.posterous.com/nmrshiftdb-1006-contributions-and-counting\">1000th NMR spectrum</a> to the database, I thought it was time to\nfinally reinstall <a href=\"http://www.openlinksw.com/wiki/main/Main/VOSDownload\">Virtuoso</a>.</p>\n<p>There are precompiled binaries for <a href=\"https://launchpad.net/~wdaniels/+archive/ppa\">Ubuntu</a> and <a href=\"http://bugs.debian.org/cgi-bin/bugreport.cgi?bug=508048\">Debian</a>,\nbut Michel encouraged me to use version 6 when <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/06/26/michel-dumontier-at-uppsala-university.html\">he visited us <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.\nAnd so I compiled and install <a href=\"https://sourceforge.net/projects/virtuoso/files/virtuoso-devel/6.0.0-TP1/\">6.0.0.TP1</a> on the public server, while I do have the\nbinary debs for 5.0.12 on my laptop. With some basic Apache magic, I hooked up the SPARQL end point of the server to the web:</p>\n<div class=\"language-xml highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code><span class=\"nt\">&lt;Proxy</span> <span class=\"err\">/nmrshiftdb/sparql</span><span class=\"nt\">&gt;</span>\n  RewriteEngine On\n  Allow from all\n  ProxyPass        http://localhost:8890/sparql\n  ProxyPassReverse http://localhost:8890/sparql\n<span class=\"nt\">&lt;/Proxy&gt;</span>\n</code></pre></div></div>\n<p>Nice thing about this is, that I can set up multiple servers, allowing me to keep incompatibly licensed data sets apart (see\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2009/05/18/open-data-license-rights-aggregation.html\">Open Data: license, rights, aggregation, clean interfaces? <i class=\"fa-solid fa-recycle fa-xs\"></i></a>), which is\nthe same approach Bio2RDF is taking.</p>\n<p>The <a href=\"http://pele.farmbio.uu.se/nmrshiftdb/sparql\">end point</a> now offers about <a href=\"http://pele.farmbio.uu.se/nmrshiftdb/sparql?default-graph-uri=&amp;query=SELECT+count%28*%29+WHERE+{\\%0D%0A++%3Fs+%3Fp+%3Fo+.%0D%0A}&amp;format=text%2Fhtml&amp;debug=on\">278887</a>\ntriples, but this will soon rise as I make more content from the database available in the original SQL database. The data is from the\n<a href=\"https://sourceforge.net/projects/nmrshiftdb/files/nmrshiftdb/1.3.3/\">1.3.3 release</a> by <a href=\"http://www.steinbeck-molecular.de/steinblog/\">Chris</a>'\nteam, and does not include my 1000th spectrum.</p>\n<p>Getting the data into the database was not trivial either. The documentation suggests WebDAV, and that indeed worked for me once, after\nusing the <a href=\"http://www.snee.com/bobdc.blog/2009/02/getting-started-using-virtuoso.html\">curl approach suggested here</a>. But upon a second upload, it\ndid again not enter the store. The ultimate solution was to use the iSQL interface, with the following SQL</p>\n<div class=\"language-plaintext highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code>DB.DBA.RDF_LOAD_RDFXML_MT(\n  file_to_string_output('/tmp/nmrshiftdb.rdf'), '',\n  'http://pele.farmbio.uu.se/nmrshiftdb'\n);\n</code></pre></div></div>\n<p>Scientifically, this progress is not overly interesting, although it makes it very clear that you really should not have to be happy with proprietary\nand non-semantic formats for anything. But, to me, this is mostly a technological success of great importance: I can now share really large sets of\nRDF data.</p>\n<p>Querying this data is a simple with SPARQL, and the results are available in various formats, such as JSON, which makes it easy to integrate in\nthird-party applications or <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/09/02/google-wave-robot-for-cdk-functionality.html\">Google Wave robots <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\n(did I hear someone say <a href=\"http://nmrshifty.appspot.com/\">NMRShifty</a>?). As I have <a href=\"http://chem-bla-ics.blogspot.com/search?q=sparql\">blogged before</a>,\nSPARQL is an excellent tool to aggregate scientific data prior to data analysis. And I will demo more interesting queries later this month.</p>","doi":"https://doi.org/10.59350/nv925-tje87","guid":"https://doi.org/10.59350/nv925-tje87","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1252022400,"rid":"rgwsf-1q277","summary":"About 6 months ago I reported about my efforts to RDF-ize the data from the NMRShiftDB.","tags":["Rdf","Sparql","Nmrshiftdb","Cheminf"],"title":"NMRShiftDB enters rdf.openmolecules.net #2: SPARQL end point with Virtuoso","updated_at":1785871018,"url":"https://chem-bla-ics.linkedchemistry.info/2009/09/04/nmrshiftdb-enters-rdfopenmoleculesnet-2.html","version":"v1"}},{"document":{"authors":[{"affiliation":[{"id":"https://ror.org/048a87296","name":"Uppsala University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>Yeah, it's my turn. Standing on the shoulders of <a href=\"http://chemjobber.blogspot.com/2010/05/my-favorite-things-about-chemistry.html\">ChemJobber</a>,\n<a href=\"http://www.chemistry-blog.com/2010/05/07/a-few-of-my-favorite-chemistry-things/\">Azmanam</a>, and\n<a href=\"http://www.sciencebase.com/science-blog/my-favourite-chemistry-things.html\">ScienceBase</a>, here's list of things I like about chemistry.\nTo put things into perspective first, a bit, I note that ChemJobber and Azmanam focused on wet-lab chemistry, and David on fancy\nmolecules. Now, I am a theoretical chemist, and was thinking on what to orient the things I like, and on how general to make them.\nThis meme is not easy, you now. But here goes:</p>\n<h3 id=\"1-chemical-graph-theory\">1. chemical graph theory</h3>\n<p>Chemical graph theory is one of the common theoretical models chemists work with to make sense of chemical properties.\nI like it because the graph theory is fairly straightforward, but chemistry adds enough color (literally!) to create a\nnice complexity that kept the cheminformatics field going strong for more than 50 years now :) For example, how to adapt\nthe theory to <a href=\"https://chem-bla-ics.linkedchemistry.info/2006/12/30/modern-chemistry-in-cdk-beyond-two.html\">deal with mutli-atom bonds <i class=\"fa-solid fa-recycle fa-xs\"></i></a> :)</p>\n<h3 id=\"2-rare-nuclei-in-the-nmrshiftdb\">2. rare nuclei in the NMRShiftDB</h3>\n<p>The <a href=\"http://www.nmrshiftdb.org/\">NMRShiftDB</a> is an Open Data repository for annotated NMR spectra. The fun here is to\nadd NMR spectra of <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/09/05/nmrshiftdb-rdf-2-some-statistics.html?q=nmr+nuclei+sparql\">rare nuclei <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.\nDon't you just love a\nmolecule with NMR shifts for all atoms?</p>\n<h3 id=\"3-metabolomics\">3. metabolomics</h3>\n<p>Metabolomics is the research field that studies the small molecules of life. Plant metabolomics is particularly fun.\nTens of thousands of molecules, and a lot of metabolite identification to be done, and much more. Lot's of cool stuff\nto do here, and I am trying to secure funding for it. This is\n<a href=\"http://chem-bla-ics.blogspot.com/search?q=metabolomics\">what I blogged about metabolomics before</a>. What about\nhis nice <a href=\"http://en.wikipedia.org/wiki/Secondary_metabolite\">secondary metabolite</a> (source:\n<a href=\"http://en.wikipedia.org/\">Wikipedia</a>, <a href=\"http://en.wikipedia.org/wiki/File:Discodermolide_Structure.png\">CC0</a>):</p>\n<p><img alt=\"\" src=\"https://upload.wikimedia.org/wikipedia/commons/e/e4/Discodermolide_Structure.png\"/></p>\n<h3 id=\"4-hexavalent-carbon\">4. hexavalent carbon</h3>\n<p>Atom types is another theoretical model for chemistry. <a href=\"https://chem-bla-ics.linkedchemistry.info/2007/07/01/atom-typing-in-cdk.html\">Atom typing <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\nis one of the underlying technologies of <a href=\"http://en.wikipedia.org/wiki/Force_field_%28chemistry%29\">force fields</a>, which are used\nin many, many fields in chemistry. Now, force fields typically take only a subset of atom types. New atom types, consequently, need\nto be added. One such new atom type was the <a href=\"http://www.ch.ic.ac.uk/rzepa/blog/?p=811\">hexavalent carbon</a>. Rare, very rare, but\njust the amount of complexity I like about chemistry.</p>\n<!-- Image lost -->\n<h3 id=\"5-self-organizing-maps\">5. self-organizing maps</h3>\n<p>Kohonen maps, or <a href=\"http://en.wikipedia.org/wiki/Self-organizing_map\">self-organizing maps</a> (SOM), are a machine learning\nmethod that have interesting visualization features. They have numerous applications, and also in chemistry. The group\nwhere I did my PhD developed a supervised SOM, which I used them to classify crystal structures\n(doi:<a href=\"https://doi.org/10.1021/cg060872y\">10.1021/cg060872y</a>). Another of my favorites is the\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2006/04/04/mining-kegg-pathway-database-with-self.html\">reaction classification by Aires-de-Sousa <em>et al.</em> <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\nusing unsupervised SOMs.</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/som_thesis_image.png\"/></p>\n<h3 id=\"6-the-maillard-reaction\">6. the Maillard reaction</h3>\n<p>People who know me personally, know that I like tasting things. That also makes me have to worry about overweight.\nTaste is to a large extend governed by cooking, and the <a href=\"http://en.wikipedia.org/wiki/Maillard_reaction\">Maillard reaction</a>\nplays an important role here. If you like to learn more about the chemistry of cooking, checkout these\n<a href=\"http://chemistandcook.blogspot.com/\">two</a> <a href=\"http://blog.khymos.org/\">blogs</a>.</p>\n<h3 id=\"7-cb\">7. Cb</h3>\n<p>Cb is a new element on the world wide web. Well, not so new anymore, and the full name is likely more familiar:\n<a href=\"http://cb.openmolecules.net/\">Chemical blogspace</a>. This social web application brings together blogging chemists\nworld wide. Oh, and this meme is picked up nicely:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/chemMeme.png\"/></p>\n<h3 id=\"8-chemical-abstracts\">8. chemical abstracts</h3>\n<p>No, not the database, but the nice graphical article abstracts in chemistry journals. <a href=\"http://www.chemfeeds.com/\">ChemFeeds</a>\ngets is all together. BTW, there remains very much to be done about improving publishing chemistry.\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2006/09/08/chemical-archeology-oscar3-to.html\">I <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\n<a href=\"http://chem-bla-ics.blogspot.com/2007/02/rsc-first-publisher-to-go-semantic.html\">blogged</a>\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2009/03/22/journal-of-cheminformatics-i-hope.html\">about <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\n<a href=\"http://chem-bla-ics.blogspot.com/2007/10/how-blogosphere-changes-publishing.html\">that</a>\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2009/03/19/nature-chemistry-improves-publishing.html\">repeatedly <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.</p>\n<h3 id=\"9-organometallics\">9. organometallics</h3>\n<p><a href=\"http://en.wikipedia.org/wiki/Organometallic_chemistry\">Organometallics</a> is, like metabolomics, a really\ninteresting area, with lots of complexities (pun intended :). Actually, I am not even aware of a\norganometallics/metabolomics mashup. Anyone with some nice pointers? I have not blogged about it much,\nand <a href=\"https://chem-bla-ics.linkedchemistry.info/2006/12/30/modern-chemistry-in-cdk-beyond-two.html\">the one time <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\nI did was in relation\nto chemical graph theory.</p>\n<h3 id=\"10-sparkling-fire\">10. sparkling fire</h3>\n<p>Burning things. Nothing more to say about that, I guess. Well, perhaps. Chemists like burning things;\nothers might too, but chemists at least. Blowing up things too. When I was a student, I had a very\nfriendly colleague who liked blowing up things and made TNT himself and took that to university too\n(stabilized, mind you :). Cool!</p>\n<p>Anyway, while googling for something to spice up this tenth item, I ran into this book:\n<a href=\"http://www.amazon.com/Caveman-Chemistry-Projects-Creation-Production/dp/1581125666?ie=UTF8&amp;link_code=btl&amp;camp=213689&amp;creative=392969\">Caveman Chemistry: 28 Projects, from the Creation of Fire to the Production of Plastics</a>.\nThe <a href=\"http://www.cavemanchemistry.com/browse.html\">prologue</a> nicely writes up that you need to sparkle\nsome fire in education to get the students enlightened:</p>\n<blockquote>\n<p>I teach chemistry at Hampden-Sydney College, a small liberal-arts college in central Virginia. The students\nhere, by and large, do not come equipped with insatiable curiosity about my discipline and experience has\nconvinced me that the profession of professing has more to do with motivation than with explanation; a student\nwho is not curious will resist even the most valiant attempts at compulsory education; conversely, inquiring\nminds want to know. A great deal of my time, then, has been spent devising tricks, gimmicks, schemes and plots\nfor leading stubborn horses to water, knowing full well that I can't make them think.</p>\n</blockquote>\n<p>Now, that leaves me with tagging a few further blogs to tag to continue the meme. The meme is spreading fast,\nso I hope I do not tag someone who already is tagged. <a href=\"http://usefulchem.blogspot.com/\">Jean-Claude</a>,\n<a href=\"http://wwmm.ch.cam.ac.uk/blogs/murrayrust/\">Peter</a>, <a href=\"http://baoilleach.blogspot.com/\">Noel</a>,\n<a href=\"http://depth-first.com/\">Rich</a>, <a href=\"http://www.chemspider.com/blog/\">Antony</a>, would you mind letting us know your\nten favourite chemistry things?</p>","doi":"https://doi.org/10.59350/be00d-tn533","guid":"https://doi.org/10.59350/be00d-tn533","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1273708800,"reference":[{"id":"https://doi.org/10.1021/cg060872y","unstructured":"Unknown title"}],"rid":"k2113-tjf27","summary":"Yeah, it's my turn. Standing on the shoulders of ChemJobber, Azmanam, and ScienceBase, here's list of things I like about chemistry. To put things into perspective first, a bit, I note that ChemJobber and Azmanam focused on wet-lab chemistry, and David on fancy molecules. Now, I am a theoretical chemist, and was thinking on what to orient the things I like, and on how general to make them. This meme is not easy, you now.","tags":["Chemistry"],"title":"My favourite chemistry things","updated_at":1785871017,"url":"https://chem-bla-ics.linkedchemistry.info/2010/05/13/my-favourite-chemistry-things.html","version":"v1"}},{"document":{"authors":[{"affiliation":[{"id":"https://ror.org/048a87296","name":"Uppsala University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>I had to dig deep to find posts on QSAR modeling. There are quite a few on <a href=\"http://chem-bla-ics.blogspot.com/search?q=qsar+bioclipse\">QSAR in Bioclipse</a>,\nbut that focuses on the descriptor calculation. In a quick scan, I could only spot two modeling posts:</p>\n<ul>\n<li><a href=\"http://chem-bla-ics.blogspot.com/2008/04/cdkmetabolomicschemometrics.html\">The CDK/Metabolomics/Chemometrics Unconference results</a></li>\n<li><a href=\"https://chem-bla-ics.linkedchemistry.info/2005/11/08/when-to-stop-including-qsar-model.html\">When to stop including QSAR model variables\u2026 <i class=\"fa-solid fa-recycle fa-xs\"></i></a></li>\n</ul>\n<p>Given the prominent place <a href=\"https://chem-bla-ics.linkedchemistry.info/2008/03/01/todo-april-2nd-defend-my-phd-work.html\">QSAR has in my thesis <i class=\"fa-solid fa-recycle fa-xs\"></i></a>,\nthis is somewhat surprising. Anyway, here is some more QSAR modeling talk.</p>\n<p><a href=\"http://gilleain.blogspot.com/\">Gilleain</a> <a href=\"http://www.blogger.com/github.com/gilleain/signatures\">implemented</a> the signature descriptors developed by Faulon et al.\n(see doi:<a href=\"https://doi.org/10.1021/ci020345w\">10.1021/ci020345w</a>; I <a href=\"http://chem-bla-ics.blogspot.com/2006/02/novel-qsar-and-qspr-descriptors_24.html\">mentioned the paper in 2006</a>),\nand the <a href=\"http://sourceforge.net/tracker/?func=detail&amp;aid=3017759&amp;group_id=20024&amp;atid=320024\">CDK patch</a> is currently being reviewed.\nWith some transformations, the atomic signatures for a molecule can be transformed into a fixed-length numerical representation:\n<code class=\"language-plaintext highlighter-rouge\">[70:1, 54:1, 23:1, 22:1, 9:9, 45:2]</code>. This string means that atomic signature 70 occurs once in this molecule and signature 9 occurs\nnine times. At this moment, I am not yet concerned about the actual signature, but just checking how well these signature can be used\nin QSPR modeling.</p>\n<p><a href=\"http://blog.rguha.net/\">Rajarshi</a>'s <a href=\"http://cran.r-project.org/web/packages/fingerprint/index.html\">fingerprint</a> code provides a good\ntemplate to parse this into a X matrix in <a href=\"http://www.r-project.org/\">R</a>:</p>\n<p>For my test case, I have used the boiling point data I used in my thesis paper <em>On the Use of 1H and 13C 1D NMR Spectra as QSPR Descriptors</em>\n(see doi:<a href=\"https://doi.org/10.1021/ci050282s\">10.1021/ci050282s</a>). Some of this data is actually <a href=\"http://www.chemspider.com/blog/gathering-physicochemical-data-onto-chemspider.html\">available from ChemSpider</a>,\nbut I do not think I ever uploaded the boiling point data. This constitutes a data set with 277 molecules, and my paper provides\nsome reference model quality statistics; that way, I have something to compare against. Moreover, I can use my previous scripts\nto do the PLS modeling (there are many <a href=\"http://www.google.se/search?q=tutorial+partial+least+squares\">tutorials online</a>, but you\ncan always buy an expensive book like the one shown on the right, if you really have to), (10-fold) cross-validation (CV), and\n5 repeats of random sampling.</p>\n<p>I strongly suggest people interested in statistical modeling to read this\n<a href=\"http://baoilleach.blogspot.com/2010/06/non-random-method-to-improve-your-qsar.html\">interesting post from Noel</a>: whatever test\nset sampling method you use, you <strong><em>must</em></strong> do some repeats to learn about the sensitivity of your modeling approach to changes\nin the data set. Depending on the actual sampling approach, you might see different sizes of variance, but until you measure it,\nyou will not know. For my application, these are the numbers:</p>\n<div class=\"language-R highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code><span class=\"c1\"># source(\"pls.R\")</span><span class=\"w\">\n</span><span class=\"n\">Read</span><span class=\"w\"> </span><span class=\"m\">277</span><span class=\"w\"> </span><span class=\"n\">items</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">66</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.987</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.921</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">31.37</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">66</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.983</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.924</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">12.405</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">65</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.985</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.949</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">38.503</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">63</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.983</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.948</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">36.981</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">65</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.986</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.923</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">21.49</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">64</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.983</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.91</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">17.759</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">64</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.983</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.921</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">17.062</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">66</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.986</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.94</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">40.311</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">66</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.982</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.927</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">13</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">68</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.986</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.929</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">16.23</span><span class=\"w\">\n</span></code></pre></div></div>\n<p>I know 42 is the answer to the universe, but 42 latent variables (LVs)?!? Well, it's just a start. A more accurate number of LVs\nseems to be around 15, but my script had to make the transition from the old pls.pcr package to the newer pls package. And I have\nyet to discover how I can get the new package to return me the lowest number of LVs for which the CV statistic is no longer\nsignificantly different from the best (see my paper how that works). Actually, I have set the maximum LVs to consider to 1/5th of\nthe number of objects (which is about the accepted ratio in the QSAR community); otherwise, it would have happily increased.</p>\n<p>However, the five repeats nicely show the variance in the quality statistics, R\u00b2, Q\u00b2, and root mean square error of prediction\n(RMSEP). From the numbers, a model with Q\u00b2 = 0.94 is <strong>not</strong> better than one with Q\u00b2 = 0.93 (and I have seen the variance quite some\nlarger). Bottom line: just measure that variability, and put it in the publication, will you??</p>\n<p>Anyway, what we all have been waiting for: the prediction results visualized (in black the CV predictions; in red the test set\npredictions):</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/signaturePrediction.png\"/></p>\n<p>Well, there is still much work to do, and you can expect the result to get better. Part of statistical modeling is to find\nthe source of variance, and I have yet to explore a few of them. For example, what are the effects of:</p>\n<ul>\n<li>creating signature from the hydrogen-depleted graph</li>\n<li>effect of tautomerism (see <a href=\"http://www.springerlink.com/content/l3p3t7066645/?p=bff6cd9b91bd40c59aa0d7afe11cf78a&amp;pi=0\">this special issue</a>)</li>\n<li>effect of the height of the signature</li>\n</ul>\n<p>And there are so many other things I like to do. But this will do for now.</p>","doi":"https://doi.org/10.59350/8d6w8-avm05","guid":"https://doi.org/10.59350/8d6w8-avm05","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1276992000,"reference":[{"id":"https://doi.org/10.1021/ci020345w","unstructured":"Unknown title"},{"id":"https://doi.org/10.1021/ci050282s","unstructured":"Unknown title"}],"rid":"3pvr9-k3592","summary":"I had to dig deep to find posts on QSAR modeling. There are quite a few on QSAR in Bioclipse, but that focuses on the descriptor calculation.","tags":["Cdk","Chemometrics"],"title":"QSPR modeling with signatures","updated_at":1785871016,"url":"https://chem-bla-ics.linkedchemistry.info/2010/06/20/qspr-modeling-with-signatures.html","version":"v1"}},{"document":{"authors":[{"affiliation":[{"id":"https://ror.org/02jz4aj89","name":"Maastricht University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>About four and a half years ago, I started <a href=\"http://rdf.openmolecules.net/\">OpenMolecules RDF</a>, a spin off from\n<a href=\"http://cb.openmolecules.net/\">Chemical blogspace</a> (Cb, which is still up and running thanks to Peter Maas!) where\nI started <a href=\"https://chem-bla-ics.linkedchemistry.info/2007/07/31/rdf-ing-molecular-space.html\">using InChIs in URIs <i class=\"fa-solid fa-recycle fa-xs\"></i></a>. My\ninterest came from the dereferencability, the ability to take an InChI and find information about the chemical\nstructure representated by it. Because information about anything is scattered around the internet, and we need\nsomething <a href=\"http://chem-bla-ics.blogspot.com/2007/08/centralized-or-decentralized.html\">decentralized</a>. Moreover,\nat the time searching of InChIs with search engines like Google did not work well at all: InChIs were tokenized\nin inconvenient ways.</p>\n<p>Originally, these URIs for InChIs were provided (and still are) by Cb, this July five years ago:</p>\n<div class=\"language-plaintext highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code>http://cb.openmolecules.net/rdf/?InChI=1/CH4/h1H4\n</code></pre></div></div>\n<p>for which soon after a separate domain was instantiated (thanx to Geoff!):</p>\n<div class=\"language-plaintext highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code>http://rdf.openmolecules.net/?InChI=1/CH4/h1H4\n</code></pre></div></div>\n<p>Mind you, <strong>OpenMolecules RDF</strong> is a decent citizen of the Linked Open Data network, though not much linked to.\nThe <a href=\"https://github.com/egonw/chembl.rdf\">ChEMBL-RDF</a> data is, and love to hear if there are other link sets\npointing there. On the outlinking side, it points to <a href=\"http://www.ebi.ac.uk/chebi/\">ChEBI</a> (via\n<a href=\"http://www.bio2rdf.org/\">Bio2RDF</a>), <a href=\"http://dbpedia.org/\">DBPedia</a>, <a href=\"http://www.chemspider.com/\">ChemSpider</a>\n(for 10k structures), the <a href=\"http://chem-bla-ics.blogspot.com/2009/03/nmrshiftdb-enters-rdfopenmoleculesnet.html\">NMRShiftDB</a>,\nand Cb itself. This post describes the adding of the <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/02/17/dbpedia-enters-rdfopenmoleculesnet.html\">link to DBPedia <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.</p>\n<p>In the past few years, I have written up bits on OpenMolecules RDF. The main reference is our chapter in <em>Beautiful Data</em> (Willighagen, 2010),\nwhere I used the <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/02/27/solubility-data-in-bioclipse-3-finding.html\">URIs for the solubility data <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.\nIt was later also described in the <em>Linking the Resource Description Framework to cheminformatics and proteochemometrics paper</em> (Willighagen, 2011),\nand another book chapter (Guha, 2011).</p>\n<p>This blog features a few more use cases, such as the ability to use these URIs to bookmark molecules or to\n<a href=\"http://chem-bla-ics.blogspot.com/2007/09/tagging-molecules-mashup-of-connotea.html\">annotate them with tags with Connotea</a>\n(which resulted in a nice <a href=\"http://chem-bla-ics.blogspot.com/2007/10/lunch-at-nature-hq-with-euan-joanna-ian.html\">lunch with the Nature people at the time</a>).\nThe link to Connotea is disabled at the moment, though.</p>\n<p>At this moment the system still holds, though there is problem in that browsers can put practical limits on\nURIs length, which limits the maximum size of the InChI. Virtuoso does this too.</p>","doi":"https://doi.org/10.59350/eg04s-efd96","guid":"https://doi.org/10.59350/eg04s-efd96","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1334448000,"reference":[{"id":"https://doi.org/10.1038/npre.2010.4918.1","unstructured":"Unknown title"},{"id":"https://doi.org/10.1002/9781118026038.ch24","unstructured":"Unknown title"},{"id":"https://doi.org/10.1186/2041-1480-2-s1-s6","unstructured":"Unknown title"}],"rid":"kpw7g-39w56","summary":"About four and a half years ago, I started OpenMolecules RDF, a spin off from Chemical blogspace (Cb, which is still up and running thanks to Peter Maas!) where I started using InChIs in URIs . My interest came from the dereferencability, the ability to take an InChI and find information about the chemical structure representated by it. Because information about anything is scattered around the internet, and we need something decentralized.","tags":["Chemistry","Rdf","Inchi","Opendata"],"title":"Dereferencable InChIs: OpenMolecules RDF","updated_at":1785871013,"url":"https://chem-bla-ics.linkedchemistry.info/2012/04/15/dereferencable-inchis-openmolecules-rdf.html","version":"v1"}},{"document":{"authors":[{"affiliation":[{"id":"https://ror.org/02jz4aj89","name":"Maastricht University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>A good number of years ago, a colleague and I explored if we could get access to the <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/02/16/retraction-data-in-wikidata.html/retractiondatabase.org/\">Retraction Watch Database</a>,\nbut we could not afford it. We have been using data on retractions for curate our databases, like\n<a href=\"https://www.wikipathways.org/\">WikiPathways</a>. A database should not contain knowledge based on (only) a retracted article.\nWikidata, btw, has a small number (499) of statements supported by retracted articles. Similarly, it turns out that I am\n<a href=\"https://w.wiki/8pwe\">citing retracted articles in two papers</a> (and a preprint of one of them).</p>\n<p><a href=\"https://www.wikidata.org/\">Wikidata</a> has a good number of retracted articles in their database\n(<a href=\"https://scholia.toolforge.org/statistics\">some 21 thousand at the time of writing</a>). A lot of this data\ncomes from CrossRef, that recently <a href=\"https://www.crossref.org/blog/news-crossref-and-retraction-watch/\">acquired the Retraction Watch Database</a>\n(doi:<a href=\"https://doi.org/10.13003/c23rw1d9\">10.13003/c23rw1d9</a>)) and started providing the content as FAIR and Open data.\nWith <a href=\"https://github.com/egonw/ons-wikidata/blob/main/RetractionWatch/quickstatements.groovy\">a Bacting-based script</a>\nI am regularly updating Wikidata with annotations from CrossRef, giving a rich dataset in Wikidata around\nthe queries. Over the past few years I have written various SPARQL queries to show the results which today\nI <a href=\"https://bigcat-um.github.io/sparql-examples/examples/WikidataRetractions/\">collected under a single home</a>:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/retraction_SPARQL.png\"/></p>","doi":"https://doi.org/10.59350/w4zj3-mbw53","guid":"https://doi.org/10.59350/w4zj3-mbw53","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1739664000,"reference":[{"id":"https://doi.org/10.1093/nar/gkad960","unstructured":"Agrawal, A., Balc\u0131, H., Hanspers, K., Coort, S. L., Martens, M., Slenter, D. N., Ehrhart, F., Digles, D., Waagmeester, A., Wassink, I., Abbassi-Daloii, T., Lopes, E. N., Iyer, A., Acosta, J. M., Willighagen, L. G., Nishida, K., Riutta, A., Basaric, H., Evelo, C. T., \u2026 Pico, A. R. (2023). WikiPathways 2024: next generation pathway database. <i>Nucleic Acids Research</i>, <i>52</i>(D1), D679\u2013D689."},{"id":"https://doi.org/10.13003/c23rw1d9","unstructured":"Crossref, Hendricks, G., Center for Scientific Integrity&amp; Lammey, R. (2023). <i>Crossref acquires Retraction Watch data and opens it for the scientific community</i>. Crossref.  <b>[cito:citesAsEvidence]</b>"}],"rid":"e8vfg-wqz89","summary":"A good number of years ago, a colleague and I explored if we could get access to the Retraction Watch Database, but we could not afford it. We have been using data on retractions for curate our databases, like WikiPathways. A database should not contain knowledge based on (only) a retracted article. Wikidata, btw, has a small number (499) of statements supported by retracted articles.","tags":["Wikidata","Wikipathways"],"title":"Retracted articles in Wikidata","updated_at":1785870031,"url":"https://chem-bla-ics.linkedchemistry.info/2025/02/16/retraction-data-in-wikidata.html","version":"v1"}},{"document":{"authors":[{"affiliation":[{"id":"https://ror.org/02jz4aj89","name":"Maastricht University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>Names of chemicals are part of the human user experience when browsing a chemical database. And literature too,\nof course. Chemical names are also not easy to use, and what a chemical name means is not always clear.\nThis is why the <a href=\"https://en.wikipedia.org/wiki/International_Union_of_Pure_and_Applied_Chemistry\">IUPAC</a>\nstarted a standardizing nomenclature in chemistry, the <em>IUPAC names</em>. Each IUPAC name uniquely defines\nthe chemical structure it defines. For example, <em>methane</em> is the IUPAC name for the chemical CH<sub>4</sub>.</p>\n<p>So, when propagating chemical structures from the <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/02/13/beiltein-journal-has-bioschemas.html\">Beilstein Bioschemas feed</a>,\nI was looking for names, IUPAC or not, ideally the name used in the article. When I asked about this,\nthe question came up if they could autogenerate IUPAC names, for which\n<a href=\"https://doi.org/10.1038/s41598-021-94082-y\">various</a>\n<a href=\"https://doi.org/10.1186/s13321-021-00535-x\">new</a>\n<a href=\"https://doi.org/10.1186/s13321-021-00512-4\">tools</a>\n<a href=\"https://doi.org/10.1186/s13321-024-00941-x\">exist</a>\n(I think I am missing one from an American team, but cannot find the reference),\nalong with multiple established commerical tools.\nBecause the IUPAC nomenclature is a long list of naming rules, priorities, etc, a rule-based\nalgorithm is logical, but newer methods take a deep-learning approach.</p>\n<p>Back to the chemical annotation of chemistry literature. This is of obvious interest: you want\nto know where we can read more about a certain chemical. We need the chemical structures in\na database for that, linked to the articles. This is, of course, one of the original studies\nof <em>cheminformatics</em>. And when authors of the chemical literature do not provide this routinely\n(<a href=\"https://chem-bla-ics.linkedchemistry.info/2025/02/13/beiltein-journal-has-bioschemas.html\">this post</a>\nshows a few exceptions, but it is still all too rare). And then manual and automated curation\nis needed, e.g. done by <a href=\"https://en.wikipedia.org/wiki/Chemical_Abstracts_Service\">Chemical Abstracts</a>.</p>\n<p>Third, <a href=\"https://wikidata.org/\">Wikidata</a> has <a href=\"https://scholia.toolforge.org/chemical/\">about 1.4 million</a>\nchemical compounds and many names. A <a href=\"https://www.wikidata.org/wiki/Wikidata:Property_proposal/Pending#IUPAC_name\">property propoal for IUPAC names</a>\nhas been long pending, but once accepted in one form or another, will require IUPAC names too.</p>\n<h2 id=\"one-million-iupac-names\">One million IUPAC names</h2>\n<p>Thus, the idea came up, can we create a set of 1 million unique IUPAC names found in literature?\nI asked on the <a href=\"https://elixir-europe.org/\">ELIXIR Europe</a> slack channel if <a href=\"https://europepmc.org/\">Europe PMC</a>\nhad such a dataset (doi:<a href=\"https://doi.org/10.1093/nar/gkad1085\">10.1093/nar/gkad1085</a>). I knew they had been adding chemical\n<a href=\"https://scholia.toolforge.org/topic/Q403574\">named-entity recognition</a> (NER) results in\n<a href=\"https://europepmc.org/Annotations\">their annotation API</a>. I learned they used <a href=\"https://www.ebi.ac.uk/chebi/\">ChEBI</a>.\nMelanie Vollmar and Summer Rosonovski or Europe PMC gave useful information and support.\n<a href=\"https://cpm.lumc.nl/research/bioinformatics-224/magnus-palmblad-5\">Magnus Palmblad</a> also replied\nand provided Python code to use the Europe PMC API to fetch names it returns and see if those\nare IUPAC names. Well, that's easy. We have <a href=\"https://opsin.ch.cam.ac.uk/\">OPSIN</a> for that\n(see doi:<a href=\"https://doi.org/10.1021/ci100384d\">10.1021/ci100384d</a>).</p>\n<p>Unfortunately, the Europe PMC NER results are not ideal for IUPAC names. Just scanning\nsome 5, 6 organic chemistry journals returned some 8 thousand IUPAC names in open access\narticles. But it quickly started to be too limited: each set of articles returned\nincreasingly few new names. The reason is simple: the NER is too <em>greedy</em> and as a\nresult, does not easily recognize longer IUPAC names. It is too happy with a substring\nof the IUPAC name. For example, when it encounters the IUPAC name <em>5-Bromo-1H-indole-3-carboxylic acid</em>,\nit settles for <em>indole-3-carboxylic acid</em>:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/greedy.png\"/></p>\n<h2 id=\"open-source-chemistry-analysis-routines\">Open-Source Chemistry Analysis Routines</h2>\n<p>During my PhD, in 2003, when I worked a few months with Prof. <a href=\"https://scholia.toolforge.org/author/Q908710\">Peter Murray-Rust</a> (University of Cambridge)\nand Prof. Janet Thornthon (EMBL-EBI), I learned about the research by <a href=\"https://scholia.toolforge.org/author/Q28946549\">Sam Adams</a>\n(doi:<a href=\"https://doi.org/10.1039/B411699M\">10.1039/B411699M</a>), <a href=\"https://scholia.toolforge.org/author/Q133040220\">Joe Townsend</a>\n(doi:<a href=\"https://doi.org/10.1039/B411033A\">10.1039/B411033A</a>), and <a href=\"https://scholia.toolforge.org/author/Q90318722\">Peter Corbett</a>\n(doi:<a href=\"https://doi.org/10.1007/11875741_11\">10.1007/11875741_11</a>). One of the tools that used\nthis research was (is) <a href=\"https://scholia.toolforge.org/topic/Q133037490\">OSCAR</a>,\nshort for <em>Open-Source Chemistry Analysis Routines</em> (see <a href=\"https://blogs.ch.cam.ac.uk/pmr/2009/05/16/opsin-and-oscar-chemical-language-processing/\">this detailed write up by Peter MR</a>).\nLater, in 2010 I visted Peter again, as postdoc, in Cambridge, and then\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2010/10/15/working-on-oscar-for-three-months.html\">worked on the OSCAR project</a> too.\nAnd while OSCAR did a lot more, the integration of <a href=\"https://chem-bla-ics.linkedchemistry.info/2010/12/26/oscar-training-data-models-etc.html\">Corbett's NER research</a>\nmade OSCAR the obvious follow-up step in finding IUPAC names in literature.</p>\n<p>And because <a href=\"https://chem-bla-ics.linkedchemistry.info/2011/09/27/almost-year-ago-i-started-position-with.html\">OSCAR4 had been integrated into Bioclipse</a>\n(doi:<a href=\"https://doi.org/10.1186/1758-2946-3-41\">10.1186/1758-2946-3-41</a>) and I had this ported to Bacting already\n(doi:<a href=\"https://doi.org/10.21105/joss.02558\">10.21105/joss.02558</a>), using this was trivial.\nThe use of Europe PMC is different now, however, and we are no longer using the Annotations API,\nbut just using it to find open access articles, and to get the full text in XML format.\nThat allows a simple XPath search on <code class=\"language-plaintext highlighter-rouge\">&lt;p&gt;</code> elements, pass the resulting string to OSCAR4,\nand the recognized names are checked with OPSIN.\nAnd with this approach, processing two of the five or six journals we earlier explored,\nwe find another 40+ thousand IUPAC names. Quite a success, I am tempted to say.</p>\n<h2 id=\"a-blue-obelisk-project\">A Blue Obelisk project</h2>\n<p>So, I started a new <a href=\"https://blueobelisk.github.io/\">Blue Obelisk</a> project,\n<a href=\"https://github.com/BlueObelisk/iupac-names\">iupac-names</a>, to collect 1M IUPAC names. For researchers\nto use, learn from, etc. Just IUPAC names. Not even the chemical structure, nor the link to the\narticles. The first is trivial to do with OPSIN, so the matching SMILES do not need to be stored.\nLinks to literature is tricky because of the aforementioned issues, and we only want to know\nwhich (partial) IUPAC names occur in literature. If you really want to know in which articles\nthat IUPAC name is found, you can simply do a search in Europe PMC.</p>\n<p>And because we only store IUPAC names, this are very basic facts (this is an IUPAC name, as defined\nby OPSIN being able to generate a SMILES for this structure) and that that string occurs in\nsome article) and we can share them as CCZero. We <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/03/08/iupac-names.html/is:issue\" title=\"milestone release\">defined various milestones</a>,\nand I am happy that the first two have been reached within two weeks:</p>\n<ul>\n<li><a href=\"https://github.com/BlueObelisk/iupac-names/releases/tag/milestone-10k\">Milestone 10k</a> (doi:<a href=\"https://doi.org/10.5281/zenodo.14965762\">10.5281/zenodo.14965762</a>)</li>\n<li><a href=\"https://github.com/BlueObelisk/iupac-names/releases/tag/milestone-50k\">Milestone 50k</a> (doi:<a href=\"https://doi.org/10.5281/zenodo.14978557\">10.5281/zenodo.14978557</a>)</li>\n</ul>\n<p>This second milestone has 53848 unique names, but as literature goes, there are interesting\nvariations, some likely because of typesetting leading to spaces added and missing. If\nwe ignore spaces and hyphens, we have 50534 names left (hence the milestone). But IUPAC\nnames are also not fully unique, partly because of Unicode character variations and greek\nletter alternatives, and you may wonder how many different chemical structures this set\nreflects. While not perfect, the Standard InChI gives some lower limit, and we find 36528\nInChIKeys in this second milestone.</p>\n<p>Now, we need twenty times as much to reach the 1M IUPAC names, but given we have many, many\nmore open access articles to process. The bottleneck seems to be mostly our workflow.</p>\n<h3 id=\"can-you-contribute\">Can you contribute?</h3>\n<p>Yes, of course! This is an open science project. But please keep in mind the narrow focus of this\nproject: only IUPAC names which can be found in (open access) literature. This project doed not accept\nautogenerated names (PubChem would have given use many millions already), nor IUPAC names from existing\ndatabases. Ideally, you are able to show the code you use to extract/find those names in literature.</p>\n<h3 id=\"can-i-use-these-names\">Can I use these names?</h3>\n<p>First of all, this is what the CCZero license and open science nature of this project is about: reuse.\nWe love to hear how you are using these names, tho, and we encourage you to write up how you\nare using them. You can use <a href=\"https://datacite.org/\">DataCite</a> to cite the release you used,\nand citing this blog post by DOI is also possible.</p>\n<h3 id=\"does-it-support-my-language-too\">Does it support my language too?</h3>\n<p>No, at this moment it only support IUPAC names in English. Dutch, French, Spanish, or Chinese\nIUPAC names are valid, but currently not supported. See also\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2010/12/30/text-mining-chemistry-from-dutch-or.html\">this post</a>.</p>\n<h3 id=\"will-there-be-a-publication\">Will there be a publication?</h3>\n<p>Magnus and I intend so. We already submitted an abstract to the <a href=\"https://iccs-nl.org/\">International Conference on Chemical Structures</a>,\nwhich has <a href=\"https://www.biomedcentral.com/collections/ICCS25\">a Collection in the Journal of Cheminformatics</a>.\nIf the abstract gets accepted, of course, we can submit there. Otherwise, we will look for another venue,\nlikely <a href=\"https://en.wikipedia.org/wiki/Diamond_open_access\">diamond open access</a>.</p>\n<h3 id=\"where-is-your-script\">Where is your script?</h3>\n<p>Ah, fair point. We did not decide on the final license yet. I have used two scripts based on the template\nby Magnus. As soon as we have finalized the license, we will make those available.</p>","doi":"https://doi.org/10.59350/tjkf2-k1608","guid":"https://doi.org/10.59350/tjkf2-k1608","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1741392000,"reference":[{"id":"https://doi.org/10.1038/s41598-021-94082-y","unstructured":"Krasnov, L., Khokhlov, I., Fedorov, M. V.&amp; Sosnin, S. (2021). Transformer-based artificial neural networks for the conversion between chemical notations. <i>Scientific Reports</i>, <i>11</i>(1)."},{"id":"https://doi.org/10.1186/s13321-021-00512-4","unstructured":"Rajan, K., Zielesny, A.&amp; Steinbeck, C. (2021). STOUT: SMILES to IUPAC names using neural machine translation. <i>Journal of Cheminformatics</i>, <i>13</i>(1)."},{"id":"https://doi.org/10.1186/s13321-021-00535-x","unstructured":"Handsel, J., Matthews, B., Knight, N. J.&amp; Coles, S. J. (2021). Translating the InChI: adapting neural machine translation to predict IUPAC names from a chemical identifier. <i>Journal of Cheminformatics</i>, <i>13</i>(1)."},{"id":"https://doi.org/10.1186/s13321-024-00941-x","unstructured":"Rajan, K., Zielesny, A.&amp; Steinbeck, C. (2024). STOUT V2.0: SMILES to IUPAC name conversion using transformer models. <i>Journal of Cheminformatics</i>, <i>16</i>(1)."},{"id":"https://doi.org/10.1021/ci100384d","unstructured":"Lowe, D. M., Corbett, P. T., Murray-Rust, P.&amp; Glen, R. C. (2011). Chemical Name to Structure: OPSIN, an Open Source Solution. <i>Journal of Chemical Information and Modeling</i>, <i>51</i>(3), 739\u2013753."},{"id":"https://doi.org/10.1039/b411699m","unstructured":"Adams, S. E., Goodman, J. M., Kidd, R. J., McNaught, A. D., Murray-Rust, P., Norton, F. R., Townsend, J. A.&amp; Waudby, C. A. (2004). Experimental data checker: better information for organic chemists. <i>Organic & Biomolecular Chemistry</i>, <i>2</i>(21), 3067."},{"id":"https://doi.org/10.1039/b411033a","unstructured":"Townsend, J. A., Adams, S. E., Waudby, C. A., de Souza, V. K., Goodman, J. M.&amp; Murray-Rust, P. (2004). Chemical documents: machine understanding and automated information extraction. <i>Organic & Biomolecular Chemistry</i>, <i>2</i>(22), 3294."},{"id":"https://doi.org/10.1007/11875741_11","unstructured":"Corbett, P.&amp; Murray-Rust, P. (2006). High-Throughput Identification of Chemistry in Life Science Texts. In <i>Lecture Notes in Computer Science</i> (pp. 107\u2013118). Springer Berlin Heidelberg."},{"id":"https://doi.org/10.1186/1758-2946-3-41","unstructured":"Jessop, D. M., Adams, S. E., Willighagen, E. L., Hawizy, L.&amp; Murray-Rust, P. (2011). OSCAR4: a flexible architecture for chemical text-mining. <i>Journal of Cheminformatics</i>, <i>3</i>(1).  <b>[cito:usesMethodIn]</b>"},{"id":"https://doi.org/10.21105/joss.02558","unstructured":"Willighagen, E. (2021). Bacting: a next generation, command line version of Bioclipse. <i>Journal of Open Source Software</i>, <i>6</i>(62), 2558.  <b>[cito:usesMethodIn]</b>"},{"id":"https://doi.org/10.1093/nar/gkad1085","unstructured":"Rosonovski, S., Levchenko, M., Bhatnagar, R., Chandrasekaran, U., Faulk, L., Hassan, I., Jeffryes, M., Mubashar, S. I., Nassar, M., Jayaprabha\u00a0Palanisamy, M., Parkin, M., Poluru, J., Rogers, F., Saha, S., Selim, M., Shafique, Z., Ide-Smith, M., Stephenson, D., Tirunagari, S., \u2026 Harrison, M. (2023). Europe PMC in 2023. <i>Nucleic Acids Research</i>, <i>52</i>(D1), D1668\u2013D1676.  <b>[cito:usesMethodIn]</b>"},{"id":"https://doi.org/10.5281/zenodo.14965762","unstructured":"Egon Willighagen. (2025). <i>BlueObelisk/iupac-names: Milestone 10k</i> (Version milestone-10k) [Dataset]. Zenodo.  <b>[cito:citesAsEvidence]</b>"},{"id":"https://doi.org/10.5281/zenodo.14978557","unstructured":"Egon Willighagen. (2025). <i>BlueObelisk/iupac-names: Milestone 50k</i> (Version milestone-50k) [Dataset]. Zenodo.  <b>[cito:citesAsEvidence]</b>"}],"rid":"a2g45-bjb85","summary":"Names of chemicals are part of the human user experience when browsing a chemical database. And literature too, of course. Chemical names are also not easy to use, and what a chemical name means is not always clear. This is why the IUPAC started a standardizing nomenclature in chemistry, the IUPAC names. Each IUPAC name uniquely defines the chemical structure it defines. For example, methane is the IUPAC name for the chemical CH4.","tags":["Iupac","Cheminf","Oscar","Textmining","Europepmc"],"title":"One Million IUPAC names","updated_at":1785870029,"url":"https://chem-bla-ics.linkedchemistry.info/2025/03/08/iupac-names.html","version":"v1"}},{"document":{"authors":[{"affiliation":[{"id":"https://ror.org/02jz4aj89","name":"Maastricht University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>As part of our <a href=\"https://www.nwo.nl/en/\">Dutch Research Council</a> (NWO) <a href=\"https://www.nwo.nl/en/projects/osf232097\">Open Science grant</a>,\nwe organized a <a href=\"https://cdk.github.io/nwo-openscience-2024/\">Chemistry Development Kit User Group Meeting</a>\n(<a href=\"https://hashtags-hub.toolforge.org/CDK25UGM\">#CDK25UGM</a>), of which yesterday was the \"conference\" day, and today a hackathon.</p>\n<p>I opened the session with a few slides welcoming everyone at Maastricht University (and our\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2025/01/27/translational-genomics.html\">Dept of Translational Genomics</a>,\nand explaining the NWO grant.\n<a href=\"https://orcid.org/0000-0001-7730-2646\">John Mayfield</a> (<a href=\"https://www.nextmovesoftware.com/\">NextMove</a>) spoke about\n\"What's New\" in the Chemistry Development Kit 2.10, e.g. explaining more about the new (much faster) <code class=\"language-plaintext highlighter-rouge\">AtomContainer</code>,\nSMIRKS, and more.</p>\n<p>After lunch, <a href=\"https://orcid.org/0000-0003-1554-6666\">Jonas Schaub</a> (<a href=\"https://www.uni-jena.de/en/\">Friedrich Schiller University Jena</a>)\nshowed various projects where the CDK is used, titled  \"Scaffolds, Functional Groups, Aglycones: Algorithmic Substructure Identification with CDK\"\n(see doi:<a href=\"https://doi.org/10.1186/s13321-023-00762-4\">10.1186/s13321-023-00762-4</a>, doi:<a href=\"https://doi.org/10.1186/s13321-022-00656-x\">10.1186/s13321-022-00656-x</a>,\nand doi:<a href=\"https://doi.org/10.1186/s13321-020-00467-y\">10.1186/s13321-020-00467-y</a>).\nLyudvika Radeva (<a href=\"https://www.ideaconsult.net/\">Ideaconsult Ltd</a>, <a href=\"https://uni-plovdiv.bg/en/\">University of Plovdiv</a>) showed\nwhat SYBYL Line Notation (SLN) is and how this is implemented in Ambit (see doi:<a href=\"https://doi.org/10.1002/minf.202100027\">10.1002/minf.202100027</a>).\n<a href=\"https://orcid.org/0000-0002-4354-4353\">Sonja Herres-Pawlis</a> (<a href=\"https://www.rwth-aachen.de/\">RWTH Aachen University</a>)\nupdated us with \"News from the InChI: making the InChI FAIR and including inorganics\", e.g. showing how\nthey worked out how the InChI is going to handle organometalics, where the bonds and the stereochemistry\nas aspects that were not handled by the current InChI.</p>\n<p>After the afternoon coffee break, <a href=\"https://orcid.org/0000-0003-3662-2621\">Zhixu Ni</a> (<a href=\"https://fedorovalab.net/team/zhixu-ni/\">TU Dresden</a>)\nshowed his work on lipid maps characterization and identification. We previously met a few times\nat EpiLipidNET COST action meetings, and it was great to see his continued research on representation\nof lipids and lipid classes in hit \"A Fuzzy Solution for Lipid Structures Using CXSMILES\".</p>\n<p>Finally, <a href=\"https://www.linkedin.com/in/matthiasmailaender/\">Matthias Mail\u00e4nder</a> (<a href=\"https://www.lablicate.com/\">Lablicate GmbH</a>)\ngave a \"Live demo of where <a href=\"https://github.com/OpenChrom\">OpenChrom</a> uses the CDK\", and\n<a href=\"https://research.rug.nl/en/persons/yajie-ding\">Yajie Ding</a> (University of Groningen) told her about her\nglycoscience research. There, cheminformatics can also greatly help and the CDK may provide\nthem with solutions.</p>\n<p>This really doesn't do justice to all the discussions, examples, use cases, etc. But it gives you\nan idea. We had 11 people in the room, and were joined online by an additional 6 people.</p>","doi":"https://doi.org/10.59350/e08pe-thb38","funding_references":[{"awardNumber":"osf232097","awardTitle":"The Chemistry Development Kit in 2024: improving cheminformatics research","awardUri":"https://www.nwo.nl/en/projects/osf232097","funderIdentifier":"https://ror.org/04jsz6e67","funderIdentifierType":"ROR","funderName":"Dutch Research Council"}],"guid":"https://doi.org/10.59350/e08pe-thb38","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1741651200,"reference":[{"id":"https://doi.org/10.1002/minf.202100027","unstructured":"Kochev, N., Jeliazkova, N.&amp; Tancheva, G. (2021). Ambit\u2010SLN: an Open Source Software Library for Processing of Chemical Objects via SLN Linear Notation. <i>Molecular Informatics</i>, <i>40</i>(11)."},{"id":"https://doi.org/10.1186/s13321-022-00656-x","unstructured":"Schaub, J., Zander, J., Zielesny, A.&amp; Steinbeck, C. (2022). Scaffold Generator: a Java library implementing molecular scaffold functionalities in the Chemistry Development Kit (CDK). <i>Journal of Cheminformatics</i>, <i>14</i>(1)."},{"id":"https://doi.org/10.1186/s13321-023-00762-4","unstructured":"Chandrasekhar, V., Sharma, N., Schaub, J., Steinbeck, C.&amp; Rajan, K. (2023). Cheminformatics Microservice: unifying access to open cheminformatics toolkits. <i>Journal of Cheminformatics</i>, <i>15</i>(1)."},{"id":"https://doi.org/10.1186/s13321-020-00467-y","unstructured":"Schaub, J., Zielesny, A., Steinbeck, C.&amp; Sorokina, M. (2020). Too sweet: cheminformatics for deglycosylation in natural products. <i>Journal of Cheminformatics</i>, <i>12</i>(1)."}],"rid":"smhnn-nr982","summary":"As part of our Dutch Research Council (NWO) Open Science grant, we organized a Chemistry Development Kit User Group Meeting (#CDK25UGM), of which yesterday was the \"conference\" day, and today a hackathon.","tags":["Cdk","Openscience","Cdk2024"],"title":"cdk2024 #4: Chemistry Development Kit User Group Meeting - Day 1","updated_at":1785870028,"url":"https://chem-bla-ics.linkedchemistry.info/2025/03/11/CDK-UGM.html","version":"v1"}}],"items":[{"authors":[{"affiliation":[{"id":"https://ror.org/02mb95055","name":"Birkbeck, University of London"}],"contributor_roles":[],"family":"Eve","given":"Martin Paul","url":"https://orcid.org/0000-0002-5589-8511"}],"blog":{"authors":[{"name":"Martin Paul Eve","url":"https://orcid.org/0000-0002-5589-8511"}],"community_id":"9224b0d7-fc03-497c-9c6f-85c9fd1e72da","created":1690329600,"current_feed_url":null,"description":null,"doi":"https://doi.org/10.59348/eve","favicon":"https://rogue-scholar.org/api/communities/9224b0d7-fc03-497c-9c6f-85c9fd1e72da/logo","feed_format":"application/atom+xml","feed_url":"https://eve.gd/feed_all.xml","filter":null,"generator":"Jekyll","home_page_url":"https://eve.gd","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59348","relative_url":null,"secure":true,"slug":"eve","status":"active","subfield":"1208","title":"Martin Paul Eve","updated":1785660546,"use_api":true},"blog_name":"Martin Paul Eve","blog_slug":"eve","content_html":"<p>This morning, I made a last-minute incorporation of Kornbrodt, Joseph, and David Zappone, 'Special Feature on <em>To the Journey: Looking Back at</em> Star Trek: Voyager', interview with Kate Mulgrew, 455 Films, 2026, Bluray into my <em>Voyager</em> book. I had watched the documentary itself in a digital copy earlier in the year. It's OK. Nothing very new. You do get a VERY enthusiastic Garrett Wang shepherding the program along with a slightly strange foray into a zero-gravity plane dive. I am unsure this has anything to do with <em>Voyager</em>, but it was mildly entertaining. And it's Harry Kim! And he clearly loves it and the <em>Voyager</em> fan community and all that goes with it. It's all quite heartwarming.</p>\n<p>But yesterday, the Bluray arrived and that has special features. And I really enjoyed the special feature interviews on this documentary! There was also some bad editing. Mulgrew's interview has a section that repeats itself and that is very frustrating as a viewer. However, I suppose that this is mostly b-reel footage. Also, the interviewer is not very good at moving her on beyond the \"women still struggle to have it all\" question (not that that isn't important).</p>\n<p>However, there is some great stuff in here!</p>\n<p>Just Mulgrew's work ethic. 3am starts, midnight finishes. Every part of her was invested in The Work. (In fact, so many of the interviewees talk with reverence about The Work. It's all there is, for them. A total dedication to it.) She talks of how she had to brace every part of her body's musculature in a specific way when playing Janeway. Her acting is a total body investment. She speaks of how playing Janeway for seven years meant that she was not there for her children, ever. She made a conscious choice to do this and, to this day, feels guilt. Her children have never seen her play Janeway. They have watched none of <em>Voyager</em> because, she says, they consider it the reason they had a motherless childhood. This seems to have had very deep, lasting family psychological problems. This is, then, an incredible commitment. A life given over to the role, really.</p>\n<p>She also states that she does not believe that introducing Janeway would work now. She believes that \"as a result of the digital age\" audiences \"are attracted to darkness\" and, for Mulgrew, \"Janeway was not dark. Janeway was light\". <em>Voyager</em> was, for her, an optimistic Trek, not a dark, cynical Trek. This is such a crucial point for my book. It is also very heartening to hear Mulgrew talk of \"the Janeway effect\" that has encouraged women in STEM and space. She talks of how female fans come up to her at conventions and spill their souls about how they went into science because of her. My heart did sink a little, as a man. Because I would love (but am too shy) to meet her at a convention and tell her how inspirational <em>I</em>, even though I am a man, found her. I never questioned her authority or thought it out of place that she was in command. It was clear. Obvious. Just how it should be, I thought (although I was in my teenage years when the later series of <em>Voyager</em> was airing for the first time.) So, Kate, if you ever read this, please know: men took inspiration from you, too. That's why <a href=\"https://janeway.systems/\">I named our software platform after you</a>.</p>\n<p>There's also a lovely, if not terribly informative, interview with Jeri Taylor (RIP). I suppose there IS lots of interest here; especially about the open story submission process that they ran. They simply didn't have enough ideas and so invited anyone to pitch to them. She does say that this means she sat through a lot of VERY bad pitches! But the main thing I took away from her interview was a sense of sadness. She talked of the intense workload and non-stop need for fresh material. But, as closing remarks, she said she would never, ever write for fun now. The fun has been taken from her by decades of writing to a schedule. This was somewhat terrible, even though she said it half-jokingly.</p>\n<p>There are interviews with Michael Piller's son and wife as well on the disc that give an interesting background to his life. Although I just thought more about how Shawn Piller's life had been extraordinary. Being introduced to Gene Roddenberry! Writing stories/scripts for <em>Star Trek: The Next Generation</em>! All opportunities opened up through his father (although he had to prove himself, even though he had this leg-up.) An amazing life. Few have that experience.</p>\n<p>There's also a fun section on directing (with a lot of Armin Shimerman featured, who is GREAT and hilarious!) but I was disappointed not to hear, in this section, from Roxann Dawson, who is doubtless the most successful crossover actor/director from the <em>Voyager</em> cast. But hey. I think the documentary had varying levels of commitment from the cast and was not high on their priority list, sadly. (Although they had lots of Mulgrew's time, which was generous of her.)</p>\n<p>Anyway, all in all, this was an informative and enjoyable set of interviews. It's probably only of great interest to dedicated fans. It can otherwise feel niche and repetitive. But for the hardcore fan, this is a worthwhile set of bonuses. I have still to watch the sections with the Paramount Executives. They're not generally liked! Mulgrew points out: the bottom line is always the money! But I suppose their executive position is worth knowing about. If there's anything significant, I will probably update this post.</p>\n<p><a href=\"https://eve.gd/2026/08/02/the-extended-interviews-on-ito-the-journey-looking-back-ati-star-trek-voyager/\">The extended interviews on <i>To the Journey: Looking Back at</i> Star Trek: Voyager</a> was originally published by Martin Paul Eve at <a href=\"https://eve.gd\">eve.gd: Martin Paul Eve</a> on August 02, 2026.</p>","doi":"https://doi.org/10.59348/qygeq-6n059","guid":"https://doi.org/10.59348/qygeq-6n059","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1785628800,"rid":"aszdc-x4j13","summary":"This morning, I made a last-minute incorporation of Kornbrodt, Joseph, and David Zappone, 'Special Feature on <em> To the Journey: Looking Back at </em> Star Trek: Voyager', interview with Kate Mulgrew, 455 Films, 2026, Bluray into my <em> Voyager </em> book. I had watched the documentary itself in a digital copy earlier in the year. It's OK. Nothing very new.","title":"The extended interviews on <i>To the Journey: Looking Back at</i> Star Trek: Voyager","updated_at":1785873800,"url":"https://eve.gd/2026/08/02/the-extended-interviews-on-ito-the-journey-looking-back-ati-star-trek-voyager/","version":"v1"},{"authors":[{"affiliation":[{"id":"https://ror.org/048a87296","name":"Uppsala University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p><a href=\"http://www.nature.com/nchem/\">Nature Chemistry</a> just released the first issue with a few free papers,\nlike <em>Asymmetric total syntheses of (+)- and (-)-versicolamide B and biosynthetic implications</em> by Miller et al.\n(DOI:<a href=\"https://doi.org/10.1038/nchem.110\">10.1038/nchem.110</a>).</p>\n<p>Now, we've seen the Royal Society of Chemistry's <a href=\"http://chem-bla-ics.blogspot.com/search?q=project+prospect\">Project Prospect</a> <!-- keep link -->\n(see <a href=\"https://chem-bla-ics.linkedchemistry.info/2007/02/01/rsc-first-publisher-to-go-semantic.html\">RSC: the first publisher to go semantic! <i class=\"fa-solid fa-recycle fa-xs\"></i></a>)\nand ChemSpiders recent <a href=\"http://www.chemmantis.com/\">ChemMantis</a> system which enriches\nthe papers with machine readable representations of the molecules discussed in those\npapers. The new Nature publication has been in the works for a while, and they\n<a href=\"http://blogs.nature.com/thescepticalchymist/2008/05/jj_day_98_service_with_a_simpl.html\">asked</a>\nthe community before what a Nature Chemistry paper should like like, and I replied in\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2008/05/08/re-what-should-nature-chemistry-paper.html\">Re: What should a Nature Chemistry paper look like? <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.</p>\n<h2 id=\"the-verdict\">The verdict</h2>\n<p>So, have the been listening? Is the HTML they produce semantic? Is it data rich? Or is it\njust another hamburger? Well, I am very happy to see some of the suggestions I made picked\nup (though I do not fool myself in believing I am the only one that suggested those\nfeatures). A tour of good things, and points for improvement.</p>\n<p>The first impression is not shocking; it looks like any other interface, with molecules drawn as images in the paper:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem3.png\"/></p>\n<p>All structures that are numbered and linked (as in <em>C6-epi-stephacidin A (Compound <strong>13</strong>)</em>\nhave a hover-over function to popup a drawing of the structure:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem4.png\"/></p>\n<p>The popup image is a nice gimmick, but not really sematically useful. The link, however,\nis! It points to a separate supplementary page with further information which include\na image of the 2D structure and, following a link, the 3D structure in <a href=\"http://www.jmol.org/\">Jmol</a>.\nMoreover, it comes with the machine readable representations:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem5.png\"/></p>\n<p>This is indeed interesting, and a big step forward, though please do note my comments later.\nFor convenience, all molecules with such supplementary information is available from the\nspecial Chemical Compounds section of the paper:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem2.png\"/></p>\n<p>Excellent! This really is a step forward towards a data-rich paper! Indeed, I will shortly\nwrite up a <a href=\"http://www.bioclipse.net/\">Bioclipse</a> plugin for Nature Chemistry, which\nwill download all molecular structures based on the DOI! Anyway, more on that later\u2026\nFor this article, that table looks like:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem1.png\"/></p>\n<p>By now, you likely also noted the links to <a href=\"http://pubchem.ncbi.nlm.nih.gov/\">PubChem</a>, and\nindeed, upon publication of a paper, all structures are deposited in the public domain:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem6.png\"/></p>\n<p>At last but not least, each molecule is available in the <a href=\"http://en.wikipedia.org/wiki/Chemical_Markup_Language\">Chemical Markup Language</a>\n(with 2D coordinates)! And you know I am a very happy CML user for a long time (see e.g.\nPeter's recent blog <a href=\"https://doi.org/10.59350/jbq7c-szw40\">Egon Willighagen and CML <i class=\"fa-solid fa-recycle fa-xs\"></i></a>).\nBTW, one comment on the CML: the namespace used is the outdated namespace, <strong>not</strong>\nthe current one (see <a href=\"http://cmlexplained.blogspot.com/2007/06/there-can-be-only-one-namespace.html\">There can be only one (namespace)</a>).\n(But the <a href=\"http://cdk.sf.net/\">CDK</a> and Bioclipse will read it anyway.)</p>\n<h2 id=\"details-matter\">Details matter</h2>\n<p>So, while the first impression was not shocking, it was a bit deceptive. <em>Nature Chemistry</em>\nreally changes publishing of chemistry. But I have bad news too. They need to improve the\nHTML they produce.</p>\n<p>But before pointing out some missed chances, let me reply <em>inter alia</em> to Peter's recent\nwork on the Open Source plugin for including semantic chemistry in MS-Word documents\n(see <a href=\"https://doi.org/10.59350/wn2pv-gef13\">How can we publish semantic chemical documents? <i class=\"fa-solid fa-recycle fa-xs\"></i></a>):\nNature Chemistry seems to have done a great job with existing tools. Nevertheless, I fully\nback up Peters comment that while the plugin is useless without Word, the results produced\nwith the plugin are extremely Open Standard, and enormously reusable! Indeed, while the\nWord file format is only formally an true Open Standard, the file format is plain XML, and\nextracting content bearing the CML namespace is trivial.</p>\n<p>Which reminds me, if someone from the Nature Chemistry team is reading this, please point\nme to a blog what tools actually <em>are</em> involved in publishing a Nature Chemistry paper!\nI think we all like to know.</p>\n<p>Now, the <a href=\"http://en.wikipedia.org/wiki/HTML\">HTML</a> has room for improvement. First of all,\na look at the metadata defined for the web page of the article shows a <em>description</em>\nand <em>keywords</em> about the journal, not the article, and the same goes for the web pages for\nthe molecules:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem7.png\"/></p>\n<p>Additionally, the compound details web page has no special markup for the machine readable\ninformation:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem8.png\"/></p>\n<p>Or, if it does, it's still mixed with markup for visual pleasing output:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/nchem9.png\"/></p>\n<p>Still, the HTML is clean enough to have some regular expressions extract a good deal of\ninformation, and there is also still the PubChem deposition.</p>\n<h2 id=\"beyond-connection-tables\">Beyond connection tables</h2>\n<p>Like many other chemistry journals, Nature Chemistry does not consider properties of\nthe molecule interesting, and NMR spectra are hidden in the Supplementary Information.\nThis paper in particular, disregards a lot of machine readable facts by putting all\nexperimental section bits in a PDF document. So, the next challenge for Nature Chemistry\nwill be to get the authors of papers contribute the original spectra (JCAMP-DX, CMLSpect,\netc) in the supplementary information section. Better, have the raw data or even the NMR\npeak-atom annotations deposited in public repositories such (see \n<a href=\"https://chem-bla-ics.linkedchemistry.info/2009/03/04/open-nmr-data-raw-curves-and-annotated.html\">Open NMR data: raw curves and annotated peak lists <i class=\"fa-solid fa-recycle fa-xs\"></i></a>).</p>\n<p>All in all, I am rather positive about the first Nature Chemistry issue, and like to\nthank the editors and paper authors for there efforts on improving publishing chemistry!</p>","doi":"https://doi.org/10.59350/40377-hz881","guid":"https://doi.org/10.59350/40377-hz881","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1237420800,"reference":[{"id":"https://doi.org/10.1038/nchem.110","unstructured":"Unknown title"},{"id":"https://doi.org/10.59350/jbq7c-szw40","unstructured":"Unknown title"},{"id":"https://doi.org/10.59350/wn2pv-gef13","unstructured":"Unknown title"}],"rid":"wasde-08n67","summary":"Nature Chemistry just released the first issue with a few free papers, like Asymmetric total syntheses of (+)- and (-)-versicolamide B and biosynthetic implications by Miller et al. (DOI:10.1038/nchem.110).","tags":["Inchi","Chemistry","Jmol"],"title":"Nature Chemistry improves publishing chemistry: a detailed analysis","updated_at":1785873774,"url":"https://chem-bla-ics.linkedchemistry.info/2009/03/19/nature-chemistry-improves-publishing.html","version":"v1"},{"authors":[{"affiliation":[{"id":"https://ror.org/048a87296","name":"Uppsala University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>I have blogged about two Molecular Chemometrics principles so far:</p>\n<ul>\n<li><a href=\"https://chem-bla-ics.linkedchemistry.info/2010/08/09/molecular-chemometrics-principles-1.html\">McPrinciple #1: access to data</a></li>\n<li><a href=\"https://chem-bla-ics.linkedchemistry.info/2010/08/12/molecular-chemometrics-principles-2-be.html\">McPrinciple #2: be clear in what you mean</a></li>\n</ul>\n<p>Peter's post <a href=\"https://doi.org/10.59350/hphjc-qgr72\">#solo10: Green Chain Reaction; where to store the data? DSR? IR? BioTorrent, OKF or ??? <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\ngives me enough basis to write up a third principle:</p>\n<p><strong>Molecular Chemometrics Principles #3</strong>: We make scientific progress if we build on past achievements.</p>\n<p>Sounds logical, right? Practically, the way we share our cheminformatics knowledge makes this standing on shoulders pretty difficult.\nBut there is one particular aspect I would like to ask your attention for: you can contribute by making clear what shoulders\nyou would like to stand on. That is, where do you prefer to put your effort, and what message would you like to give to your user community.</p>\n<p>In the aforelinked post, Peter asks where he should upload his data, and he suggest <a href=\"http://www.biotorrents.net/\">BioTorrent</a> (see my review\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2010/04/18/bittorrents-for-science.html\">BitTorrents for Science <i class=\"fa-solid fa-recycle fa-xs\"></i></a>), DSpace, and <a href=\"http://www.ckan.net/\">CKAN</a>.\nNow, his <a href=\"http://www.google.se/search?sourceid=chrome&amp;client=ubuntu&amp;channel=cs&amp;ie=UTF-8&amp;q=%22Green+Chain+Reaction%22\">Green Chain Reaction</a>\nis picked up (see <a href=\"http://researchremix.wordpress.com/2010/08/11/green-chain-reaction-project-putting-my-minutes-where-my-mouth-is/\">these</a>\n<a href=\"http://scienceonlinelondon.wikidot.com/topics:green-chain-reaction\">few</a> <a href=\"https://doi.org/10.59350/h2jq5-3np88\">blog <i class=\"fa-solid fa-recycle fa-xs\"></i></a> posts),\nand the resulting data should be distributed as much as possible. The exact location does not really matter\u2026</p>\n<p>But\u2026</p>\n<p>By picking where you upload, you make a statement to your community: \"<em>Look guys, we are distributing our data via Foo, because we believe those guys are doing good work! Perhaps you can support them too.</em>\".</p>\n<p>This principle does not only apply to data, it applies to things too. For example, when\n<a href=\"http://www.chemspider.com/blog/ichemlabs-and-rsc-chemspider-announce-partnership.html\">iChemLabs and RSC ChemSpider Announce Partnership</a>\nthey do not just improve the user experience of ChemSpider (which I certainly won't object against), but they also imply\n\"<em>Look dudes, your product is just not good enough and we do not want to help you improve it either</em>\".\nOf course, ChemSpider has every right, and for them to succeed it is crucial to make decisions like this. Fortunately,\n<a href=\"http://web.chemdoodle.com/installation.php\">ChemDoodle is GPL</a>.</p>\n<p>Every project with a user base has the opportunity to support shoulders, if they only visibly stand on them. By merely discussion the\n<em>Green Chain Reaction</em>, I show to support this social web experiment. You can too. Use these powers wisely. May the McPrinciples be with you.</p>","doi":"https://doi.org/10.59350/832vn-qwh10","guid":"https://doi.org/10.59350/832vn-qwh10","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1281744000,"reference":[{"id":"https://doi.org/10.59350/hphjc-qgr72","unstructured":"Unknown title"},{"id":"https://doi.org/10.59350/h2jq5-3np88","unstructured":"Unknown title"}],"rid":"55qe3-cbt36","summary":"I have blogged about two Molecular Chemometrics principles so far:","tags":["Mcprinciples","Solo10","Chemdoodle","Chemspider","Javascript"],"title":"The Molecular Chemometrics Principles #3: stand on shoulders","updated_at":1785873773,"url":"https://chem-bla-ics.linkedchemistry.info/2010/08/14/molecular-chemometrics-principles-3.html","version":"v1"},{"authors":[{"affiliation":[{"id":"https://ror.org/048a87296","name":"Uppsala University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>About 6 months ago I <a href=\"http://chem-bla-ics.blogspot.com/2009/03/nmrshiftdb-enters-rdfopenmoleculesnet.html\">reported</a> about my efforts to RDF-ize the data from the\n<a href=\"http://www.nmrshiftdb.org/\">NMRShiftDB</a>. Since then, time was consumed by many other things, but now that <a href=\"http://www.bioclipse.net/\">Bioclipse</a> can query\n<a href=\"http://en.wikipedia.org/wiki/SPARQL\">SPARQL</a> end points, that I want to contribute the triple set (it is <a href=\"http://www.gnu.org/copyleft/fdl.html\">GNU FDL</a>-licensed)\nto <a href=\"http://www.bio2rdf.org/\">Bio2RDF</a>, that a student started working in my group (now larger than just me :) on reasoning on life sciences data, and that I\nrecently contributed my <a href=\"http://egonw.posterous.com/nmrshiftdb-1006-contributions-and-counting\">1000th NMR spectrum</a> to the database, I thought it was time to\nfinally reinstall <a href=\"http://www.openlinksw.com/wiki/main/Main/VOSDownload\">Virtuoso</a>.</p>\n<p>There are precompiled binaries for <a href=\"https://launchpad.net/~wdaniels/+archive/ppa\">Ubuntu</a> and <a href=\"http://bugs.debian.org/cgi-bin/bugreport.cgi?bug=508048\">Debian</a>,\nbut Michel encouraged me to use version 6 when <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/06/26/michel-dumontier-at-uppsala-university.html\">he visited us <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.\nAnd so I compiled and install <a href=\"https://sourceforge.net/projects/virtuoso/files/virtuoso-devel/6.0.0-TP1/\">6.0.0.TP1</a> on the public server, while I do have the\nbinary debs for 5.0.12 on my laptop. With some basic Apache magic, I hooked up the SPARQL end point of the server to the web:</p>\n<div class=\"language-xml highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code><span class=\"nt\">&lt;Proxy</span> <span class=\"err\">/nmrshiftdb/sparql</span><span class=\"nt\">&gt;</span>\n  RewriteEngine On\n  Allow from all\n  ProxyPass        http://localhost:8890/sparql\n  ProxyPassReverse http://localhost:8890/sparql\n<span class=\"nt\">&lt;/Proxy&gt;</span>\n</code></pre></div></div>\n<p>Nice thing about this is, that I can set up multiple servers, allowing me to keep incompatibly licensed data sets apart (see\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2009/05/18/open-data-license-rights-aggregation.html\">Open Data: license, rights, aggregation, clean interfaces? <i class=\"fa-solid fa-recycle fa-xs\"></i></a>), which is\nthe same approach Bio2RDF is taking.</p>\n<p>The <a href=\"http://pele.farmbio.uu.se/nmrshiftdb/sparql\">end point</a> now offers about <a href=\"http://pele.farmbio.uu.se/nmrshiftdb/sparql?default-graph-uri=&amp;query=SELECT+count%28*%29+WHERE+{\\%0D%0A++%3Fs+%3Fp+%3Fo+.%0D%0A}&amp;format=text%2Fhtml&amp;debug=on\">278887</a>\ntriples, but this will soon rise as I make more content from the database available in the original SQL database. The data is from the\n<a href=\"https://sourceforge.net/projects/nmrshiftdb/files/nmrshiftdb/1.3.3/\">1.3.3 release</a> by <a href=\"http://www.steinbeck-molecular.de/steinblog/\">Chris</a>'\nteam, and does not include my 1000th spectrum.</p>\n<p>Getting the data into the database was not trivial either. The documentation suggests WebDAV, and that indeed worked for me once, after\nusing the <a href=\"http://www.snee.com/bobdc.blog/2009/02/getting-started-using-virtuoso.html\">curl approach suggested here</a>. But upon a second upload, it\ndid again not enter the store. The ultimate solution was to use the iSQL interface, with the following SQL</p>\n<div class=\"language-plaintext highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code>DB.DBA.RDF_LOAD_RDFXML_MT(\n  file_to_string_output('/tmp/nmrshiftdb.rdf'), '',\n  'http://pele.farmbio.uu.se/nmrshiftdb'\n);\n</code></pre></div></div>\n<p>Scientifically, this progress is not overly interesting, although it makes it very clear that you really should not have to be happy with proprietary\nand non-semantic formats for anything. But, to me, this is mostly a technological success of great importance: I can now share really large sets of\nRDF data.</p>\n<p>Querying this data is a simple with SPARQL, and the results are available in various formats, such as JSON, which makes it easy to integrate in\nthird-party applications or <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/09/02/google-wave-robot-for-cdk-functionality.html\">Google Wave robots <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\n(did I hear someone say <a href=\"http://nmrshifty.appspot.com/\">NMRShifty</a>?). As I have <a href=\"http://chem-bla-ics.blogspot.com/search?q=sparql\">blogged before</a>,\nSPARQL is an excellent tool to aggregate scientific data prior to data analysis. And I will demo more interesting queries later this month.</p>","doi":"https://doi.org/10.59350/nv925-tje87","guid":"https://doi.org/10.59350/nv925-tje87","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1252022400,"rid":"rgwsf-1q277","summary":"About 6 months ago I reported about my efforts to RDF-ize the data from the NMRShiftDB.","tags":["Rdf","Sparql","Nmrshiftdb","Cheminf"],"title":"NMRShiftDB enters rdf.openmolecules.net #2: SPARQL end point with Virtuoso","updated_at":1785871018,"url":"https://chem-bla-ics.linkedchemistry.info/2009/09/04/nmrshiftdb-enters-rdfopenmoleculesnet-2.html","version":"v1"},{"authors":[{"affiliation":[{"id":"https://ror.org/048a87296","name":"Uppsala University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>Yeah, it's my turn. Standing on the shoulders of <a href=\"http://chemjobber.blogspot.com/2010/05/my-favorite-things-about-chemistry.html\">ChemJobber</a>,\n<a href=\"http://www.chemistry-blog.com/2010/05/07/a-few-of-my-favorite-chemistry-things/\">Azmanam</a>, and\n<a href=\"http://www.sciencebase.com/science-blog/my-favourite-chemistry-things.html\">ScienceBase</a>, here's list of things I like about chemistry.\nTo put things into perspective first, a bit, I note that ChemJobber and Azmanam focused on wet-lab chemistry, and David on fancy\nmolecules. Now, I am a theoretical chemist, and was thinking on what to orient the things I like, and on how general to make them.\nThis meme is not easy, you now. But here goes:</p>\n<h3 id=\"1-chemical-graph-theory\">1. chemical graph theory</h3>\n<p>Chemical graph theory is one of the common theoretical models chemists work with to make sense of chemical properties.\nI like it because the graph theory is fairly straightforward, but chemistry adds enough color (literally!) to create a\nnice complexity that kept the cheminformatics field going strong for more than 50 years now :) For example, how to adapt\nthe theory to <a href=\"https://chem-bla-ics.linkedchemistry.info/2006/12/30/modern-chemistry-in-cdk-beyond-two.html\">deal with mutli-atom bonds <i class=\"fa-solid fa-recycle fa-xs\"></i></a> :)</p>\n<h3 id=\"2-rare-nuclei-in-the-nmrshiftdb\">2. rare nuclei in the NMRShiftDB</h3>\n<p>The <a href=\"http://www.nmrshiftdb.org/\">NMRShiftDB</a> is an Open Data repository for annotated NMR spectra. The fun here is to\nadd NMR spectra of <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/09/05/nmrshiftdb-rdf-2-some-statistics.html?q=nmr+nuclei+sparql\">rare nuclei <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.\nDon't you just love a\nmolecule with NMR shifts for all atoms?</p>\n<h3 id=\"3-metabolomics\">3. metabolomics</h3>\n<p>Metabolomics is the research field that studies the small molecules of life. Plant metabolomics is particularly fun.\nTens of thousands of molecules, and a lot of metabolite identification to be done, and much more. Lot's of cool stuff\nto do here, and I am trying to secure funding for it. This is\n<a href=\"http://chem-bla-ics.blogspot.com/search?q=metabolomics\">what I blogged about metabolomics before</a>. What about\nhis nice <a href=\"http://en.wikipedia.org/wiki/Secondary_metabolite\">secondary metabolite</a> (source:\n<a href=\"http://en.wikipedia.org/\">Wikipedia</a>, <a href=\"http://en.wikipedia.org/wiki/File:Discodermolide_Structure.png\">CC0</a>):</p>\n<p><img alt=\"\" src=\"https://upload.wikimedia.org/wikipedia/commons/e/e4/Discodermolide_Structure.png\"/></p>\n<h3 id=\"4-hexavalent-carbon\">4. hexavalent carbon</h3>\n<p>Atom types is another theoretical model for chemistry. <a href=\"https://chem-bla-ics.linkedchemistry.info/2007/07/01/atom-typing-in-cdk.html\">Atom typing <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\nis one of the underlying technologies of <a href=\"http://en.wikipedia.org/wiki/Force_field_%28chemistry%29\">force fields</a>, which are used\nin many, many fields in chemistry. Now, force fields typically take only a subset of atom types. New atom types, consequently, need\nto be added. One such new atom type was the <a href=\"http://www.ch.ic.ac.uk/rzepa/blog/?p=811\">hexavalent carbon</a>. Rare, very rare, but\njust the amount of complexity I like about chemistry.</p>\n<!-- Image lost -->\n<h3 id=\"5-self-organizing-maps\">5. self-organizing maps</h3>\n<p>Kohonen maps, or <a href=\"http://en.wikipedia.org/wiki/Self-organizing_map\">self-organizing maps</a> (SOM), are a machine learning\nmethod that have interesting visualization features. They have numerous applications, and also in chemistry. The group\nwhere I did my PhD developed a supervised SOM, which I used them to classify crystal structures\n(doi:<a href=\"https://doi.org/10.1021/cg060872y\">10.1021/cg060872y</a>). Another of my favorites is the\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2006/04/04/mining-kegg-pathway-database-with-self.html\">reaction classification by Aires-de-Sousa <em>et al.</em> <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\nusing unsupervised SOMs.</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/som_thesis_image.png\"/></p>\n<h3 id=\"6-the-maillard-reaction\">6. the Maillard reaction</h3>\n<p>People who know me personally, know that I like tasting things. That also makes me have to worry about overweight.\nTaste is to a large extend governed by cooking, and the <a href=\"http://en.wikipedia.org/wiki/Maillard_reaction\">Maillard reaction</a>\nplays an important role here. If you like to learn more about the chemistry of cooking, checkout these\n<a href=\"http://chemistandcook.blogspot.com/\">two</a> <a href=\"http://blog.khymos.org/\">blogs</a>.</p>\n<h3 id=\"7-cb\">7. Cb</h3>\n<p>Cb is a new element on the world wide web. Well, not so new anymore, and the full name is likely more familiar:\n<a href=\"http://cb.openmolecules.net/\">Chemical blogspace</a>. This social web application brings together blogging chemists\nworld wide. Oh, and this meme is picked up nicely:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/chemMeme.png\"/></p>\n<h3 id=\"8-chemical-abstracts\">8. chemical abstracts</h3>\n<p>No, not the database, but the nice graphical article abstracts in chemistry journals. <a href=\"http://www.chemfeeds.com/\">ChemFeeds</a>\ngets is all together. BTW, there remains very much to be done about improving publishing chemistry.\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2006/09/08/chemical-archeology-oscar3-to.html\">I <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\n<a href=\"http://chem-bla-ics.blogspot.com/2007/02/rsc-first-publisher-to-go-semantic.html\">blogged</a>\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2009/03/22/journal-of-cheminformatics-i-hope.html\">about <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\n<a href=\"http://chem-bla-ics.blogspot.com/2007/10/how-blogosphere-changes-publishing.html\">that</a>\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2009/03/19/nature-chemistry-improves-publishing.html\">repeatedly <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.</p>\n<h3 id=\"9-organometallics\">9. organometallics</h3>\n<p><a href=\"http://en.wikipedia.org/wiki/Organometallic_chemistry\">Organometallics</a> is, like metabolomics, a really\ninteresting area, with lots of complexities (pun intended :). Actually, I am not even aware of a\norganometallics/metabolomics mashup. Anyone with some nice pointers? I have not blogged about it much,\nand <a href=\"https://chem-bla-ics.linkedchemistry.info/2006/12/30/modern-chemistry-in-cdk-beyond-two.html\">the one time <i class=\"fa-solid fa-recycle fa-xs\"></i></a>\nI did was in relation\nto chemical graph theory.</p>\n<h3 id=\"10-sparkling-fire\">10. sparkling fire</h3>\n<p>Burning things. Nothing more to say about that, I guess. Well, perhaps. Chemists like burning things;\nothers might too, but chemists at least. Blowing up things too. When I was a student, I had a very\nfriendly colleague who liked blowing up things and made TNT himself and took that to university too\n(stabilized, mind you :). Cool!</p>\n<p>Anyway, while googling for something to spice up this tenth item, I ran into this book:\n<a href=\"http://www.amazon.com/Caveman-Chemistry-Projects-Creation-Production/dp/1581125666?ie=UTF8&amp;link_code=btl&amp;camp=213689&amp;creative=392969\">Caveman Chemistry: 28 Projects, from the Creation of Fire to the Production of Plastics</a>.\nThe <a href=\"http://www.cavemanchemistry.com/browse.html\">prologue</a> nicely writes up that you need to sparkle\nsome fire in education to get the students enlightened:</p>\n<blockquote>\n<p>I teach chemistry at Hampden-Sydney College, a small liberal-arts college in central Virginia. The students\nhere, by and large, do not come equipped with insatiable curiosity about my discipline and experience has\nconvinced me that the profession of professing has more to do with motivation than with explanation; a student\nwho is not curious will resist even the most valiant attempts at compulsory education; conversely, inquiring\nminds want to know. A great deal of my time, then, has been spent devising tricks, gimmicks, schemes and plots\nfor leading stubborn horses to water, knowing full well that I can't make them think.</p>\n</blockquote>\n<p>Now, that leaves me with tagging a few further blogs to tag to continue the meme. The meme is spreading fast,\nso I hope I do not tag someone who already is tagged. <a href=\"http://usefulchem.blogspot.com/\">Jean-Claude</a>,\n<a href=\"http://wwmm.ch.cam.ac.uk/blogs/murrayrust/\">Peter</a>, <a href=\"http://baoilleach.blogspot.com/\">Noel</a>,\n<a href=\"http://depth-first.com/\">Rich</a>, <a href=\"http://www.chemspider.com/blog/\">Antony</a>, would you mind letting us know your\nten favourite chemistry things?</p>","doi":"https://doi.org/10.59350/be00d-tn533","guid":"https://doi.org/10.59350/be00d-tn533","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1273708800,"reference":[{"id":"https://doi.org/10.1021/cg060872y","unstructured":"Unknown title"}],"rid":"k2113-tjf27","summary":"Yeah, it's my turn. Standing on the shoulders of ChemJobber, Azmanam, and ScienceBase, here's list of things I like about chemistry. To put things into perspective first, a bit, I note that ChemJobber and Azmanam focused on wet-lab chemistry, and David on fancy molecules. Now, I am a theoretical chemist, and was thinking on what to orient the things I like, and on how general to make them. This meme is not easy, you now.","tags":["Chemistry"],"title":"My favourite chemistry things","updated_at":1785871017,"url":"https://chem-bla-ics.linkedchemistry.info/2010/05/13/my-favourite-chemistry-things.html","version":"v1"},{"authors":[{"affiliation":[{"id":"https://ror.org/048a87296","name":"Uppsala University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>I had to dig deep to find posts on QSAR modeling. There are quite a few on <a href=\"http://chem-bla-ics.blogspot.com/search?q=qsar+bioclipse\">QSAR in Bioclipse</a>,\nbut that focuses on the descriptor calculation. In a quick scan, I could only spot two modeling posts:</p>\n<ul>\n<li><a href=\"http://chem-bla-ics.blogspot.com/2008/04/cdkmetabolomicschemometrics.html\">The CDK/Metabolomics/Chemometrics Unconference results</a></li>\n<li><a href=\"https://chem-bla-ics.linkedchemistry.info/2005/11/08/when-to-stop-including-qsar-model.html\">When to stop including QSAR model variables\u2026 <i class=\"fa-solid fa-recycle fa-xs\"></i></a></li>\n</ul>\n<p>Given the prominent place <a href=\"https://chem-bla-ics.linkedchemistry.info/2008/03/01/todo-april-2nd-defend-my-phd-work.html\">QSAR has in my thesis <i class=\"fa-solid fa-recycle fa-xs\"></i></a>,\nthis is somewhat surprising. Anyway, here is some more QSAR modeling talk.</p>\n<p><a href=\"http://gilleain.blogspot.com/\">Gilleain</a> <a href=\"http://www.blogger.com/github.com/gilleain/signatures\">implemented</a> the signature descriptors developed by Faulon et al.\n(see doi:<a href=\"https://doi.org/10.1021/ci020345w\">10.1021/ci020345w</a>; I <a href=\"http://chem-bla-ics.blogspot.com/2006/02/novel-qsar-and-qspr-descriptors_24.html\">mentioned the paper in 2006</a>),\nand the <a href=\"http://sourceforge.net/tracker/?func=detail&amp;aid=3017759&amp;group_id=20024&amp;atid=320024\">CDK patch</a> is currently being reviewed.\nWith some transformations, the atomic signatures for a molecule can be transformed into a fixed-length numerical representation:\n<code class=\"language-plaintext highlighter-rouge\">[70:1, 54:1, 23:1, 22:1, 9:9, 45:2]</code>. This string means that atomic signature 70 occurs once in this molecule and signature 9 occurs\nnine times. At this moment, I am not yet concerned about the actual signature, but just checking how well these signature can be used\nin QSPR modeling.</p>\n<p><a href=\"http://blog.rguha.net/\">Rajarshi</a>'s <a href=\"http://cran.r-project.org/web/packages/fingerprint/index.html\">fingerprint</a> code provides a good\ntemplate to parse this into a X matrix in <a href=\"http://www.r-project.org/\">R</a>:</p>\n<p>For my test case, I have used the boiling point data I used in my thesis paper <em>On the Use of 1H and 13C 1D NMR Spectra as QSPR Descriptors</em>\n(see doi:<a href=\"https://doi.org/10.1021/ci050282s\">10.1021/ci050282s</a>). Some of this data is actually <a href=\"http://www.chemspider.com/blog/gathering-physicochemical-data-onto-chemspider.html\">available from ChemSpider</a>,\nbut I do not think I ever uploaded the boiling point data. This constitutes a data set with 277 molecules, and my paper provides\nsome reference model quality statistics; that way, I have something to compare against. Moreover, I can use my previous scripts\nto do the PLS modeling (there are many <a href=\"http://www.google.se/search?q=tutorial+partial+least+squares\">tutorials online</a>, but you\ncan always buy an expensive book like the one shown on the right, if you really have to), (10-fold) cross-validation (CV), and\n5 repeats of random sampling.</p>\n<p>I strongly suggest people interested in statistical modeling to read this\n<a href=\"http://baoilleach.blogspot.com/2010/06/non-random-method-to-improve-your-qsar.html\">interesting post from Noel</a>: whatever test\nset sampling method you use, you <strong><em>must</em></strong> do some repeats to learn about the sensitivity of your modeling approach to changes\nin the data set. Depending on the actual sampling approach, you might see different sizes of variance, but until you measure it,\nyou will not know. For my application, these are the numbers:</p>\n<div class=\"language-R highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code><span class=\"c1\"># source(\"pls.R\")</span><span class=\"w\">\n</span><span class=\"n\">Read</span><span class=\"w\"> </span><span class=\"m\">277</span><span class=\"w\"> </span><span class=\"n\">items</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">66</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.987</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.921</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">31.37</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">66</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.983</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.924</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">12.405</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">65</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.985</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.949</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">38.503</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">63</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.983</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.948</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">36.981</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">65</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.986</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.923</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">21.49</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">64</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.983</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.91</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">17.759</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">64</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.983</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.921</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">17.062</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">66</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.986</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.94</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">40.311</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">66</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.982</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.927</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">13</span><span class=\"w\">\n   </span><span class=\"n\">rank</span><span class=\"o\">=</span><span class=\"m\">68</span><span class=\"w\">     </span><span class=\"n\">LV</span><span class=\"o\">=</span><span class=\"m\">42</span><span class=\"w\">     </span><span class=\"n\">R2</span><span class=\"o\">=</span><span class=\"m\">0.986</span><span class=\"w\">     </span><span class=\"n\">Q2</span><span class=\"o\">=</span><span class=\"m\">0.929</span><span class=\"w\">  </span><span class=\"n\">RMSEP</span><span class=\"o\">=</span><span class=\"m\">16.23</span><span class=\"w\">\n</span></code></pre></div></div>\n<p>I know 42 is the answer to the universe, but 42 latent variables (LVs)?!? Well, it's just a start. A more accurate number of LVs\nseems to be around 15, but my script had to make the transition from the old pls.pcr package to the newer pls package. And I have\nyet to discover how I can get the new package to return me the lowest number of LVs for which the CV statistic is no longer\nsignificantly different from the best (see my paper how that works). Actually, I have set the maximum LVs to consider to 1/5th of\nthe number of objects (which is about the accepted ratio in the QSAR community); otherwise, it would have happily increased.</p>\n<p>However, the five repeats nicely show the variance in the quality statistics, R\u00b2, Q\u00b2, and root mean square error of prediction\n(RMSEP). From the numbers, a model with Q\u00b2 = 0.94 is <strong>not</strong> better than one with Q\u00b2 = 0.93 (and I have seen the variance quite some\nlarger). Bottom line: just measure that variability, and put it in the publication, will you??</p>\n<p>Anyway, what we all have been waiting for: the prediction results visualized (in black the CV predictions; in red the test set\npredictions):</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/signaturePrediction.png\"/></p>\n<p>Well, there is still much work to do, and you can expect the result to get better. Part of statistical modeling is to find\nthe source of variance, and I have yet to explore a few of them. For example, what are the effects of:</p>\n<ul>\n<li>creating signature from the hydrogen-depleted graph</li>\n<li>effect of tautomerism (see <a href=\"http://www.springerlink.com/content/l3p3t7066645/?p=bff6cd9b91bd40c59aa0d7afe11cf78a&amp;pi=0\">this special issue</a>)</li>\n<li>effect of the height of the signature</li>\n</ul>\n<p>And there are so many other things I like to do. But this will do for now.</p>","doi":"https://doi.org/10.59350/8d6w8-avm05","guid":"https://doi.org/10.59350/8d6w8-avm05","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1276992000,"reference":[{"id":"https://doi.org/10.1021/ci020345w","unstructured":"Unknown title"},{"id":"https://doi.org/10.1021/ci050282s","unstructured":"Unknown title"}],"rid":"3pvr9-k3592","summary":"I had to dig deep to find posts on QSAR modeling. There are quite a few on QSAR in Bioclipse, but that focuses on the descriptor calculation.","tags":["Cdk","Chemometrics"],"title":"QSPR modeling with signatures","updated_at":1785871016,"url":"https://chem-bla-ics.linkedchemistry.info/2010/06/20/qspr-modeling-with-signatures.html","version":"v1"},{"authors":[{"affiliation":[{"id":"https://ror.org/02jz4aj89","name":"Maastricht University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>About four and a half years ago, I started <a href=\"http://rdf.openmolecules.net/\">OpenMolecules RDF</a>, a spin off from\n<a href=\"http://cb.openmolecules.net/\">Chemical blogspace</a> (Cb, which is still up and running thanks to Peter Maas!) where\nI started <a href=\"https://chem-bla-ics.linkedchemistry.info/2007/07/31/rdf-ing-molecular-space.html\">using InChIs in URIs <i class=\"fa-solid fa-recycle fa-xs\"></i></a>. My\ninterest came from the dereferencability, the ability to take an InChI and find information about the chemical\nstructure representated by it. Because information about anything is scattered around the internet, and we need\nsomething <a href=\"http://chem-bla-ics.blogspot.com/2007/08/centralized-or-decentralized.html\">decentralized</a>. Moreover,\nat the time searching of InChIs with search engines like Google did not work well at all: InChIs were tokenized\nin inconvenient ways.</p>\n<p>Originally, these URIs for InChIs were provided (and still are) by Cb, this July five years ago:</p>\n<div class=\"language-plaintext highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code>http://cb.openmolecules.net/rdf/?InChI=1/CH4/h1H4\n</code></pre></div></div>\n<p>for which soon after a separate domain was instantiated (thanx to Geoff!):</p>\n<div class=\"language-plaintext highlighter-rouge\"><div class=\"highlight\"><pre class=\"highlight\"><code>http://rdf.openmolecules.net/?InChI=1/CH4/h1H4\n</code></pre></div></div>\n<p>Mind you, <strong>OpenMolecules RDF</strong> is a decent citizen of the Linked Open Data network, though not much linked to.\nThe <a href=\"https://github.com/egonw/chembl.rdf\">ChEMBL-RDF</a> data is, and love to hear if there are other link sets\npointing there. On the outlinking side, it points to <a href=\"http://www.ebi.ac.uk/chebi/\">ChEBI</a> (via\n<a href=\"http://www.bio2rdf.org/\">Bio2RDF</a>), <a href=\"http://dbpedia.org/\">DBPedia</a>, <a href=\"http://www.chemspider.com/\">ChemSpider</a>\n(for 10k structures), the <a href=\"http://chem-bla-ics.blogspot.com/2009/03/nmrshiftdb-enters-rdfopenmoleculesnet.html\">NMRShiftDB</a>,\nand Cb itself. This post describes the adding of the <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/02/17/dbpedia-enters-rdfopenmoleculesnet.html\">link to DBPedia <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.</p>\n<p>In the past few years, I have written up bits on OpenMolecules RDF. The main reference is our chapter in <em>Beautiful Data</em> (Willighagen, 2010),\nwhere I used the <a href=\"https://chem-bla-ics.linkedchemistry.info/2009/02/27/solubility-data-in-bioclipse-3-finding.html\">URIs for the solubility data <i class=\"fa-solid fa-recycle fa-xs\"></i></a>.\nIt was later also described in the <em>Linking the Resource Description Framework to cheminformatics and proteochemometrics paper</em> (Willighagen, 2011),\nand another book chapter (Guha, 2011).</p>\n<p>This blog features a few more use cases, such as the ability to use these URIs to bookmark molecules or to\n<a href=\"http://chem-bla-ics.blogspot.com/2007/09/tagging-molecules-mashup-of-connotea.html\">annotate them with tags with Connotea</a>\n(which resulted in a nice <a href=\"http://chem-bla-ics.blogspot.com/2007/10/lunch-at-nature-hq-with-euan-joanna-ian.html\">lunch with the Nature people at the time</a>).\nThe link to Connotea is disabled at the moment, though.</p>\n<p>At this moment the system still holds, though there is problem in that browsers can put practical limits on\nURIs length, which limits the maximum size of the InChI. Virtuoso does this too.</p>","doi":"https://doi.org/10.59350/eg04s-efd96","guid":"https://doi.org/10.59350/eg04s-efd96","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1334448000,"reference":[{"id":"https://doi.org/10.1038/npre.2010.4918.1","unstructured":"Unknown title"},{"id":"https://doi.org/10.1002/9781118026038.ch24","unstructured":"Unknown title"},{"id":"https://doi.org/10.1186/2041-1480-2-s1-s6","unstructured":"Unknown title"}],"rid":"kpw7g-39w56","summary":"About four and a half years ago, I started OpenMolecules RDF, a spin off from Chemical blogspace (Cb, which is still up and running thanks to Peter Maas!) where I started using InChIs in URIs . My interest came from the dereferencability, the ability to take an InChI and find information about the chemical structure representated by it. Because information about anything is scattered around the internet, and we need something decentralized.","tags":["Chemistry","Rdf","Inchi","Opendata"],"title":"Dereferencable InChIs: OpenMolecules RDF","updated_at":1785871013,"url":"https://chem-bla-ics.linkedchemistry.info/2012/04/15/dereferencable-inchis-openmolecules-rdf.html","version":"v1"},{"authors":[{"affiliation":[{"id":"https://ror.org/02jz4aj89","name":"Maastricht University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>A good number of years ago, a colleague and I explored if we could get access to the <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/02/16/retraction-data-in-wikidata.html/retractiondatabase.org/\">Retraction Watch Database</a>,\nbut we could not afford it. We have been using data on retractions for curate our databases, like\n<a href=\"https://www.wikipathways.org/\">WikiPathways</a>. A database should not contain knowledge based on (only) a retracted article.\nWikidata, btw, has a small number (499) of statements supported by retracted articles. Similarly, it turns out that I am\n<a href=\"https://w.wiki/8pwe\">citing retracted articles in two papers</a> (and a preprint of one of them).</p>\n<p><a href=\"https://www.wikidata.org/\">Wikidata</a> has a good number of retracted articles in their database\n(<a href=\"https://scholia.toolforge.org/statistics\">some 21 thousand at the time of writing</a>). A lot of this data\ncomes from CrossRef, that recently <a href=\"https://www.crossref.org/blog/news-crossref-and-retraction-watch/\">acquired the Retraction Watch Database</a>\n(doi:<a href=\"https://doi.org/10.13003/c23rw1d9\">10.13003/c23rw1d9</a>)) and started providing the content as FAIR and Open data.\nWith <a href=\"https://github.com/egonw/ons-wikidata/blob/main/RetractionWatch/quickstatements.groovy\">a Bacting-based script</a>\nI am regularly updating Wikidata with annotations from CrossRef, giving a rich dataset in Wikidata around\nthe queries. Over the past few years I have written various SPARQL queries to show the results which today\nI <a href=\"https://bigcat-um.github.io/sparql-examples/examples/WikidataRetractions/\">collected under a single home</a>:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/retraction_SPARQL.png\"/></p>","doi":"https://doi.org/10.59350/w4zj3-mbw53","guid":"https://doi.org/10.59350/w4zj3-mbw53","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1739664000,"reference":[{"id":"https://doi.org/10.1093/nar/gkad960","unstructured":"Agrawal, A., Balc\u0131, H., Hanspers, K., Coort, S. L., Martens, M., Slenter, D. N., Ehrhart, F., Digles, D., Waagmeester, A., Wassink, I., Abbassi-Daloii, T., Lopes, E. N., Iyer, A., Acosta, J. M., Willighagen, L. G., Nishida, K., Riutta, A., Basaric, H., Evelo, C. T., \u2026 Pico, A. R. (2023). WikiPathways 2024: next generation pathway database. <i>Nucleic Acids Research</i>, <i>52</i>(D1), D679\u2013D689."},{"id":"https://doi.org/10.13003/c23rw1d9","unstructured":"Crossref, Hendricks, G., Center for Scientific Integrity&amp; Lammey, R. (2023). <i>Crossref acquires Retraction Watch data and opens it for the scientific community</i>. Crossref.  <b>[cito:citesAsEvidence]</b>"}],"rid":"e8vfg-wqz89","summary":"A good number of years ago, a colleague and I explored if we could get access to the Retraction Watch Database, but we could not afford it. We have been using data on retractions for curate our databases, like WikiPathways. A database should not contain knowledge based on (only) a retracted article. Wikidata, btw, has a small number (499) of statements supported by retracted articles.","tags":["Wikidata","Wikipathways"],"title":"Retracted articles in Wikidata","updated_at":1785870031,"url":"https://chem-bla-ics.linkedchemistry.info/2025/02/16/retraction-data-in-wikidata.html","version":"v1"},{"authors":[{"affiliation":[{"id":"https://ror.org/02jz4aj89","name":"Maastricht University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>Names of chemicals are part of the human user experience when browsing a chemical database. And literature too,\nof course. Chemical names are also not easy to use, and what a chemical name means is not always clear.\nThis is why the <a href=\"https://en.wikipedia.org/wiki/International_Union_of_Pure_and_Applied_Chemistry\">IUPAC</a>\nstarted a standardizing nomenclature in chemistry, the <em>IUPAC names</em>. Each IUPAC name uniquely defines\nthe chemical structure it defines. For example, <em>methane</em> is the IUPAC name for the chemical CH<sub>4</sub>.</p>\n<p>So, when propagating chemical structures from the <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/02/13/beiltein-journal-has-bioschemas.html\">Beilstein Bioschemas feed</a>,\nI was looking for names, IUPAC or not, ideally the name used in the article. When I asked about this,\nthe question came up if they could autogenerate IUPAC names, for which\n<a href=\"https://doi.org/10.1038/s41598-021-94082-y\">various</a>\n<a href=\"https://doi.org/10.1186/s13321-021-00535-x\">new</a>\n<a href=\"https://doi.org/10.1186/s13321-021-00512-4\">tools</a>\n<a href=\"https://doi.org/10.1186/s13321-024-00941-x\">exist</a>\n(I think I am missing one from an American team, but cannot find the reference),\nalong with multiple established commerical tools.\nBecause the IUPAC nomenclature is a long list of naming rules, priorities, etc, a rule-based\nalgorithm is logical, but newer methods take a deep-learning approach.</p>\n<p>Back to the chemical annotation of chemistry literature. This is of obvious interest: you want\nto know where we can read more about a certain chemical. We need the chemical structures in\na database for that, linked to the articles. This is, of course, one of the original studies\nof <em>cheminformatics</em>. And when authors of the chemical literature do not provide this routinely\n(<a href=\"https://chem-bla-ics.linkedchemistry.info/2025/02/13/beiltein-journal-has-bioschemas.html\">this post</a>\nshows a few exceptions, but it is still all too rare). And then manual and automated curation\nis needed, e.g. done by <a href=\"https://en.wikipedia.org/wiki/Chemical_Abstracts_Service\">Chemical Abstracts</a>.</p>\n<p>Third, <a href=\"https://wikidata.org/\">Wikidata</a> has <a href=\"https://scholia.toolforge.org/chemical/\">about 1.4 million</a>\nchemical compounds and many names. A <a href=\"https://www.wikidata.org/wiki/Wikidata:Property_proposal/Pending#IUPAC_name\">property propoal for IUPAC names</a>\nhas been long pending, but once accepted in one form or another, will require IUPAC names too.</p>\n<h2 id=\"one-million-iupac-names\">One million IUPAC names</h2>\n<p>Thus, the idea came up, can we create a set of 1 million unique IUPAC names found in literature?\nI asked on the <a href=\"https://elixir-europe.org/\">ELIXIR Europe</a> slack channel if <a href=\"https://europepmc.org/\">Europe PMC</a>\nhad such a dataset (doi:<a href=\"https://doi.org/10.1093/nar/gkad1085\">10.1093/nar/gkad1085</a>). I knew they had been adding chemical\n<a href=\"https://scholia.toolforge.org/topic/Q403574\">named-entity recognition</a> (NER) results in\n<a href=\"https://europepmc.org/Annotations\">their annotation API</a>. I learned they used <a href=\"https://www.ebi.ac.uk/chebi/\">ChEBI</a>.\nMelanie Vollmar and Summer Rosonovski or Europe PMC gave useful information and support.\n<a href=\"https://cpm.lumc.nl/research/bioinformatics-224/magnus-palmblad-5\">Magnus Palmblad</a> also replied\nand provided Python code to use the Europe PMC API to fetch names it returns and see if those\nare IUPAC names. Well, that's easy. We have <a href=\"https://opsin.ch.cam.ac.uk/\">OPSIN</a> for that\n(see doi:<a href=\"https://doi.org/10.1021/ci100384d\">10.1021/ci100384d</a>).</p>\n<p>Unfortunately, the Europe PMC NER results are not ideal for IUPAC names. Just scanning\nsome 5, 6 organic chemistry journals returned some 8 thousand IUPAC names in open access\narticles. But it quickly started to be too limited: each set of articles returned\nincreasingly few new names. The reason is simple: the NER is too <em>greedy</em> and as a\nresult, does not easily recognize longer IUPAC names. It is too happy with a substring\nof the IUPAC name. For example, when it encounters the IUPAC name <em>5-Bromo-1H-indole-3-carboxylic acid</em>,\nit settles for <em>indole-3-carboxylic acid</em>:</p>\n<p><img alt=\"\" src=\"https://chem-bla-ics.linkedchemistry.info/assets/images/greedy.png\"/></p>\n<h2 id=\"open-source-chemistry-analysis-routines\">Open-Source Chemistry Analysis Routines</h2>\n<p>During my PhD, in 2003, when I worked a few months with Prof. <a href=\"https://scholia.toolforge.org/author/Q908710\">Peter Murray-Rust</a> (University of Cambridge)\nand Prof. Janet Thornthon (EMBL-EBI), I learned about the research by <a href=\"https://scholia.toolforge.org/author/Q28946549\">Sam Adams</a>\n(doi:<a href=\"https://doi.org/10.1039/B411699M\">10.1039/B411699M</a>), <a href=\"https://scholia.toolforge.org/author/Q133040220\">Joe Townsend</a>\n(doi:<a href=\"https://doi.org/10.1039/B411033A\">10.1039/B411033A</a>), and <a href=\"https://scholia.toolforge.org/author/Q90318722\">Peter Corbett</a>\n(doi:<a href=\"https://doi.org/10.1007/11875741_11\">10.1007/11875741_11</a>). One of the tools that used\nthis research was (is) <a href=\"https://scholia.toolforge.org/topic/Q133037490\">OSCAR</a>,\nshort for <em>Open-Source Chemistry Analysis Routines</em> (see <a href=\"https://blogs.ch.cam.ac.uk/pmr/2009/05/16/opsin-and-oscar-chemical-language-processing/\">this detailed write up by Peter MR</a>).\nLater, in 2010 I visted Peter again, as postdoc, in Cambridge, and then\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2010/10/15/working-on-oscar-for-three-months.html\">worked on the OSCAR project</a> too.\nAnd while OSCAR did a lot more, the integration of <a href=\"https://chem-bla-ics.linkedchemistry.info/2010/12/26/oscar-training-data-models-etc.html\">Corbett's NER research</a>\nmade OSCAR the obvious follow-up step in finding IUPAC names in literature.</p>\n<p>And because <a href=\"https://chem-bla-ics.linkedchemistry.info/2011/09/27/almost-year-ago-i-started-position-with.html\">OSCAR4 had been integrated into Bioclipse</a>\n(doi:<a href=\"https://doi.org/10.1186/1758-2946-3-41\">10.1186/1758-2946-3-41</a>) and I had this ported to Bacting already\n(doi:<a href=\"https://doi.org/10.21105/joss.02558\">10.21105/joss.02558</a>), using this was trivial.\nThe use of Europe PMC is different now, however, and we are no longer using the Annotations API,\nbut just using it to find open access articles, and to get the full text in XML format.\nThat allows a simple XPath search on <code class=\"language-plaintext highlighter-rouge\">&lt;p&gt;</code> elements, pass the resulting string to OSCAR4,\nand the recognized names are checked with OPSIN.\nAnd with this approach, processing two of the five or six journals we earlier explored,\nwe find another 40+ thousand IUPAC names. Quite a success, I am tempted to say.</p>\n<h2 id=\"a-blue-obelisk-project\">A Blue Obelisk project</h2>\n<p>So, I started a new <a href=\"https://blueobelisk.github.io/\">Blue Obelisk</a> project,\n<a href=\"https://github.com/BlueObelisk/iupac-names\">iupac-names</a>, to collect 1M IUPAC names. For researchers\nto use, learn from, etc. Just IUPAC names. Not even the chemical structure, nor the link to the\narticles. The first is trivial to do with OPSIN, so the matching SMILES do not need to be stored.\nLinks to literature is tricky because of the aforementioned issues, and we only want to know\nwhich (partial) IUPAC names occur in literature. If you really want to know in which articles\nthat IUPAC name is found, you can simply do a search in Europe PMC.</p>\n<p>And because we only store IUPAC names, this are very basic facts (this is an IUPAC name, as defined\nby OPSIN being able to generate a SMILES for this structure) and that that string occurs in\nsome article) and we can share them as CCZero. We <a href=\"https://chem-bla-ics.linkedchemistry.info/2025/03/08/iupac-names.html/is:issue\" title=\"milestone release\">defined various milestones</a>,\nand I am happy that the first two have been reached within two weeks:</p>\n<ul>\n<li><a href=\"https://github.com/BlueObelisk/iupac-names/releases/tag/milestone-10k\">Milestone 10k</a> (doi:<a href=\"https://doi.org/10.5281/zenodo.14965762\">10.5281/zenodo.14965762</a>)</li>\n<li><a href=\"https://github.com/BlueObelisk/iupac-names/releases/tag/milestone-50k\">Milestone 50k</a> (doi:<a href=\"https://doi.org/10.5281/zenodo.14978557\">10.5281/zenodo.14978557</a>)</li>\n</ul>\n<p>This second milestone has 53848 unique names, but as literature goes, there are interesting\nvariations, some likely because of typesetting leading to spaces added and missing. If\nwe ignore spaces and hyphens, we have 50534 names left (hence the milestone). But IUPAC\nnames are also not fully unique, partly because of Unicode character variations and greek\nletter alternatives, and you may wonder how many different chemical structures this set\nreflects. While not perfect, the Standard InChI gives some lower limit, and we find 36528\nInChIKeys in this second milestone.</p>\n<p>Now, we need twenty times as much to reach the 1M IUPAC names, but given we have many, many\nmore open access articles to process. The bottleneck seems to be mostly our workflow.</p>\n<h3 id=\"can-you-contribute\">Can you contribute?</h3>\n<p>Yes, of course! This is an open science project. But please keep in mind the narrow focus of this\nproject: only IUPAC names which can be found in (open access) literature. This project doed not accept\nautogenerated names (PubChem would have given use many millions already), nor IUPAC names from existing\ndatabases. Ideally, you are able to show the code you use to extract/find those names in literature.</p>\n<h3 id=\"can-i-use-these-names\">Can I use these names?</h3>\n<p>First of all, this is what the CCZero license and open science nature of this project is about: reuse.\nWe love to hear how you are using these names, tho, and we encourage you to write up how you\nare using them. You can use <a href=\"https://datacite.org/\">DataCite</a> to cite the release you used,\nand citing this blog post by DOI is also possible.</p>\n<h3 id=\"does-it-support-my-language-too\">Does it support my language too?</h3>\n<p>No, at this moment it only support IUPAC names in English. Dutch, French, Spanish, or Chinese\nIUPAC names are valid, but currently not supported. See also\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2010/12/30/text-mining-chemistry-from-dutch-or.html\">this post</a>.</p>\n<h3 id=\"will-there-be-a-publication\">Will there be a publication?</h3>\n<p>Magnus and I intend so. We already submitted an abstract to the <a href=\"https://iccs-nl.org/\">International Conference on Chemical Structures</a>,\nwhich has <a href=\"https://www.biomedcentral.com/collections/ICCS25\">a Collection in the Journal of Cheminformatics</a>.\nIf the abstract gets accepted, of course, we can submit there. Otherwise, we will look for another venue,\nlikely <a href=\"https://en.wikipedia.org/wiki/Diamond_open_access\">diamond open access</a>.</p>\n<h3 id=\"where-is-your-script\">Where is your script?</h3>\n<p>Ah, fair point. We did not decide on the final license yet. I have used two scripts based on the template\nby Magnus. As soon as we have finalized the license, we will make those available.</p>","doi":"https://doi.org/10.59350/tjkf2-k1608","guid":"https://doi.org/10.59350/tjkf2-k1608","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1741392000,"reference":[{"id":"https://doi.org/10.1038/s41598-021-94082-y","unstructured":"Krasnov, L., Khokhlov, I., Fedorov, M. V.&amp; Sosnin, S. (2021). Transformer-based artificial neural networks for the conversion between chemical notations. <i>Scientific Reports</i>, <i>11</i>(1)."},{"id":"https://doi.org/10.1186/s13321-021-00512-4","unstructured":"Rajan, K., Zielesny, A.&amp; Steinbeck, C. (2021). STOUT: SMILES to IUPAC names using neural machine translation. <i>Journal of Cheminformatics</i>, <i>13</i>(1)."},{"id":"https://doi.org/10.1186/s13321-021-00535-x","unstructured":"Handsel, J., Matthews, B., Knight, N. J.&amp; Coles, S. J. (2021). Translating the InChI: adapting neural machine translation to predict IUPAC names from a chemical identifier. <i>Journal of Cheminformatics</i>, <i>13</i>(1)."},{"id":"https://doi.org/10.1186/s13321-024-00941-x","unstructured":"Rajan, K., Zielesny, A.&amp; Steinbeck, C. (2024). STOUT V2.0: SMILES to IUPAC name conversion using transformer models. <i>Journal of Cheminformatics</i>, <i>16</i>(1)."},{"id":"https://doi.org/10.1021/ci100384d","unstructured":"Lowe, D. M., Corbett, P. T., Murray-Rust, P.&amp; Glen, R. C. (2011). Chemical Name to Structure: OPSIN, an Open Source Solution. <i>Journal of Chemical Information and Modeling</i>, <i>51</i>(3), 739\u2013753."},{"id":"https://doi.org/10.1039/b411699m","unstructured":"Adams, S. E., Goodman, J. M., Kidd, R. J., McNaught, A. D., Murray-Rust, P., Norton, F. R., Townsend, J. A.&amp; Waudby, C. A. (2004). Experimental data checker: better information for organic chemists. <i>Organic & Biomolecular Chemistry</i>, <i>2</i>(21), 3067."},{"id":"https://doi.org/10.1039/b411033a","unstructured":"Townsend, J. A., Adams, S. E., Waudby, C. A., de Souza, V. K., Goodman, J. M.&amp; Murray-Rust, P. (2004). Chemical documents: machine understanding and automated information extraction. <i>Organic & Biomolecular Chemistry</i>, <i>2</i>(22), 3294."},{"id":"https://doi.org/10.1007/11875741_11","unstructured":"Corbett, P.&amp; Murray-Rust, P. (2006). High-Throughput Identification of Chemistry in Life Science Texts. In <i>Lecture Notes in Computer Science</i> (pp. 107\u2013118). Springer Berlin Heidelberg."},{"id":"https://doi.org/10.1186/1758-2946-3-41","unstructured":"Jessop, D. M., Adams, S. E., Willighagen, E. L., Hawizy, L.&amp; Murray-Rust, P. (2011). OSCAR4: a flexible architecture for chemical text-mining. <i>Journal of Cheminformatics</i>, <i>3</i>(1).  <b>[cito:usesMethodIn]</b>"},{"id":"https://doi.org/10.21105/joss.02558","unstructured":"Willighagen, E. (2021). Bacting: a next generation, command line version of Bioclipse. <i>Journal of Open Source Software</i>, <i>6</i>(62), 2558.  <b>[cito:usesMethodIn]</b>"},{"id":"https://doi.org/10.1093/nar/gkad1085","unstructured":"Rosonovski, S., Levchenko, M., Bhatnagar, R., Chandrasekaran, U., Faulk, L., Hassan, I., Jeffryes, M., Mubashar, S. I., Nassar, M., Jayaprabha\u00a0Palanisamy, M., Parkin, M., Poluru, J., Rogers, F., Saha, S., Selim, M., Shafique, Z., Ide-Smith, M., Stephenson, D., Tirunagari, S., \u2026 Harrison, M. (2023). Europe PMC in 2023. <i>Nucleic Acids Research</i>, <i>52</i>(D1), D1668\u2013D1676.  <b>[cito:usesMethodIn]</b>"},{"id":"https://doi.org/10.5281/zenodo.14965762","unstructured":"Egon Willighagen. (2025). <i>BlueObelisk/iupac-names: Milestone 10k</i> (Version milestone-10k) [Dataset]. Zenodo.  <b>[cito:citesAsEvidence]</b>"},{"id":"https://doi.org/10.5281/zenodo.14978557","unstructured":"Egon Willighagen. (2025). <i>BlueObelisk/iupac-names: Milestone 50k</i> (Version milestone-50k) [Dataset]. Zenodo.  <b>[cito:citesAsEvidence]</b>"}],"rid":"a2g45-bjb85","summary":"Names of chemicals are part of the human user experience when browsing a chemical database. And literature too, of course. Chemical names are also not easy to use, and what a chemical name means is not always clear. This is why the IUPAC started a standardizing nomenclature in chemistry, the IUPAC names. Each IUPAC name uniquely defines the chemical structure it defines. For example, methane is the IUPAC name for the chemical CH4.","tags":["Iupac","Cheminf","Oscar","Textmining","Europepmc"],"title":"One Million IUPAC names","updated_at":1785870029,"url":"https://chem-bla-ics.linkedchemistry.info/2025/03/08/iupac-names.html","version":"v1"},{"authors":[{"affiliation":[{"id":"https://ror.org/02jz4aj89","name":"Maastricht University"}],"contributor_roles":[],"family":"Willighagen","given":"Egon","url":"https://orcid.org/0000-0001-7542-0286"}],"blog":{"authors":[{"name":"Egon Willighagen"}],"community_id":"7f57028e-9d03-489c-b3b4-3d60de06bc9e","created":1710288000,"current_feed_url":"https://chem-bla-ics.linkedchemistry.info/feed.json","description":"Chemblaics (pronounced chem-bla-ics) is the science that uses open science and computers to solve problems in chemistry, biochemistry and related fields.","doi":"https://doi.org/10.59350/chem_bla_ics","favicon":"https://rogue-scholar.org/api/communities/7f57028e-9d03-489c-b3b4-3d60de06bc9e/logo","feed_format":"application/feed+json","feed_url":"https://chem-bla-ics.linkedchemistry.info/archive.json","filter":null,"generator":"Jekyll","home_page_url":"https://chem-bla-ics.linkedchemistry.info","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"chem_bla_ics","status":"active","subfield":"1606","title":"chem-bla-ics","updated":1785628800,"use_api":true},"blog_name":"chem-bla-ics","blog_slug":"chem_bla_ics","content_html":"<p>As part of our <a href=\"https://www.nwo.nl/en/\">Dutch Research Council</a> (NWO) <a href=\"https://www.nwo.nl/en/projects/osf232097\">Open Science grant</a>,\nwe organized a <a href=\"https://cdk.github.io/nwo-openscience-2024/\">Chemistry Development Kit User Group Meeting</a>\n(<a href=\"https://hashtags-hub.toolforge.org/CDK25UGM\">#CDK25UGM</a>), of which yesterday was the \"conference\" day, and today a hackathon.</p>\n<p>I opened the session with a few slides welcoming everyone at Maastricht University (and our\n<a href=\"https://chem-bla-ics.linkedchemistry.info/2025/01/27/translational-genomics.html\">Dept of Translational Genomics</a>,\nand explaining the NWO grant.\n<a href=\"https://orcid.org/0000-0001-7730-2646\">John Mayfield</a> (<a href=\"https://www.nextmovesoftware.com/\">NextMove</a>) spoke about\n\"What's New\" in the Chemistry Development Kit 2.10, e.g. explaining more about the new (much faster) <code class=\"language-plaintext highlighter-rouge\">AtomContainer</code>,\nSMIRKS, and more.</p>\n<p>After lunch, <a href=\"https://orcid.org/0000-0003-1554-6666\">Jonas Schaub</a> (<a href=\"https://www.uni-jena.de/en/\">Friedrich Schiller University Jena</a>)\nshowed various projects where the CDK is used, titled  \"Scaffolds, Functional Groups, Aglycones: Algorithmic Substructure Identification with CDK\"\n(see doi:<a href=\"https://doi.org/10.1186/s13321-023-00762-4\">10.1186/s13321-023-00762-4</a>, doi:<a href=\"https://doi.org/10.1186/s13321-022-00656-x\">10.1186/s13321-022-00656-x</a>,\nand doi:<a href=\"https://doi.org/10.1186/s13321-020-00467-y\">10.1186/s13321-020-00467-y</a>).\nLyudvika Radeva (<a href=\"https://www.ideaconsult.net/\">Ideaconsult Ltd</a>, <a href=\"https://uni-plovdiv.bg/en/\">University of Plovdiv</a>) showed\nwhat SYBYL Line Notation (SLN) is and how this is implemented in Ambit (see doi:<a href=\"https://doi.org/10.1002/minf.202100027\">10.1002/minf.202100027</a>).\n<a href=\"https://orcid.org/0000-0002-4354-4353\">Sonja Herres-Pawlis</a> (<a href=\"https://www.rwth-aachen.de/\">RWTH Aachen University</a>)\nupdated us with \"News from the InChI: making the InChI FAIR and including inorganics\", e.g. showing how\nthey worked out how the InChI is going to handle organometalics, where the bonds and the stereochemistry\nas aspects that were not handled by the current InChI.</p>\n<p>After the afternoon coffee break, <a href=\"https://orcid.org/0000-0003-3662-2621\">Zhixu Ni</a> (<a href=\"https://fedorovalab.net/team/zhixu-ni/\">TU Dresden</a>)\nshowed his work on lipid maps characterization and identification. We previously met a few times\nat EpiLipidNET COST action meetings, and it was great to see his continued research on representation\nof lipids and lipid classes in hit \"A Fuzzy Solution for Lipid Structures Using CXSMILES\".</p>\n<p>Finally, <a href=\"https://www.linkedin.com/in/matthiasmailaender/\">Matthias Mail\u00e4nder</a> (<a href=\"https://www.lablicate.com/\">Lablicate GmbH</a>)\ngave a \"Live demo of where <a href=\"https://github.com/OpenChrom\">OpenChrom</a> uses the CDK\", and\n<a href=\"https://research.rug.nl/en/persons/yajie-ding\">Yajie Ding</a> (University of Groningen) told her about her\nglycoscience research. There, cheminformatics can also greatly help and the CDK may provide\nthem with solutions.</p>\n<p>This really doesn't do justice to all the discussions, examples, use cases, etc. But it gives you\nan idea. We had 11 people in the room, and were joined online by an additional 6 people.</p>","doi":"https://doi.org/10.59350/e08pe-thb38","funding_references":[{"awardNumber":"osf232097","awardTitle":"The Chemistry Development Kit in 2024: improving cheminformatics research","awardUri":"https://www.nwo.nl/en/projects/osf232097","funderIdentifier":"https://ror.org/04jsz6e67","funderIdentifierType":"ROR","funderName":"Dutch Research Council"}],"guid":"https://doi.org/10.59350/e08pe-thb38","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1741651200,"reference":[{"id":"https://doi.org/10.1002/minf.202100027","unstructured":"Kochev, N., Jeliazkova, N.&amp; Tancheva, G. (2021). Ambit\u2010SLN: an Open Source Software Library for Processing of Chemical Objects via SLN Linear Notation. <i>Molecular Informatics</i>, <i>40</i>(11)."},{"id":"https://doi.org/10.1186/s13321-022-00656-x","unstructured":"Schaub, J., Zander, J., Zielesny, A.&amp; Steinbeck, C. (2022). Scaffold Generator: a Java library implementing molecular scaffold functionalities in the Chemistry Development Kit (CDK). <i>Journal of Cheminformatics</i>, <i>14</i>(1)."},{"id":"https://doi.org/10.1186/s13321-023-00762-4","unstructured":"Chandrasekhar, V., Sharma, N., Schaub, J., Steinbeck, C.&amp; Rajan, K. (2023). Cheminformatics Microservice: unifying access to open cheminformatics toolkits. <i>Journal of Cheminformatics</i>, <i>15</i>(1)."},{"id":"https://doi.org/10.1186/s13321-020-00467-y","unstructured":"Schaub, J., Zielesny, A., Steinbeck, C.&amp; Sorokina, M. (2020). Too sweet: cheminformatics for deglycosylation in natural products. <i>Journal of Cheminformatics</i>, <i>12</i>(1)."}],"rid":"smhnn-nr982","summary":"As part of our Dutch Research Council (NWO) Open Science grant, we organized a Chemistry Development Kit User Group Meeting (#CDK25UGM), of which yesterday was the \"conference\" day, and today a hackathon.","tags":["Cdk","Openscience","Cdk2024"],"title":"cdk2024 #4: Chemistry Development Kit User Group Meeting - Day 1","updated_at":1785870028,"url":"https://chem-bla-ics.linkedchemistry.info/2025/03/11/CDK-UGM.html","version":"v1"}],"out_of":53415,"page":1,"per_page":10,"total-results":53415}
