{"found":57638,"hits":[{"document":{"authors":[{"contributor_roles":[],"family":"Eden","given":"Terence","url":"https://orcid.org/0000-0002-9265-9069"}],"blog":{"authors":null,"community_id":"61ce553a-bafd-4aba-a952-d3bab5e85bcc","created":1788652800,"current_feed_url":null,"description":"Regular nonsense about tech and its effects \ud83d\ude43","doi":"https://doi.org/10.59350/shkspr","favicon":"https://rogue-scholar.org/api/communities/61ce553a-bafd-4aba-a952-d3bab5e85bcc/logo","feed_format":"application/atom+xml","feed_url":"https://shkspr.mobi/blog/feed/DOI","filter":"category:-1982","generator":"WordPress","home_page_url":"https://shkspr.mobi/blog","issn":"2753-1570","language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"shkspr","status":"active","subfield":"1712","title":"Terence Eden's Blog","updated":1790106552,"use_api":true},"blog_name":"Terence Eden's Blog","blog_slug":"shkspr","content_html":"<img src=\"https://shkspr.mobi/blog/wp-content/uploads/2026/09/how-to-build-a-space-station.webp\" alt=\"Book cover featuring astronauts on Mars looking at a habitat.\" width=\"256\" height=\"384\" class=\"alignleft\">\n\n<p>This book, by The Times' former Architecture Correspondent, stands in direct opposition to <a href=\"https://shkspr.mobi/blog/2026/07/book-review-a-city-on-mars-by-dr-kelly-weinersmith-and-zach-weinersmith/\">A City on Mars</a>. Whereas that book presented a (perhaps too) sceptical look at the realities of living on other planets, Jonathan Morrison's book goes (perhaps too far) in the opposite direction.</p>\n\n<!-- \n\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\nWelcome to CommentClub! This content is only available to people who read my HTML comments. Why would you do that? What are you hoping to learn? So, the three rules of comment club are\u2026\n\n1. You must not tell anyone about Comment Club - let them find out about it themselves.\n2. You *must* add some Comment Club Content to at least one of your blog posts. Doesn't have to be much, just a little note will do.\n3. You should drop an email to anyone whose comments you've read. Start by sending me a friendly message at comment.club@shkspr.mobi\n\nKeep your eyes peeled for more comments in future blog posts \ud83d\ude03\n\nTTFN.\n-->\n\n<p>How to Build a Space Station is a beautiful examination of just how important architecture will be to our colonisation of other worlds. Not just in terms of physical safety - but psychological safety as well. It is a direct and forceful rebuttal to those who say it cannot be done.</p>\n\n<p>It is, in my opinion, just a touch too credulous about some of the ludicrous claims from the hype merchants. I want to believe that Martian igloos can be conjured out of the ice and that Musk's rockets will deliver a steady stream of supplies to distant worlds. But the evidence presented is rather thin. The book works best when it focuses on what architecture can bring to the table when it comes to designing the future.</p>\n\n<blockquote><p>In short, a spacecraft is not just a machine; it is also a home, an office, a refuge. If people are asked to go to the most remote environments, to live in spaces scarcely larger than a few rooms, and to perform work of immense complexity and risk, then comfort, efficiency and ergonomics are not just luxuries.</p></blockquote>\n\n<p>Space has to be <em>worth</em> living in. Putting people into a tin-can with no windows, blank walls, and an infernal background hum will drive them mad. All this is backed up with extensive descriptions of the engineering challenges of polar research bases, spaceports, and previous craft.</p>\n\n<p>Despite being rightly scathing about Wernher von Braun's involvement in atrocities and his eventual political rehabilitation - he is somewhat more muted in his criticism of Messrs Musk &amp; Bezos. There's a <em>lot</em> of praise for celebrity architects and designers - without any real examination of whether their designs are practical rather than just award fodder.</p>\n\n<p>Similarly, the book takes on trust that autonomous robots <em>can</em> ingest extraterrestrial soil, process it, and 3D print structures from it all while in a hostile environment. The fact that we don't have swarms of drones prefabbing houses in the relatively benign atmosphere of our planet should be evidence that maybe these claims aren't quite matched with reality.</p>\n\n<p>Finally, the \"why?\" question. A City on Mars points out that the cost of mining gold from asteroids would be more profitably spent improving mining technology here on Earth. How to Build a Space Station takes a different approach; it'll improve things here:</p>\n\n<blockquote><p>Space architecture is not just escapism, a thrilling sci-\u00adfi fantasy \u2013 it is a forge for creating the tools we need at home. These include but are not limited to circular systems, low-\u00adenergy fabrication, modular construction and buildings that take psychology seriously.</p></blockquote>\n\n<p>I have a lot of sympathy for that. Except\u2026 the Internation Space Station has shown us how to endlessly recycle water relatively cheaply. Yet every modern building on Earth pays only lip-service to reusing grey-water. 3D printing is amazing, but the number of structures built using autonomous robots extruding concrete is approximately zero.</p>\n\n<p>We have the technology - but we don't seem to be interested in using it.</p>\n\n<p>The book is mostly well illustrated - with some gorgeous drawings of actual craft and possible future inventions. Sadly no photos, maps, or anything to help illuminate some of the other challenges faced by living and working in space.</p>\n\n<p>This book is endlessly fascinating and bang up to date, with lots of talk of events that happened in 2025. The way it brings together the sciences of engineering and psychology is marvellous.  But, as much as I'd like to believe in a Martian habitat built by robot trebuchets flinging microwave sintered tetrapods into each other, I just don't find it convincing.</p>\n\n<p>I <em>really</em> hope I'm wrong.</p>\n\n<p>Many thanks to Netgalley for the review copy - the book is available to buy now.</p><img src=\"https://shkspr.mobi/blog/wp-content/themes/edent-wordpress-theme/info/okgo.php?ID=75653&HTTP_REFERER=DOI\" alt width=1 height=1 loading=eager>","doi":"https://doi.org/10.59350/0vqd8-3qy21","guid":"https://shkspr.mobi/blog/?p=75653","image":"https://shkspr.mobi/blog/wp-content/uploads/2026/09/how-to-build-a-space-station.webp","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1789862400,"rid":"8s9qs-35530","summary":"This book, by The Times' former Architecture Correspondent, stands in direct opposition to A City on Mars. Whereas that book presented a (perhaps too) sceptical look at the realities of living on other planets, Jonathan Morrison's book goes (perhaps too far) in the opposite direction.","tags":["/etc/","Book Review","NetGalley"],"title":"Book Review: How to Build a Space Station by Jonathan Morrison","updated_at":1790106912,"url":"https://shkspr.mobi/blog/2026/09/book-review-how-to-build-a-space-station-by-jonathan-morrison/","version":"v1"}},{"document":{"authors":[{"affiliation":[{"name":"Front Matter"}],"contributor_roles":[],"family":"Fenner","given":"Martin","url":"https://orcid.org/0000-0003-1419-2405"}],"blog":{"authors":[{"name":"Martin Fenner","url":"https://orcid.org/0000-0003-1419-2405"}],"community_id":"15a362ea-8138-42b8-917f-1840a92addf8","created":1672531200,"current_feed_url":null,"description":"The Front Matter Blog covers the intersection of science and technology since 2007.","doi":"https://doi.org/10.53731/front_matter","favicon":"https://rogue-scholar.org/api/communities/15a362ea-8138-42b8-917f-1840a92addf8/logo","feed_format":"application/atom+xml","feed_url":"https://blog.front-matter.de/atom","filter":null,"generator":"Ghost","home_page_url":"https://blog.front-matter.de/","issn":"2749-9952","language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.53731","relative_url":null,"secure":true,"slug":"front_matter","status":"active","subfield":"1710","title":"Front Matter","updated":1790088325,"use_api":true},"blog_name":"Front Matter","blog_slug":"front_matter","content_html":"<p>The science blog archive Rogue Scholar is launching an Editorial Board to handle submissions of new blogs, and oversee editorial management of existing blogs, e.g. blogs or individual posts that should be retracted or no longer archived.</p><p>The Rogue Scholar science blog archive accepts new blogs via a submission form and after acceptance doesn't interfere with what participating blogs publish or how often. If participating blogs stop publishing it doesn't expect a notification. While this adhoc workflow generally worked for the 200 blogs accepted into Rogue Scholar so far, there are some issues that can be improved:</p><ul><li>Processing of submissions of new blogs takes too long and with no clear feedback loop,</li><li>Blog posts written in languages other than German and English are hard to understand for me, and I am more comfortable in some scientific disciplines than others,</li><li>The decision for acceptance of a new blog is ultimately the decision of a single person (me),</li><li>The requirements in the submission form are frequently ignored (e.g. supported language, number of posts, or use of AI),</li><li>The kinds of new blogs we want to see has not been clearly articulated, and no systematic active blog recruitment is happening.</li></ul><p>My hope is that the new Rogue Scholar Editorial Board will address these issues. Starting this week all new submissions will be processed once a month on the 15th, with time for additional questions before the decision if a submission is unclear. The submission form has added a question to describe the blog in a free-text field to provide more context. And I am looking for volunteers to join the Editorial Board, so that the workload and governance can be shared by more people. I hope to have an Editorial Board with 3-5 additional people setup until the end of the year, and I am looking for volunteers comfortable in other languages and scientific disciplines. This Editorial Board will supplement the work of the Rogue Scholar Advisory, and the Board of the non-profit organization once the Rogue Scholar non-profit launches. This is potentially two half days of volunteer work every month, but it is hopefully rewarding work. Until the Editorial Board is in place I will process blog submissions alone at the 15th of every month.</p><p>Please reach out via&nbsp;<a href=\"https://join.slack.com/t/rogue-scholar/shared_invite/zt-2ylpq1yoy-o~TkxDarfz5LSMhGSCYtiA\" rel=\"noreferrer\">Slack</a>,&nbsp;<a href=\"mailto:info@rogue-scholar.org\" rel=\"noreferrer\">email</a>,&nbsp;<a href=\"https://wisskomm.social/@rogue_scholar\" rel=\"noreferrer\">Mastodon</a>, or&nbsp;<a href=\"https://bsky.app/profile/rogue-scholar.bsky.social\" rel=\"noreferrer\">Bluesky</a>&nbsp;if you want to join the new Editorial Board.</p><div class=\"kg-card kg-callout-card kg-callout-card-blue\"><div class=\"kg-callout-text\">Rogue Scholar is a scholarly infrastructure that is free for all authors and readers. You can support Rogue Scholar with a one-time or recurring&nbsp;<a href=\"https://ko-fi.com/rogue_scholar\" rel=\"noreferrer\">donation</a>&nbsp;or by becoming a sponsor.</div></div>","doi":"https://doi.org/10.53731/wtf9v-y0n52","guid":"https://doi.org/10.53731/wtf9v-y0n52","image":"https://images.unsplash.com/photo-1613799591389-0160e094db32?crop=entropy&cs=tinysrgb&fit=max&fm=jpg&ixid=M3wxMTc3M3wwfDF8c2VhcmNofDI4N3x8RWRpdG9yaWFsJTIwYm9hcmR8ZW58MHx8fHwxNzkwMDg1MzY3fDA&ixlib=rb-4.1.0&q=80&w=2000","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1790035200,"rid":"89y6z-bd709","summary":"The science blog archive Rogue Scholar is launching an Editorial Board to handle submissions of new blogs, and oversee editorial management of existing blogs, e.g. blogs or individual posts that should be retracted or no longer archived. The Rogue Scholar science blog archive accepts new blogs via a submission form and after acceptance doesn't interfere with what participating blogs publish or how often.","tags":["Rogue Scholar"],"title":"Rogue Scholar is launching an Editorial Board","updated_at":1790092816,"url":"https://blog.front-matter.de/posts/rogue-scholar-is-launching-an-editorial-board/","version":"v1"}},{"document":{"authors":[{"affiliation":[{"id":"https://ror.org/0153tk833","name":"University of Virginia"}],"contributor_roles":[],"family":"Turner","given":"Stephen","url":"https://orcid.org/0000-0001-9140-9028"}],"blog":{"authors":[{"name":"Stephen Turner"}],"community_id":"382941a7-2ffa-41df-8bbb-5f772188517f","created":1780876800,"current_feed_url":null,"description":"A practicing data scientist's take on AI, genomics, biosecurity, and the ways AI is reshaping how science gets done. Weekly updates from the field. Occasional notes on programming.","doi":"https://doi.org/10.59350/stephenturner","favicon":"https://rogue-scholar.org/api/communities/382941a7-2ffa-41df-8bbb-5f772188517f/logo","feed_format":"application/rss+xml","feed_url":"https://blog.stephenturner.us/feed","filter":null,"generator":"Substack","home_page_url":"https://blog.stephenturner.us","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"stephenturner","status":"active","subfield":"1311","title":"Paired Ends","updated":1790091521,"use_api":true},"blog_name":"Paired Ends","blog_slug":"stephenturner","content_html":"<p>I went on CBS19 TV news here in Charlottesville last night to talk about AI and biosecurity. I thought I was walking into a session that was going to be taped and played over a slow news day. I was pleasantly surprised (and a little thrown off) when I found out the segment was live on the evening news! <a href=\"https://www.cbs19news.com/news/inside-the-numbers-09-21/video_e59405cb-29cd-5b0d-8d85-929e19b2bf36.html\">Here's the recording</a>. </p><p>Biosecurity, and AI safety in general, is a really nuanced topic to try to get right in 240 seconds!</p><div class=\"captioned-image-container\"><figure><a class=\"image-link image2 is-viewable-img\" target=\"_blank\" href=\"https://www.cbs19news.com/news/inside-the-numbers-09-21/video_e59405cb-29cd-5b0d-8d85-929e19b2bf36.html\" data-component-name=\"Image2ToDOM\"><div class=\"image2-inset\"><picture><source type=\"image/webp\" srcset=\"https://substackcdn.com/image/fetch/$s_!0d-_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 424w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 848w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 1456w\" sizes=\"100vw\"><img src=\"https://substackcdn.com/image/fetch/$s_!0d-_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg\" width=\"1184\" height=\"663\" data-attrs=\"{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:663,&quot;width&quot;:1184,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:204149,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:&quot;https://www.cbs19news.com/news/inside-the-numbers-09-21/video_e59405cb-29cd-5b0d-8d85-929e19b2bf36.html&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.stephenturner.us/i/216919562?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" class=\"sizing-normal\" alt=\"\" srcset=\"https://substackcdn.com/image/fetch/$s_!0d-_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 424w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 848w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 1456w\" sizes=\"100vw\" fetchpriority=\"high\"></picture><div class=\"image-link-expand\"><div class=\"pencraft pc-display-flex pc-gap-8 pc-reset\"><button tabindex=\"0\" type=\"button\" class=\"pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M\"><svg aria-hidden=\"true\" width=\"20\" height=\"20\" viewBox=\"0 0 20 20\" fill=\"none\" stroke-width=\"1.5\" stroke=\"var(--color-fg-primary)\" stroke-linecap=\"round\" stroke-linejoin=\"round\" xmlns=\"http://www.w3.org/2000/svg\" class=\"icon-noB79L\"><g><path d=\"M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882\"></path></g></svg></button><button tabindex=\"0\" type=\"button\" class=\"pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M\"><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"20\" height=\"20\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-maximize2 lucide-maximize-2 icon-noB79L\"><polyline points=\"15 3 21 3 21 9\"></polyline><polyline points=\"9 21 3 21 3 15\"></polyline><line x1=\"21\" x2=\"14\" y1=\"3\" y2=\"10\"></line><line x1=\"3\" x2=\"10\" y1=\"21\" y2=\"14\"></line></svg></button></div></div></div></a><figcaption class=\"image-caption\">View the full 4 minute segment on <a href=\"https://www.cbs19news.com/news/inside-the-numbers-09-21/video_e59405cb-29cd-5b0d-8d85-929e19b2bf36.html\">CBS19 News Charlottesville</a>.</figcaption></figure></div><p>Here's what I tried to get across.</p><ol><li><p><strong>The OpenAI / Hugging Face incident was a warning shot.</strong> It's a cybersecurity incident that shows that capable AI systems can persist at hard problems, use tools, escape containment, and hack into systems people thought were secure. In biosecurity, AI can lower the barriers at important steps like finding technical information, planning experiments, troubleshooting failures, and operating tools. We need to test what systems actually do in realistic settings, in a safe way, to really understand the biosecurity risks posed by AI.</p></li><li><p><strong>Be skeptical of doomsday countdowns.</strong> The anchor asked me about doomsday predictions so I advised treating these with skepticism. Some capabilities already here, but creating biothreats still requires intent, expertise, access to materials and equipment. And we here at the University of Virginia School of Data Science as well as many other <a href=\"https://blog.stephenturner.us/p/ai-biosecurity-review\">really talented people and organizations are working tirelessly</a> to make sure what just happened in cyber never happens in bio. </p></li><li><p><strong>We should take the risk seriously without overstating it.</strong> Just because AI knows a lot about biology doesn't mean some rogue AI agent can go out and create bioweapons. And, AI won't turn someone into a biologist overnight: biology still requires materials, laboratory skills, and real-world infrastructure. We need to measure this to understand it. That's one of the things we're trying to understand here: what uplift does AI provide for doing biology if you don't have a PhD in biology?</p></li><li><p><strong>We need layers of protection.</strong> <a href=\"https://www.rand.org/pubs/research_reports/RRA4999-1.html\">Defense in depth</a>, as Steph Guerra and friends at RAND call it. Realistic testing before release, limits and monitoring for high-risk AI use, strong containment for autonomous agents, and safeguards at physical chokepoints like DNA synthesis and lab access. We also need independent evaluation and prompt incident reporting. We can't rely on a single company to get this right for all of us.</p></li><li><p><strong>Pace of AI progress vs AI governance.</strong> AI capabilities are improving fast. Governance needs to keep up. But we shouldn't lock everything down with every possible restriction. We should prioritize measuring capabilities and how we mitigate threats: secure evaluation environments, independent testing, incident reporting, and screening at physical chokepoints. Without good measurement we can't manage risk effectively.</p></li><li><p><strong>It's not either or.</strong> The anchor asked me: \"Using AI to cure cancer would be good. Making bioweapons would be bad. What's the balance?\" Nobody wants to stop AI from helping with cancer, vaccines, or public health. We all want to make beneficial uses easier and dangerous uses harder or impossible. That means targeted safeguards, independent testing, and updating policy when the evidence changes.&nbsp;</p></li><li><p><strong>Measurement matters.</strong> We can't implement calibrated mitigations if we don't know what we're dealing with. Evals tell us a lot, but we're also working to go further: when AI performs well on a benchmark, to what degree does that translate into a person being more capable at doing biology in the real world? </p></li></ol><p class=\"button-wrapper\" data-attrs=\"{&quot;url&quot;:&quot;https://blog.stephenturner.us/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}\" data-component-name=\"ButtonCreateButton\"><a class=\"button primary\" href=\"https://blog.stephenturner.us/subscribe?\"><span>Subscribe now</span></a></p><p></p>","doi":"https://doi.org/10.59350/ar3zs-43b27","guid":"216919562","image":"https://substackcdn.com/image/fetch/$s_!0d-_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1790035200,"rid":"kakmx-0ep33","summary":"I got four minutes on live TV news to take on an important but nuanced topic.","tags":["Biosecurity","AI"],"title":"AI + Biosecurity in 4 minutes","updated_at":1790091756,"url":"https://blog.stephenturner.us/p/ai-biosecurity-tv-interview","version":"v1"}},{"document":{"authors":[{"contributor_roles":[],"family":"Kehm","given":"Nicole"}],"blog":{"authors":null,"community_id":"db0d8909-9e37-46d0-b16c-0551f575e86b","created":1749772800,"current_feed_url":null,"description":"Das Blog der TIB \u2013 Leibniz-Informationszentrum Technik und Naturwissenschaften und Universit\u00e4tsbibliothek","doi":"https://doi.org/10.65527/tib","favicon":"https://rogue-scholar.org/api/communities/db0d8909-9e37-46d0-b16c-0551f575e86b/logo","feed_format":"application/atom+xml","feed_url":"https://blog.tib.eu/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.tib.eu/","issn":null,"language":"deu","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.65527","relative_url":null,"secure":true,"slug":"tib","status":"active","subfield":"1802","title":"TIB-Blog","updated":1790082029,"use_api":true},"blog_name":"TIB-Blog","blog_slug":"tib","content_html":"<p>Die Retrodigitalisierung startete 2015 mit der Digitalisierung der Best\u00e4nde der TIB. Begonnen hat das Team Retrodigitalisierung mit einem kleinen Team mit drei Personen, das sich im Laufe der Zeit auf aktuell 12 Personen vergr\u00f6\u00dfert hat. In diesem Zeitraum von elf Jahren konnte die Retrodigitalisierung viele B\u00fccher, Graphiken, Pl\u00e4ne usw. mit vielen, vielen Seiten scannen.</p>\n<p>Die Retrodigitalisierung digitalisiert sowohl gemeinfreie Werke sowie noch urheberrechtgebundene Werke zur Bestandserhaltung. Gemeinfrei ist ein Werk, wenn der Urheber des Werks bereits 70 Jahre verstorben ist. Zurzeit sind mehr als 12.000 Digitalisate der TIB online im <a href=\"https://goobi.tib.eu/viewer/index/\">TIB-Viewer</a> verf\u00fcgbar. Seit ein paar Wochen befinden sich diese Digitalisate nun auch in der Deutschen Digitalen Bibliothek (DDB).</p>\n<figure id=\"attachment_33535\" aria-describedby=\"caption-attachment-33535\" style=\"width: 701px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-33535\" src=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild1.png\" alt=\"\" width=\"701\" height=\"484\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild1.png 907w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild1-300x207.png 300w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild1-768x530.png 768w\" sizes=\"auto, (max-width: 701px) 100vw, 701px\" /><figcaption id=\"caption-attachment-33535\" class=\"wp-caption-text\">Startseite Deutsche Digitale Bibliothek</figcaption></figure>\n<h2>Was ist die Deutsche Digitale Bibliothek?</h2>\n<p>Die <a href=\"https://www.deutsche-digitale-bibliothek.de/\">Deutsche Digitale Bibliothek</a> (DDB) ist ein Portal f\u00fcr Kulturobjekte, die neben B\u00fcchern auch Archivalien, Bilder, Fotografien, Skulpturen, Musikst\u00fccke, Filme, Noten, Gem\u00e4lde, Handschriften usw. beinhaltet.</p>\n<p>2012 startete die DDB mit einer Betaversion und 2014 mit der Vollversion. Finanziert wird die DDB von Bund, L\u00e4ndern und Kommunen. Getragen wird die DDB von einen Kompetenznetzwerk bestehend aus <a href=\"https://www.deutsche-digitale-bibliothek.de/content/wie-wir-organisiert-sind\">18 Kultur- und Wissenseinrichtungen</a>, wie zum Beispiel die Bayerische Staatsbibliothek, das Bundesarchiv, die Deutsche Nationalbibliothek, das FIZ Karlsruhe (Leibniz-Institut f\u00fcr Informationsinfrastruktur) oder die Stiftung Preu\u00dfischer Kulturbesitz.</p>\n<p>Das Ziel der DDB ist es, das kulturelle Erbe in Deutschland digital, kostenlos und jederzeit zug\u00e4nglich zu machen. \u00dcber <a href=\"https://www.deutsche-digitale-bibliothek.de/about-us/institutions?view=map\">5.000 Kultureinrichtungen</a> sind an der DDB beteiligt, davon fungieren fast 1.000 als Datenpartner:innen, darunter auch die TIB \u2013 Leibniz-Informationszentrum Technik und Naturwissenschaften und Universit\u00e4tsbibliothek. Zurzeit finden sich \u00fcber 65 Millionen Objekte in der DDB.</p>\n<p>Mittlerweile geh\u00f6ren drei weitere Portale zur DDB:</p>\n<ul>\n<li><a href=\"https://www.archivportal-d.de/\">Archivportal-D</a></li>\n<li><a href=\"https://www.deutsche-digitale-bibliothek.de/newspaper\">Deutsches Zeitungsportal</a></li>\n<li><a href=\"https://ccc.deutsche-digitale-bibliothek.de/de/\">Sammlungsgut aus kolonialen Kontexten</a></li>\n</ul>\n<figure id=\"attachment_33540\" aria-describedby=\"caption-attachment-33540\" style=\"width: 874px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-33540\" src=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild5-1.png\" alt=\"\" width=\"874\" height=\"107\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild5-1.png 907w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild5-1-300x37.png 300w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild5-1-768x94.png 768w\" sizes=\"auto, (max-width: 874px) 100vw, 874px\" /><figcaption id=\"caption-attachment-33540\" class=\"wp-caption-text\">Weitere Portale der Deutschen Digitalen Bibliothek</figcaption></figure>\n<p>Die Deutsche Digitale Bibliothek ist ein Datenpartner f\u00fcr Europeana. Bei der DDB stehen Kultureinrichtungen aus Deutschland im Mittelpunkt. <a href=\"https://www.europeana.eu/de\">Europeana</a> bietet eine Plattform f\u00fcr Kulturobjekte aus ganz Europa.</p>\n<h2>Deutsche Digitale Bibliothek und die TIB</h2>\n<p>Die TIB ist bereits seit etwa 2012/2013 Datenpartnerin von Europeana und der DDB. Seit diesem Zeitpunkt werden Inhalte vom <a href=\"https://av.tib.eu/\">TIB AV-Portal</a> in die DDB und Europeana eingespielt. Sp\u00e4ter wurde der DDB-Bestand um die graphischen Einzelbl\u00e4tter der Sammlung Albrecht Haupt sowie um die Reiseskizzen des Architekten Albrecht Haupt (1852-1932) erg\u00e4nzt. Diese Digitalisate des Teilbestands der Sammlung Albrecht Haupt stammen aus dem Portal <a href=\"https://sah.tib.eu/\">TIB SAH digital</a>.</p>\n<p>Dieses Jahr sind die Digitalisate der Retrodigitalisierung der TIB hinzugekommen. Daf\u00fcr mussten die Digitalisate zun\u00e4chst \u00fcber eine Schnittstelle in das Testsystem der DDB eingespielt werden. Dort wurden die Metadaten analysiert, da die Metadaten gewissen Standards gerecht werden m\u00fcssen, damit die Digitalisate auch in der DDB angezeigt werden k\u00f6nnen. Nach der Pr\u00fcfung konnten die Daten gleich ins Produktivsystem eingespielt werden.</p>\n<figure id=\"attachment_33537\" aria-describedby=\"caption-attachment-33537\" style=\"width: 701px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-33537\" src=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild3.png\" alt=\"\" width=\"701\" height=\"324\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild3.png 907w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild3-300x139.png 300w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild3-768x355.png 768w\" sizes=\"auto, (max-width: 701px) 100vw, 701px\" /><figcaption id=\"caption-attachment-33537\" class=\"wp-caption-text\">Metadaten in der Deutschen Digitalen Bibliothek</figcaption></figure>\n<p>Seit Anfang Juli 2026 befinden sich nun \u00fcber 12.000 <a href=\"https://www.deutsche-digitale-bibliothek.de/searchresults?query=&amp;offset=0&amp;rows=20&amp;facetValues%5B%5D=provider_id%3D4E6QWP2Z7ZV3XWT6D6MX7WSQN4HOKEWO&amp;isThumbnailFiltered=false&amp;facetValues%5B%5D=type_fct%3Dmediatype_003\">Digitalisate der TIB</a> in der DDB. In der DDB besteht die M\u00f6glichkeit, sich die Digitalisate anzuschauen oder auf den Datengeber des Digitalisats zu wechseln, in diesem Fall den <a href=\"https://goobi.tib.eu/viewer/index/\">TIB-Viewer</a>.</p>\n<p><a href=\"https://www.deutsche-digitale-bibliothek.de/item/236JLJM2AUG5TZUDI5M2M3BGRFP4VMWZ\">Kompendium der praktischen Toxikologie</a></p>\n<figure id=\"attachment_33536\" aria-describedby=\"caption-attachment-33536\" style=\"width: 701px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-33536\" src=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild2.png\" alt=\"\" width=\"701\" height=\"283\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild2.png 907w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild2-300x121.png 300w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild2-768x310.png 768w\" sizes=\"auto, (max-width: 701px) 100vw, 701px\" /><figcaption id=\"caption-attachment-33536\" class=\"wp-caption-text\">Bildanzeige in der Deutschen Digitalen Bibliothek</figcaption></figure>\n<p>Im n\u00e4chsten Schritt sollen die Digitalisate von der DDB an Europeana \u00fcbermittelt werden. Geplant sind viertelj\u00e4hrliche Updates an die DDB, damit auch die neuen Digitalisate der Retrodigitalisierung der TIB an die DDB \u00fcbermittelt werden k\u00f6nnen.</p>\n<figure id=\"attachment_33538\" aria-describedby=\"caption-attachment-33538\" style=\"width: 701px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-33538\" src=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild4.png\" alt=\"\" width=\"701\" height=\"327\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild4.png 907w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild4-300x140.png 300w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild4-768x358.png 768w\" sizes=\"auto, (max-width: 701px) 100vw, 701px\" /><figcaption id=\"caption-attachment-33538\" class=\"wp-caption-text\">Startseite Europeana</figcaption></figure>","doi":"https://doi.org/10.65527/hc09t-hvv39","guid":"https://blog.tib.eu/?p=33534","image":"https://blog.tib.eu/wp-content/uploads/2026/09/Bild1.png","language":"de","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1790035200,"rid":"f468w-wsq28","summary":"Die Retrodigitalisierung startete 2015 mit der Digitalisierung der TIB-Best\u00e4nde. Begonnen hat das Team Retrodigitalisierung mit einem kleinen Team mit drei Personen, inzwischen sind es zw\u00f6lf Personen. In diesem Zeitraum von elf Jahren scannten sie B\u00fccher, Graphiken, Pl\u00e4ne usw. mit vielen, vielen Seiten. Zurzeit sind mehr als 12.000 Digitalisate der TIB online im TIB-Viewer verf\u00fcgbar.","tags":["Vom Papier Zum Pixel","DEINE UNIBIB","BIBLIOTHEKSWELT","Lizenz:CC-BY-4.0-INT","Retrodigitalisierung"],"title":"Digitalisate der TIB in der Deutschen Digitalen Bibliothek","updated_at":1790082085,"url":"https://blog.tib.eu/2026/09/22/digitalisate-der-tib-in-der-deutschen-digitalen-bibliothek/","version":"v1"}},{"document":{"authors":[{"contributor_roles":[],"name":"Pensoft Editorial Team"}],"blog":{"authors":null,"community_id":"1e54953c-587a-4743-89af-ebd2cb45660c","created":1789516800,"current_feed_url":null,"description":null,"doi":"https://doi.org/10.59350/pensoft","favicon":"https://rogue-scholar.org/api/communities/1e54953c-587a-4743-89af-ebd2cb45660c/logo","feed_format":"application/atom+xml","feed_url":"https://blog.pensoft.net/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.pensoft.net/","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"pensoft","status":"active","subfield":"2303","title":"Pensoft Blog","updated":1790009683,"use_api":true},"blog_name":"Pensoft Blog","blog_slug":"pensoft","content_html":"<div class=\"twitter-share\"><a href=\"https://twitter.com/intent/tweet?via=Pensoft\" class=\"twitter-share-button\" data-size=\"large\">Tweet</a></div>\n\n<p>Digital biodiversity data has come a long way. Twenty-five years ago, the <a href=\"https://www.gbif.org/\">Global Biodiversity Information Facility</a> (GBIF) launched with 21 founding countries and a simple but ambitious goal: make the world&#8217;s biodiversity data freely and openly available to anyone, anywhere. </p>\n\n\n<div class=\"wp-block-essential-blocks-table-of-contents\"><div class=\"eb-parent-wrapper eb-parent-eb-toc-rs5x1 \"><div class=\"eb-toc-container eb-toc-rs5x1  eb-toc-is-not-sticky eb-toc-not-collapsible eb-toc-initially-not-collapsed eb-toc-scrollToTop style-1 list-style-none\" data-scroll-top=\"false\" data-scroll-top-icon=\"fas fa-angle-up\" data-collapsible=\"false\" data-sticky-hide-mobile=\"false\" data-sticky=\"false\" data-scroll-target=\"scroll_to_toc\" data-copy-link=\"false\" data-editor-type=\"\" data-hide-desktop=\"false\" data-hide-tab=\"false\" data-hide-mobile=\"false\" data-itemCollapsed=\"false\"><div class=\"eb-toc-header\"><div class=\"eb-toc-title\">Table of Contents</div></div><div class=\"eb-toc-wrapper \" data-headers=\"[{&quot;level&quot;:2,&quot;content&quot;:&quot;A Growing Global Network&quot;,&quot;text&quot;:&quot;A Growing Global Network&quot;,&quot;link&quot;:&quot;a-growing-global-network&quot;},{&quot;level&quot;:2,&quot;content&quot;:&quot;Data Supporting Science and Policy&quot;,&quot;text&quot;:&quot;Data Supporting Science and Policy&quot;,&quot;link&quot;:&quot;data-supporting-science-and-policy&quot;},{&quot;level&quot;:2,&quot;content&quot;:&quot;Adapting to New Kinds of Data&quot;,&quot;text&quot;:&quot;Adapting to New Kinds of Data&quot;,&quot;link&quot;:&quot;adapting-to-new-kinds-of-data&quot;},{&quot;level&quot;:2,&quot;content&quot;:&quot;Persistent Gaps Remain&quot;,&quot;text&quot;:&quot;Persistent Gaps Remain&quot;,&quot;link&quot;:&quot;persistent-gaps-remain&quot;},{&quot;level&quot;:2,&quot;content&quot;:&quot;Looking Ahead&quot;,&quot;text&quot;:&quot;Looking Ahead&quot;,&quot;link&quot;:&quot;looking-ahead&quot;}]\" data-visible=\"[true,true,true,true,true,true]\" data-delete-headers=\"[{&quot;label&quot;:&quot;A Growing Global Network&quot;,&quot;value&quot;:&quot;a-growing-global-network&quot;,&quot;isDelete&quot;:false},{&quot;label&quot;:&quot;Data Supporting Science and Policy&quot;,&quot;value&quot;:&quot;data-supporting-science-and-policy&quot;,&quot;isDelete&quot;:false},{&quot;label&quot;:&quot;Adapting to New Kinds of Data&quot;,&quot;value&quot;:&quot;adapting-to-new-kinds-of-data&quot;,&quot;isDelete&quot;:false},{&quot;label&quot;:&quot;Persistent Gaps Remain&quot;,&quot;value&quot;:&quot;persistent-gaps-remain&quot;,&quot;isDelete&quot;:false},{&quot;label&quot;:&quot;Looking Ahead&quot;,&quot;value&quot;:&quot;looking-ahead&quot;,&quot;isDelete&quot;:false}]\" data-smooth=\"true\" data-top-offset=\"\"><div class=\"eb-toc__list-wrap\"><ul class='eb-toc__list'><li><a href=\"#a-growing-global-network\">A Growing Global Network</a><li><a href=\"#data-supporting-science-and-policy\">Data Supporting Science and Policy</a><li><a href=\"#adapting-to-new-kinds-of-data\">Adapting to New Kinds of Data</a><li><a href=\"#persistent-gaps-remain\">Persistent Gaps Remain</a><li><a href=\"#looking-ahead\">Looking Ahead</a></ul></div></div></div></div></div>\n\n\n<p>Today, GBIF's network includes <strong>70 participating countries</strong>, along with data providers across more than 120 countries, offering open access to <strong>over 3.8 billion biodiversity data records</strong>. This milestone is explored in a <a href=\"https://doi.org/10.3897/BDJ.14.e208528\" title=\"\">new overview article published</a> in the <em><a href=\"https://bdj.pensoft.net/\" title=\"\">Biodiversity Data Journal</a></em>, which traces GBIF&#8217;s trajectory over the past quarter-century and the challenges that remain as it looks ahead.</p>\n\n\n\n<p>GBIF Executive Secretary Dr. Joe Miller remarked:</p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>As GBIF celebrates its 25th anniversary, I'm very proud to serve as Executive Secretary of the organisation and to have contributed to this milestone publication. Detailing the history of GBIF both as an infrastructure and a global network, the article pays tribute to all the people involved in the network throughout a quarter century\u2014be they staff, participants, nodes, volunteers, data publishers or data users. It represents the past, present and future of GBIF. I believe the paper will help strengthen our network and its capacity to respond to biodiversity data needs both now and in the future.</p>\n</blockquote>\n\n\n\n<p>GBIF was established in 2001, following a proposal from the Organisation for Economic Cooperation and Development's Megascience Forum which concluded that \"<em>an international mechanism is needed to make biodiversity data and information accessible worldwide</em>\". </p>\n\n\n\n<p>The call came in response to a problem facing the international community after the 1992 Rio Earth Summit: the world&#8217;s biodiversity knowledge was scattered across a small number of institutions, concentrated in wealthy countries, with no shared way to bring it together in support of the newly signed Convention on Biological Diversity.</p>\n\n\n\n<p class=\"has-electric-grass-gradient-background has-background\">Now, a quarter-century later, GBIF has become the foundational infrastructure for biodiversity science and policy worldwide.</p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>A Growing Global Network</strong></h2>\n\n\n\n<figure class=\"wp-block-image size-large\"><a href=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718328.gif?ssl=1\"><img fetchpriority=\"high\" decoding=\"async\" width=\"840\" height=\"420\" src=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718328-1024x512.gif?resize=840%2C420&#038;ssl=1\" alt=\"Biodiversity data in GBIF\" class=\"wp-image-20842\" srcset=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718328.gif?resize=1024%2C512&amp;ssl=1 1024w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718328.gif?resize=300%2C150&amp;ssl=1 300w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718328.gif?resize=768%2C384&amp;ssl=1 768w\" sizes=\"(max-width: 709px) 85vw, (max-width: 909px) 67vw, (max-width: 1362px) 62vw, 840px\" data-recalc-dims=\"1\" /></a><figcaption class=\"wp-element-caption\">Annual biodiversity data made available through the Global Biodiversity Information Facility as of July 2026. Credit to GBIF.</figcaption></figure>\n\n\n\n<p>GBIF&#8217;s work is carried out through a distributed network of 47 voting country participants and 42 organisational participants, each represented by a national or thematic &#8220;node&#8221; that mobilises data, builds local capacity, and connects biodiversity communities to GBIF&#8217;s infrastructure. More than 2,700 institutions, including museums, universities, government agencies, citizen science platforms, and a growing number of private-sector organisations, have published data through the network, with new publishers joining at a rate of more than two every three days, totalling <a href=\"https://www.gbif.org/publisher/search\">nearly 3500</a> organisations involved.</p>\n\n\n\n<figure class=\"wp-block-image size-large\"><a href=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?ssl=1\"><img decoding=\"async\" width=\"840\" height=\"964\" src=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-892x1024.jpg?resize=840%2C964&#038;ssl=1\" alt=\"key numerical metrics of the Global Biodiversity Information Facility \" class=\"wp-image-20845\" srcset=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=892%2C1024&amp;ssl=1 892w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=261%2C300&amp;ssl=1 261w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=768%2C882&amp;ssl=1 768w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=1338%2C1536&amp;ssl=1 1338w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=1784%2C2048&amp;ssl=1 1784w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=1200%2C1378&amp;ssl=1 1200w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?w=1680 1680w\" sizes=\"(max-width: 709px) 85vw, (max-width: 909px) 67vw, (max-width: 1362px) 62vw, 840px\" data-recalc-dims=\"1\" /></a><figcaption class=\"wp-element-caption\">Infographic summary of key numerical metrics of the Global Biodiversity Information Facility as of August 2026. See dynamic metrics at <a href=\"https://www.gbif.org/\" target=\"_blank\" rel=\"noreferrer noopener\">GBIF.org</a> and at <a href=\"https://www.gbif.org/analytics/global\" target=\"_blank\" rel=\"noreferrer noopener\">https://www.gbif.org/analytics/global</a>.</figcaption></figure>\n\n\n\n<p>The <em>Biodiversity Data Journal</em> itself is one example of this network in action, as it operates <a href=\"https://blog.pensoft.net/2025/03/10/the-biodiversity-data-journal-launches-its-own-data-portal-on-gbif/\" title=\"\">its <strong>own GBIF-hosted data portal </strong></a>&#8211; one of <a href=\"https://blog.pensoft.net/2025/04/03/more-than-20-journals-published-by-pensoft-with-their-own-hosted-data-portals-on-gbif-to-streamline-and-fair-ify-biodiversity-research/\" title=\"\"><strong>more than 20 </strong>such portals across Pensoft-published journals </a>&#8211; providing direct access to hundreds of datasets and hundreds of thousands of occurrence records drawn from the journal&#8217;s publications.</p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Data Supporting Science and Policy</strong></h2>\n\n\n\n<figure class=\"wp-block-image size-large\"><a href=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?ssl=1\"><img decoding=\"async\" width=\"840\" height=\"391\" src=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330-1024x477.png?resize=840%2C391&#038;ssl=1\" alt=\"GBIF data mediation\" class=\"wp-image-20847\" srcset=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=1024%2C477&amp;ssl=1 1024w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=300%2C140&amp;ssl=1 300w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=768%2C358&amp;ssl=1 768w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=1536%2C715&amp;ssl=1 1536w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=2048%2C953&amp;ssl=1 2048w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=1200%2C559&amp;ssl=1 1200w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?w=1680 1680w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?w=2520 2520w\" sizes=\"(max-width: 709px) 85vw, (max-width: 909px) 67vw, (max-width: 1362px) 62vw, 840px\" data-recalc-dims=\"1\" /></a><figcaption class=\"wp-element-caption\"><a href=\"https://doi.org/10.3897/BDJ.14.e208528.figure3\" title=\"\">The universe of GBIF data mediation. </a></figcaption></figure>\n\n\n\n<p>GBIF-mediated data now underpin an average of eight new peer-reviewed research papers per day, and have contributed to <a href=\"https://www.gbif.org/literature/search?peerReview=true&amp;literatureType=journal\">more than 15,000 publications to date</a>, spanning fields from climate change and food security to public health and invasive species management. An independent 2023 assessment <a href=\"https://www.gbif.org/news/5WZThcL928vmPnSvrGhZfE/report-reveals-return-on-investments-in-gbif\">estimated</a> that GBIF generates <strong>roughly \u20ac12 in societal benefit </strong>for<strong> every \u20ac1 invested</strong>, with researcher time savings alone valued at \u20ac35 million annually.</p>\n\n\n\n<p>The organisation&#8217;s impact also extends deeply into international policy. GBIF-mediated data support multiple indicators under the Kunming-Montreal Global Biodiversity Framework, inform assessments by the <a href=\"https://iucn.org/\">International Union for Conservation of Nature</a> (IUCN) and the <a href=\"https://www.ipbes.net/\">Intergovernmental Science-Policy Platform on Biodiversity and Ecosystem Services</a> (IPBES), and are increasingly used by the private sector to meet emerging nature-related disclosure requirements.</p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Adapting to New Kinds of Data</strong></h2>\n\n\n\n<p>Over 25 years, GBIF has continually expanded to accommodate new sources of biodiversity data &#8211; from digitised natural history specimens and citizen science observations to DNA-based records and, most recently, structured survey and monitoring data designed to support large-scale biodiversity tracking. In 2026, GBIF adopted the <a href=\"https://www.catalogueoflife.org/\">Catalogue of Life</a> as its taxonomic backbone, the result of a multi-year collaboration to build shared infrastructure for reconciling species names across datasets.</p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Persistent Gaps Remain</strong></h2>\n\n\n\n<figure class=\"wp-block-image size-large\"><a href=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?ssl=1\"><img loading=\"lazy\" decoding=\"async\" width=\"840\" height=\"491\" src=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053-1024x599.png?resize=840%2C491&#038;ssl=1\" alt=\"key events of GBIF\" class=\"wp-image-20850\" srcset=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=1024%2C599&amp;ssl=1 1024w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=300%2C176&amp;ssl=1 300w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=768%2C449&amp;ssl=1 768w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=1536%2C899&amp;ssl=1 1536w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=2048%2C1198&amp;ssl=1 2048w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=1200%2C702&amp;ssl=1 1200w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?w=1680 1680w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?w=2520 2520w\" sizes=\"(max-width: 709px) 85vw, (max-width: 909px) 67vw, (max-width: 1362px) 62vw, 840px\" data-recalc-dims=\"1\" /></a><figcaption class=\"wp-element-caption\"><br>A timeline of key events in the history of the Global Biodiversity Information Facility. Credit to Miller et al., 2026.</figcaption></figure>\n\n\n\n<p>Despite this growth, GBIF recognises that<strong> several significant challenges remain</strong>. The global loss of biodiversity continues to be described as a crisis in major international assessments, even as the data infrastructure has improved. And the <strong>data </strong>itself continues to<strong> reflect long-standing imbalances</strong> &#8211; regions with high biodiversity, particularly in the Global South, remain underrepresented, while well-studied, charismatic groups such as birds are overrepresented compared to less visible taxa; furthermore, many marine species and hard-to-identify cryptic habitats and organisms remain inadequately documented.</p>\n\n\n\n<p>Closing these gaps was part of the original motivation for creating GBIF a quarter-century ago, and it remains valid today. GBIF&#8217;s own assessment is that the network cannot resolve these asymmetries alone, doing so will depend on broader shifts in global scientific funding, data culture, and capacity, alongside GBIF&#8217;s continued efforts to expand its Participant network and to diversify the types of data it can mobilise.</p>\n\n\n\n<p>An additional, ongoing challenge is <strong>data heterogeneity</strong> and <strong>data quality</strong>, because GBIF indexes data, rather than directly vetting every record, it relies on its network of data publishers to maintain quality at source, and on users to report issues \u2013 a distributed model that keeps the network scalable but is not without friction.</p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Looking Ahead</strong></h2>\n\n\n\n<figure class=\"wp-block-image size-large\"><a href=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?ssl=1\"><img loading=\"lazy\" decoding=\"async\" width=\"840\" height=\"513\" src=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337-1024x625.png?resize=840%2C513&#038;ssl=1\" alt=\"capacity development at GBIF.\" class=\"wp-image-20853\" srcset=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=1024%2C625&amp;ssl=1 1024w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=300%2C183&amp;ssl=1 300w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=768%2C469&amp;ssl=1 768w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=1536%2C937&amp;ssl=1 1536w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=2048%2C1250&amp;ssl=1 2048w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=1200%2C732&amp;ssl=1 1200w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?w=1680 1680w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?w=2520 2520w\" sizes=\"(max-width: 709px) 85vw, (max-width: 909px) 67vw, (max-width: 1362px) 62vw, 840px\" data-recalc-dims=\"1\" /></a><figcaption class=\"wp-element-caption\">Capacity development across the Global Biodiversity Information Facility community.&nbsp; Credit to Miller et al., 2026.</figcaption></figure>\n\n\n\n<p>Guided by its 2023\u20132027 Strategic Framework, GBIF is prioritising continued growth of its Participant network, expanded support for survey and monitoring data, and deeper integration of DNA-derived biodiversity records \u2013 all aimed at closing longstanding geographic and taxonomic gaps in digital biodiversity data worldwide.</p>\n\n\n\n<p><strong>Original source:</strong></p>\n\n\n\n<p>Miller J, Mandeville C, Schigel D, Gamboa Martinez J, Nielsen AM, Sheldon S, Hahn A, Raymond M, Bundgaard-Jensen S, Bagard Laursen M, Copas K, Blissett M, H\u00f8fft M, M\u00e9ndez Hern\u00e1ndez F, Noesgaard D, Rodrigues A, Russell L, Stjernegaard Jeppesen T, Elkj\u00e6r \u00d8rum-Kristensen A, Grosjean M, Waller J, Podolskiy M, Goodson H, Novakovikj S, Fr\u00f8slev T, Ingenloff K, S\u00f8rensen Nilsson A, Svenningsen C, van der Meijden D, Suen A, Hakan Uzun A, Vaskova M, Nielsen C, Marentes Herrera E, Schaldemose Reibke N, Robertson T (2026) The Global Biodiversity Information Facility at 25. Biodiversity Data Journal 14: e208528. <a href=\"https://doi.org/10.3897/BDJ.14.e208528\" target=\"_blank\" rel=\"noreferrer noopener\">https://doi.org/10.3897/BDJ.14.e208528</a></p>\n\n<div class=\"twitter-share\"><a href=\"https://twitter.com/intent/tweet?via=Pensoft\" class=\"twitter-share-button\" data-size=\"large\">Tweet</a></div>","doi":"https://doi.org/10.59350/8k14v-ks894","guid":"https://blog.pensoft.net/?p=20832","image":"https://blog.pensoft.net/wp-content/uploads/2026/09/16_9-Journal-comms-2.png","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1790035200,"rid":"cc3mg-ynn67","summary":"Tweet Digital biodiversity data has come a long way. Twenty-five years ago, the Global Biodiversity Information Facility (GBIF) launched with 21 founding countries and a simple but ambitious goal: make the world's biodiversity data freely and openly available to anyone, anywhere.","tags":["Biodiversity Data Journal","Anniversary","Biodiversity","Citizen Science","Data"],"title":"GBIF Marks 25 Years of FAIR Biodiversity Data, Now Spanning 70 Countries and 3.8 Billion Records","updated_at":1790082047,"url":"https://blog.pensoft.net/2026/09/22/gbif-marks-25-years-of-fair-biodiversity-data-now-spanning-70-countries-and-3-8-billion-records/","version":"v1"}},{"document":{"authors":[{"contributor_roles":[],"family":"J\u00e4ckel","given":"Florian"}],"blog":{"authors":null,"community_id":"d2883c5c-8e82-4bc0-a892-92a8b752c7d9","created":1763683200,"current_feed_url":null,"description":"A blog about the dblp computer science bibliography","doi":"https://doi.org/10.59350/dblp","favicon":"https://rogue-scholar.org/api/communities/d2883c5c-8e82-4bc0-a892-92a8b752c7d9/logo","feed_format":"application/atom+xml","feed_url":"https://blog.dblp.org/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.dblp.org","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"dblp","status":"active","subfield":"1710","title":"blog.dblp.org","updated":1790003754,"use_api":false},"blog_name":"blog.dblp.org","blog_slug":"dblp","content_html":"<p>Most scientific publications name the institutions their authors were affiliated with at the time of writing. The information is recorded alongside the author names, and it often sits in the metadata that dblp processes.</p>\n<p>We are happy to announce that this information is now properly represented in our data: <strong>research institutions have become proper entities in the dblp computer science bibliography</strong>. The affiliations that connect people and publications to those institutions are now recorded as data as well. In practice, this means five new things:</p>\n<ul>\n<li>you can <strong>filter publication lists by affiliation</strong>,</li>\n<li>institutions have their own <strong>landing pages</strong> on dblp.org,</li>\n<li>you can <strong>search for institutions</strong> just like you search for authors or venues,</li>\n<li>institutions and affiliations are part of the <strong>dblp Knowledge Graph</strong>,</li>\n<li>and they are therefore also available in our <strong><a href=\"https://blog.dblp.org/2024/09/09/introducing-our-public-sparql-query-service/\">SPARQL query service</a></strong>.</li>\n</ul>\n<p>This work is a main outcome of <a href=\"https://www.dagstuhl.de/en/institute/projects/smarter-affiliations\"><strong>SmartER Affiliations</strong></a>, a joint project of Schloss Dagstuhl, <a href=\"https://www.gesis.org/en/home\">GESIS \u2013 Leibniz Institute for the Social Sciences</a>, and the <a href=\"https://www.uni-ulm.de/en/in/iui-dbis/home/\">Institute of Databases and Information Systems at Ulm University</a>, funded by the German Research Foundation (DFG).</p>\n<p>Below, we introduce the new features, explain how the data fits together, and give some numbers on coverage.</p>\n<h2 id=\"institutions-and-affiliations\">Institutions and affiliations</h2>\n<p>An <strong>institution</strong> is an <em>entity</em>. It has an identity of its own, recorded under a dblp identifier that starts with <code>inst/</code>, together with its names, acronyms, external identifiers, and other data. This data is found on the institution's landing page (cf. the screenshot further down below), just as with other entities in dblp. Institutions exist independently of whether anyone is currently affiliated with them. dblp currently knows more than 10,000 institutions in 195 countries.</p>\n<p>An <strong>affiliation</strong> is a <em>link</em> between entities. It says that some entity in dblp is connected to an institution. The relationship itself is the thing being recorded, and it carries its own metadata, such as the time span it covers and whether it is a person's current or a former affiliation.</p>\n<p>In dblp, that link comes in two variants:</p>\n<p><strong>Person-based affiliations</strong> connect a person to an institution. These are recorded manually by dblp curators, which makes them deliberate and reviewed, but also far fewer than the automatically retrieved ones: about 256,000 affiliation statements for some 195,000 people (of about 4,2 million persons total). The large majority describe a person's current institution, and around 24,000 are marked as former affiliations. On author pages, they still appear as the free-text notes you may know from before, but in dblp's data they are now links to institution entities.</p>\n<p><strong>Signature-based affiliations</strong> connect a <em>signature</em> to an institution. A signature is the link between a publication and one of its authors, the concrete act of that person authoring that paper. Signature-based affiliations therefore express something much more specific: \"in this paper, this author gave this institution as their affiliation.\" These affiliations are retrieved automatically at scale, mainly by extracting and matching affiliation statements from publication metadata. About 60% of all signatures in dblp currently carry at least one signature-based affiliation.</p>\n<h2 id=\"filtering-publications-by-affiliation\">Filtering publications by affiliation</h2>\n<p>The most immediately visible change is that publication lists can now be <strong>filtered by affiliation</strong>. On the pages where this applies, you will find affiliations as a new facet next to the filters you already use, and it can be freely combined with them.</p>\n<div class=\"wp-caption alignright\" id=\"attachment_937\" style=\"width: 316px\"><img alt=\"Filter facet\" aria-describedby=\"caption-attachment-937\" class=\"wp-image-937 size-full\" decoding=\"async\" fetchpriority=\"high\" height=\"930\" sizes=\"(max-width: 306px) 100vw, 306px\" src=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-29-26.png\" srcset=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-29-26.png 306w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-29-26-99x300.png 99w\" width=\"306\"/><p class=\"wp-caption-text\" id=\"caption-attachment-937\">New filter facet: \"refine by affiliation\"</p></div>\n<p>This filter operates on signature-based affiliations, and what it matches depends on where you are. On an author page, <strong>refine by affiliation</strong> offers that person's own affiliations and filters by those alone, so you get the papers this author wrote while at that institution. Everywhere else, in search results and on tables of contents, <strong>refine by institution</strong> matches a publication if <em>any</em> of its authors gave that institution on that publication.</p>\n<p>Either way, the result is a <strong>subset</strong> of the papers in this view that <strong>are known</strong> to have been written at the institution. Obviously, someone who has just moved there does not bring their earlier publications along with them.</p>\n<p>What any of these filters can return is limited above all by whether affiliation data exists at signature level, and very often it does not. So you get the publications for which we currently have an affiliation statement that we were able to match with sufficient confidence, rather than everything an institution has published. The data is knowingly incomplete, and we treat it as work in progress that will keep improving. More on the limits of the data below.</p>\n<h2 id=\"institution-landing-pages\">Institution landing pages</h2>\n<p>Every institution in dblp now has its own page. It collects what dblp knows about the institution and what the institution links to:</p>\n<ul>\n<li>the preferred name, plus alternative names, translations, and acronyms we know about,</li>\n<li>basic descriptive information, such as the institution's location(s) and related institutions,</li>\n<li>a <strong>visit</strong> menu with links to the institution's own web presence, its Wikipedia article, and authority control records,</li>\n<li>an <strong>export institution</strong> menu offering XML and three RDF serializations, along with the dblp key of the institution, for example <code>inst/11/687</code>,</li>\n<li>an <strong>ask others</strong> menu that hands the institution over to external search services,</li>\n<li>and the people affiliated with the institution, including those who earned their PhD there. This listing does not distinguish current from former affiliations.</li>\n</ul>\n<p>Wherever possible, the page contains links to the institution's data records in other projects and services. The underlying institution data is kept in sync with ROR, so that changes there find their way into dblp as well.</p>\n<p>Institutions are also browsable by country, in the same way you browse the rest of dblp.</p>\n<div class=\"wp-caption aligncenter\" id=\"attachment_941\" style=\"width: 760px\"><img alt=\"\" aria-describedby=\"caption-attachment-941\" class=\"wp-image-941 size-large\" decoding=\"async\" height=\"240\" sizes=\"(max-width: 750px) 100vw, 750px\" src=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41-1024x327.png\" srcset=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41-1024x327.png 1024w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41-300x96.png 300w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41-768x245.png 768w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41.png 1314w\" width=\"750\"/><p class=\"wp-caption-text\" id=\"caption-attachment-941\">Institution landing page \u2013 here: Schloss Dagstuhl</p></div>\n<h2 id=\"searching-for-institutions\">Searching for institutions</h2>\n<p>Institutions have also been added to dblp's search. You can look them up by their preferred name, by an alternative name, or by an acronym, and follow the result straight to the institution's landing page.</p>\n<div class=\"wp-caption aligncenter\" id=\"attachment_942\" style=\"width: 760px\"><img alt=\"Searching for an institution - here: combined search\" aria-describedby=\"caption-attachment-942\" class=\"wp-image-942 size-large\" decoding=\"async\" height=\"316\" sizes=\"(max-width: 750px) 100vw, 750px\" src=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-39-35-1024x431.png\" srcset=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-39-35-1024x431.png 1024w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-39-35-300x126.png 300w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-39-35-768x323.png 768w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-39-35.png 1314w\" width=\"750\"/><p class=\"wp-caption-text\" id=\"caption-attachment-942\">Searching for an institution \u2013 here: combined search</p></div>\n<h2 id=\"where-the-data-comes-from\">Where the data comes from</h2>\n<p>Two different kinds of data come together here: the affiliation statements of the publications, and the institutions they are matched against.</p>\n<p>The affiliation statements (i.e., where each author was based when working on the paper) mostly come from the data we receive from the publishers. A further source is <a href=\"https://openalex.org/\">OpenAlex</a>.</p>\n<p>Our main source for institution data is the <a href=\"https://ror.org/\"><strong>Research Organization Registry (ROR)</strong></a>. ROR is an open, community-maintained registry of research organizations, released under CC0. It provides persistent identifiers, name variants and acronyms, location information, and links to other identifier systems. Building on ROR means that institutions in dblp are identified in a way that can be joined with data from anyone else who uses ROR IDs \u2014 which by now is a large part of the scholarly metadata landscape.</p>\n<p>That said, we do not simply mirror ROR. <strong>We continue to curate institutions ourselves</strong>, because our requirements do not always match those of a general-purpose registry. Computer science has its own history of institutional mergers, renamings, and re-organisations. Some organisations that matter for our data are not (or not yet) in ROR. And matching thousands of messy affiliation strings to a registry inevitably produces cases that need a human decision. Where ROR and dblp differ, the dblp record reflects a curatorial choice we made on purpose. About four in five of our institutions carry a ROR ID; the rest are ones we maintain on our own.</p>\n<h3 id=\"no-faculties-no-departments-no-institutes\">No faculties, no departments, no institutes</h3>\n<p><strong>dblp does not model faculties, departments, institutes, or research groups as entities of their own.</strong> Generally speaking, the <strong>university is the lowest level</strong> of granularity we model. An affiliation statement naming a specific chair at a specific institute of a specific faculty of a university becomes an affiliation with that university, although the full text of the statement is kept.</p>\n<p>This loses information that some users would like to have. However, sub-institutional units are unstable. They are founded, merged, renamed, and dissolved far more often than their parent organizations, and reconstructing that history retroactively for decades of publications is not realistic. They are also inconsistently reported: the same group may appear as a chair, an institute, a department, or nothing at all, depending on the publisher's template and the author's mood. And they are largely absent from the external registries we rely on for identifiers. So we draw the line at the level we can keep consistent across the whole database.</p>\n<p>Not every institution is a university, of course. Companies, non-university research institutes, government labs and hospitals are modeled in the same way, and the same rule applies to them: we record the organization, not its individual divisions or labs.</p>\n<p>There are a few exceptions where an organization that is formally a part of a larger body is nevertheless modeled as its own institution; typically when it is a distinct legal or research entity with its own identity in ROR. As a rule of thumb, though, we model whole organizations rather than their parts.</p>\n<h2 id=\"institutions-in-the-dblp-knowledge-graph-and-sparql\">Institutions in the dblp Knowledge Graph and SPARQL</h2>\n<p>As announced in <a href=\"https://blog.dblp.org/2024/09/09/introducing-our-public-sparql-query-service/\">our post on the SPARQL query service</a>, extending the <a href=\"https://blog.dblp.org/2024/06/14/the-dblp-knowledge-graph-major-extension-and-an-update-to-the-rdf-schema/\">dblp Knowledge Graph</a> with institution entities and affiliation links was one of our declared next steps. Institutions and person-based affiliations are now part of the dblp KG, are included in our RDF dumps (starting October 2026), and are queryable at <a href=\"https://sparql.dblp.org\"><strong>https://sparql.dblp.org</strong></a>. Signature-based affiliations will follow.</p>\n<p>Institutions are modeled as <code>dblp:Institution</code>, carrying their names, their ROR ID in <code>dblp:ror</code>, and their location. Person-based affiliations connect an author to an institution through <code>dblp:affiliatedWith</code>, with the details of a single affiliation, such as the time span it covers, reachable through <code>dblp:hasAffiliation</code>. The full set of new classes and properties is documented in the <a href=\"https://dblp.org/rdf/docu/\">dblp RDF schema</a>.</p>\n<p>Because affiliations are links rather than attributes, the interesting queries are the ones that traverse them. For example, you can list the institutions with the most affiliated people in dblp:</p>\n<pre class=\"sparql\"><code>PREFIX dblp: <https: dblp.org=\"\" rdf=\"\" schema#=\"\">\nSELECT ?name (COUNT(DISTINCT ?pers) AS ?people) WHERE {\n  ?pers dblp:affiliatedWith ?inst .\n  ?inst dblp:primaryInstitutionName ?name .\n}\nGROUP BY ?name\nORDER BY DESC(?people)\nLIMIT 20</https:></code></pre>\n<p>(<a href=\"https://sparql.dblp.org/Qlzqye\">Run this query on sparql.dblp.org</a>)</p>\n<p>Institutions also carry their ROR ID in <code>dblp:ror</code> and a link to Wikidata in <code>dblp:wikidata</code>, so the data can be combined with any other source that uses those identifiers. Wikidata, for instance, knows when an institution was founded, which is something dblp does not record. The query below takes the 500 institutions with the most affiliated people in dblp and sorts them by their founding date:</p>\n<pre class=\"sparql\"><code>PREFIX dblp: <https: dblp.org=\"\" rdf=\"\" schema#=\"\">\nPREFIX wdt: <http: direct=\"\" prop=\"\" www.wikidata.org=\"\"></http:>\nPREFIX xsd: <http: 2001=\"\" www.w3.org=\"\" xmlschema#=\"\">\nSELECT ?name (year(xsd:dateTime(?inception)) as ?inceptionYear) WHERE {\n{\nSELECT ?wikiq ?name (COUNT(DISTINCT ?pers) AS ?people) WHERE {\n?inst a dblp:Institution ;\ndblp:primaryInstitutionName ?name ;\ndblp:wikidata ?wikiq .\n?pers dblp:affiliatedWith ?inst .\n}\nGROUP BY ?wikiq ?name\nORDER BY DESC(?people)\nLIMIT 500\n}\nSERVICE <https: api=\"\" qlever.dev=\"\" wikidata=\"\"> {\n?wikiq wdt:P571 ?inception .\n}\n}\nORDER BY ?inception</https:></http:></https:></code></pre>\n<p>(<a href=\"https://sparql.dblp.org/yAVdQW\">Run this query on sparql.dblp.org</a>)</p>\n<p>As we noted when the query service was launched, federated queries can be slow, and the number of results exchanged between the two endpoints needs to be kept in check. That is what the inner query with its limit is for.</p>\n<h2 id=\"limits-of-the-data\">Limits of the data</h2>\n<p><strong>Coverage is uneven.</strong> About 60% of all signatures in dblp currently carry an affiliation, 18.2 million out of 30.3 million. For most of what we index, coverage is good. The largest publishers account for roughly two thirds of all signatures, and for those, coverage is typically above 70%. What pulls the average down is a small number of clearly identifiable gaps. By far the largest is arXiv: its publications account for 2.6 million signatures in dblp, and fewer than 2,000 of them currently carry an affiliation. We are actively working on closing that gap. A few smaller gaps exist elsewhere, mostly where affiliations are not part of the metadata we receive. Coverage also peaks for publications from the mid-2010s and declines for recent years, mostly because the affiliation data for recent publications has not been processed yet.</p>\n<p><strong>Automatic extraction makes mistakes.</strong> Signature-based affiliations are machine-generated. Names are ambiguous, institutions get renamed, \"Cambridge\" is not one place, and multi-affiliation authors are common. We invest a lot of effort into matching and into feeding curator corrections back into the pipeline, but errors remain.</p>\n<p><strong>Please do not build rankings out of this.</strong> The count of publications per institution in dblp says at least as much about the coverage of affiliations in dblp as it does about the institutions themselves. dblp's scope is computer science, its indexing decisions are editorial, and affiliation coverage is a moving target. The data works well for exploration, discovery, and disambiguation but it makes a poor basis for evaluating institutions, and we would ask you not to present it that way.</p>\n<p><strong>Affiliations are historical, i.e., a</strong>\u00a0signature-based affiliation records what a paper said at the time it was published. It is not a statement about where someone works today, and an author who moves does not take their earlier publications with them. Person-based affiliations work differently. They can carry a time span, and they distinguish a person's current affiliation from former ones. They are also manually curated, and therefore far fewer in number.</p>\n<h2 id=\"corrections-and-feedback\">Corrections and feedback</h2>\n<p>Wrong affiliations and wrong institutions can be corrected like any other data error in dblp.</p>\n<p>If you notice a systematic problem \u2014 an institution that has been split into duplicates, a merger we have not caught up with, a matching error that affects a whole venue \u2014 telling us about the pattern is much more valuable than reporting individual cases. You can always reach the dblp team at dblp(at)dagstuhl.de. For questions and discussion about the knowledge graph and the SPARQL service, our <a href=\"https://github.com/dblp/kg/discussions\">GitHub Discussions forum</a> is the better place.</p>\n<h2 id=\"outlook\">Outlook</h2>\n<p>There are some things left to do, for example bringing signature-based affiliations into the knowledge graph. We will also use the new affiliation data in our own curation work, above all for author disambiguation.</p>\n<p>We are very interested in how you end up using this data: both because it helps us set priorities, and because affiliation data has a way of revealing use cases nobody anticipated.</p>\n<h2 id=\"acknowledgement\">Acknowledgement</h2>\n<p>The work described in this post has been carried out as part of the project <strong><a href=\"https://www.dagstuhl.de/en/institute/projects/smarter-affiliations\">SmartER Affiliations: Enhancing Open Repositories through Harvesting and Extracting Affiliation Data as First-class Citizen</a></strong>, a cooperation of Schloss Dagstuhl \u2013 Leibniz Center for Informatics, <a href=\"https://www.gesis.org/en/home\">GESIS \u2013 Leibniz Institute for the Social Sciences </a>(Brigitte Mathiak and Asif Suryani), and the <a href=\"https://www.uni-ulm.de/en/in/iui-dbis/home/\">Institute of Databases and Information Systems at Ulm University</a> (Ansgar Scherp and Florian Hauss). The project is funded by a grant of the German Research Foundation (DFG) within the funding program \"e-Research Technologies\" (<a href=\"https://gepris.dfg.de/gepris/projekt/515537520\">grant project number 515537520</a>).</p>\n<p>We would also like to thank <a href=\"https://ror.org/\">the ROR community</a> for maintaining an open registry of research organizations, without which this feature would have looked very different, and much worse.</p>","doi":"https://doi.org/10.59350/ew782-kf260","guid":"https://blog.dblp.org/?p=929","image":"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41-1024x327.png","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1789948800,"rid":"e8dsd-hjk21","summary":"Most scientific publications name the institutions their authors were affiliated with at the time of writing. The information is recorded alongside the author names, and it often sits in the metadata that dblp processes.","tags":["Feature Spotlight","Affiliations","Institutions","Knowledge Graph"],"title":"Institutions as first-class citizens: introducing affiliations in dblp","updated_at":1790081998,"url":"https://blog.dblp.org/2026/09/21/institutions-as-first-class-citizens-introducing-affiliations-in-dblp/","version":"v1"}},{"document":{"authors":[{"contributor_roles":[],"family":"Eden","given":"Terence"}],"blog":{"authors":null,"community_id":"61ce553a-bafd-4aba-a952-d3bab5e85bcc","created":1788652800,"current_feed_url":null,"description":"Regular nonsense about tech and its effects \ud83d\ude43","doi":"https://doi.org/10.59350/shkspr","favicon":"https://rogue-scholar.org/api/communities/61ce553a-bafd-4aba-a952-d3bab5e85bcc/logo","feed_format":"application/atom+xml","feed_url":"https://shkspr.mobi/blog/feed/DOI","filter":"category:-1982","generator":"WordPress","home_page_url":"https://shkspr.mobi/blog","issn":"2753-1570","language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"shkspr","status":"active","subfield":"1712","title":"Terence Eden's Blog","updated":1790106552,"use_api":true},"blog_name":"Terence Eden's Blog","blog_slug":"shkspr","content_html":"<p>Last year I ran <a href=\"https://shkspr.mobi/blog/2025/09/llms-are-still-surprisingly-bad-at-simple-tasks/\">an experiment to test the ability of modern LLMs</a> to correctly answer a relatively straightforward question. Every single one of them got it wrong. Some missed information, some made up false statements, none were right.</p>\n\n<p>Of course the fanbois variously claimed that I was holding it wrong, my prompts were shit, I should have chosen better defaults, and - my favourite - that it would be better next year.</p>\n\n<p>Well, next year is now. 365 days after the original experiment, let's see if these self-reinforcing-learning machines have achieved anything close to intern-levels of competence.</p>\n\n<h2 id=\"the-question-that-started-it-all\"><a href=\"#the-question-that-started-it-all\">The Question That Started It All</a></h2>\n\n<p>I asked:</p>\n\n<blockquote><p>Which TLDs have the same name as valid HTML5 elements?</p></blockquote>\n\n<h2 id=\"why-it-matters\"><a href=\"#why-it-matters\">Why It Matters</a></h2>\n\n<p>This is the sort of question that I would expect a moderately bright teenager to answer. There exists a list which comprehensively includes all TLDs. There is a separate list which contains every HTML element.</p>\n\n<p>One could either manually go through the TLD list comparing it to the HTML list, or write a tiny scrap of code to do the same.</p>\n\n<p>This isn't \"solve Navier-Stokes\" levels of superintelligence. It's a simple but boring task which people are bound to offload onto their digital sycophants.</p>\n\n<h2 id=\"the-quiet-part-out-loud\"><a href=\"#the-quiet-part-out-loud\">The Quiet Part Out Loud</a></h2>\n\n<p>How many LLM generated answers have become load-bearing without scrutiny?</p>\n\n<p>People are lazy, slapdash, and some third thing to make the rule-of-three. If an answer looks answer-shaped, they'll accept it. On a simple but boring task, no reasonable human is going to suspect that a superintelligent computer has made a mistake. All the marketing says that these machines are brainier than you - so why bother?</p>\n\n<h2 id=\"ok-ill-write-a-sub-heading-that-sounds-like-an-ai\"><a href=\"#ok-ill-write-a-sub-heading-that-sounds-like-an-ai\">OK, I'll write a sub-heading that sounds like an AI.</a></h2>\n\n<p>So how did the current crop do? First up <a href=\"https://share.gemini.google/Hta7KuW64Atb\">Google Gemini</a>.</p>\n\n<p>The default \"Flash\" model came up with this:</p>\n\n<blockquote><p>Valid top-level domains (TLDs) registered in the IANA Root Zone Database share exact names with standard HTML elements:</p>\n\n<p>.a \u2014 <code>&lt;a&gt;</code> (Anchor/Hyperlink)</p>\n\n<p>.app \u2014 <code>&lt;app&gt;</code> (Deprecated HTML tag for WebAPKs/WebApps)</p>\n\n<p>.art \u2014 <code>&lt;art&gt;</code> (Historical/non-standard HTML element)</p>\n\n<p>.audio \u2014 <code>&lt;audio&gt;</code></p>\n\n<p>.bar \u2014 <code>&lt;bar&gt;</code> (Historical SVG/HTML draft tag)</p></blockquote>\n\n<p>Then it listed a dozen more. You don't need to be a DNS expert to know that the minimum length of a TLD is two characters - <code>.a</code> simply isn't valid. HTML nerds will know that art, app, and bar have never been elements. Pathetic.</p>\n\n<p>So I tried Gemini's extended thinking model. Thankfully, it didn't make up any imaginary TLDs or elements. It did, however, miss the <code>&lt;data&gt;</code> element which has a valid <code>.data</code> TLD. It also missed <code>map</code>, <code>select</code>, and <code>search</code>.</p>\n\n<p>So, points for not making shit up. But demerits for not being able to compare two text lists.</p>\n\n<p>A friend <a href=\"https://claude.ai/share/a8a408cf-6feb-4ea8-99ea-dddd9aadafb5\">asked Claude</a>. That missed <code>search</code> and <code>select</code>. It didn't report <em>any</em> ccTLDs. You <em>could</em> argue that a country code like <code>li</code> isn't part of the original question - but I'd say that was weak justification; the set of TLDs contains ccTLDs.</p>\n\n<p>A different friend (I have many!) used <a href=\"https://claude.ai/share/c26f44bb-9efa-4b57-9f63-324ef7400cb3\">a different model</a> and, while the answers looked accurate, it included this at the end:</p>\n\n<blockquote><p>Near misses that don't count: .codes, .forum, .pictures, .market, .navy, .press, .dell, .baseball.</p></blockquote>\n\n<p>I get that there's a <code>&lt;code&gt;</code> and <code>.codes</code>, similarly <code>&lt;picture&gt;</code> and <code>.picture</code> - but what are forum, baseball, and the others doing there? This is just unnecessary verbiage designed to trick the user into thinking the task has been well-researched.</p>\n\n<p>If you want a laugh, <a href=\"https://www.perplexity.ai/search/63736333-20da-4a4b-8807-9d990260296c\">take a look at Perplexity</a> which found 54 matches - most of which were wrong.</p>\n\n<p>Finally, someone asked \"GPT Astra 6 Extra High\" (which is a bonkers bad name for any product). It seemed to get all the elements - and made a note that <a href=\"https://html.spec.whatwg.org/multipage/obsolete.html#non-conforming-features\">two were actually obsolete</a>.</p>\n\n<p>So that's a range of modern models which are either very wrong, slightly wrong, included spurious and incoherent information, or were right.</p>\n\n<p>How do you know which one to choose? How confident are you that the non-determinist computer will always produce the correct answer?</p>\n\n<h2 id=\"hello-computer\"><a href=\"#hello-computer\">Hello Computer</a></h2>\n\n<p>Another AI which got all the correct answers, didn't make anything up, didn't add extraneous information, and didn't use weasel words was\u2026</p>\n\n<p>Siri!</p>\n\n<p>FUCKING SIRI?!?!</p>\n\n<p>How did a glorified Speak 'n' Spell beat all the other AIs?</p>\n\n<img src=\"https://shkspr.mobi/blog/wp-content/uploads/2026/09/siri.webp\" alt=\"Siri warning to check sources and linking to my website.\" width=\"512\" height=\"152\" class=\"aligncenter\">\n\n<p>Oh. It just copied the answers off <a href=\"https://shkspr.mobi/blog/2023/09/false-friends-html-elements-which-are-also-top-level-domains/\">a random idiot's website</a>.</p>\n\n<h2 id=\"the-trick-which-was-hiding-in-plain-site\"><a href=\"#the-trick-which-was-hiding-in-plain-site\">The Trick Which Was Hiding In Plain Site</a></h2>\n\n<p>Note carefully the question.</p>\n\n<blockquote><p>Which TLDs have the same name as valid HTML5 elements?</p></blockquote>\n\n<p>There's a \u2014secret\u2014 and \u2014some would say\u2014 unintuitive type of element. Behold the mighty power of <a href=\"https://developer.mozilla.org/en-US/docs/Web/API/Web_components/Using_custom_elements\">The Custom Element</a>.</p>\n\n<p>Website authors can create their own elements like <code>&lt;my-custom-element&gt;</code> in order to extend the functionality of their site. But you can't go and create any old custom element. You can't have <code>&lt;mobi&gt;</code> or <code>&lt;uk&gt;</code>. No, there are <em>rules for validity</em>.</p>\n\n<p><a href=\"https://html.spec.whatwg.org/multipage/custom-elements.html#valid-custom-element-name\">The rules</a> say that custom elements must start with a lower-case letter, it must not contain any upper-case letters, and it must contain a dash.</p>\n\n<p>And that's the whole game.</p>\n\n<p>There are over <strong>one hundred and fifty</strong> Top Level Domains which match that criteria!</p>\n\n<p>The Hindi top level domain of <code>.\u0915\u0949\u092e</code> is represented in Punycode as <code>xn--11b4c3d</code>. It has been present in the list of TLDs <a href=\"https://www.iana.org/domains/root/db/xn--11b4c3d.html\">for over a decade</a>.</p>\n\n<h3 id=\"write-a-simple-piece-of-js-to-register-a-custom-element\"><a href=\"#write-a-simple-piece-of-js-to-register-a-custom-element\">Write a simple piece of JS to register a custom element.</a></h3>\n\n<p>Paste this in to your console:</p>\n\n<pre><code class=\"language-js\">class Example extends HTMLElement {\n  constructor() {\n    super();\n  }\n}\n\ncustomElements.define('xn--vermgensberatung-pwb', Example);\n</code></pre>\n\n<p>Try it again with a custom element like <code>holiday</code> (which is also a valid TLD) and it will fail with the error \"'holiday' is not a valid custom element name\". Thus it is demonstrated, Punycode TLDs <em>are</em> valid HTML5 elements.</p>\n\n<h2 id=\"one-last-thing\"><a href=\"#one-last-thing\">One Last Thing</a></h2>\n\n<p>Perhaps you think that including custom HTML elements is a cheat. A trick question set by a bitter old man to tarnish the holy name of our new machine gods?</p>\n\n<p>Verily, I submit to you one final heresy.</p>\n\n<p>HTML specifically allows <a href=\"https://html.spec.whatwg.org/multipage/embedded-content-other.html#mathml\">MathML elements</a> in its documents.</p>\n\n<p>That means we can include the following valid elements which are <em>also</em> TLDs: <code>mn</code>, <code>mo</code>, <code>ms</code>, and <code>mtr</code>!</p>\n\n<p>Amusingly, if you go back and <a href=\"https://www.perplexity.ai/search/63736333-20da-4a4b-8807-9d990260296c\">look at the Perplexity answer</a>, after it barfed up a bunch of misinformation, it said:</p>\n\n<blockquote><p>The HTML specification also includes names from embedded vocabularies\u2014<code>&lt;math&gt;</code> from MathML and <code>&lt;svg&gt;</code> from SVG\u2014but <code>.math</code> and <code>.svg</code> are not currently delegated TLDs in the public DNS root.</p></blockquote>\n\n<p>So close and yet so far!</p>\n\n<h2 id=\"youre-right-the-question-is-unfair-and-thats-on-me\"><a href=\"#youre-right-the-question-is-unfair-and-thats-on-me\">You're right, the question <em>is</em> unfair - and that's on me</a></h2>\n\n<p>If you think the original question was unfair, try asking \"<a href=\"https://share.gemini.google/xBoIpdpAqBz2\">Which TLDs have the same name as elements which are valid in an HTML document?</a>\" and see if you get better results.</p>\n\n<p>What precise wording would you use to ensure that a model would get the right answers? What assumptions are you making about how well you understand the problem? At what point do end up writing a thousand-word formal specification?</p>\n\n<h2 id=\"what-does-this-prove-other-than-you-have-too-much-time-on-your-hands-rewrite-to-be-more-friendly-and-professional\"><a href=\"#what-does-this-prove-other-than-you-have-too-much-time-on-your-hands-rewrite-to-be-more-friendly-and-professional\">What does this prove other than you have too much time on your hands? (rewrite to be more friendly and professional)</a></h2>\n\n<p>Let's delve in to the problems.</p>\n\n<ul>\n<li>Most people don't change defaults. Telling people \"you have to fiddle with the settings\" just means the normal experience is rubbish.</li>\n<li>Humans are lazy and won't check outputs. But, crucially, they shouldn't have to! If something markets itself as a genius, why should a human have to hold its hand?</li>\n<li>Sycophantic models make themselves seem less fallible by giving extraneous detail in order to misdirect overworked readers. That is despicable.</li>\n<li>The fast models are no better than they were a year ago. There's no evidence of \"trickle-down intelligence\".</li>\n<li>Some models <em>are</em> better than others! But unless you constantly validate their output, you'll have no real way of knowing which ones are capable of working at a suitable level.</li>\n</ul>\n\n<p>Look, I don't claim this question is as useful or entertaining as <a href=\"https://simonwillison.net/2025/Jun/6/six-months-in-llms/\">Simon Wilson's \"generate an SVG of a pelican riding a bicycle\"</a>. But I do think it is an example of the sort of real-world use-case where LLMs regularly fail.</p>\n\n<p>If I give a list of one thousand different numbers to Excel, I can be sure it'll add them up correctly. If I tell Photoshop to select all red pixels, I can be sure it won't imagine some of the blues are really red.</p>\n\n<p>That's people's mental model of computers - they do boring tasks quickly and accurately.</p>\n\n<p>In my opinion, LLMs are <em>still</em> surprisingly bad - but only if you know what you're looking for and if you can be bothered to check their outputs.</p>\n\n<p>(And, yes, I am <em>still</em> <a href=\"https://shkspr.mobi/blog/2026/07/im-just-so-bored-of-ai/\">just so bored of AI</a>!)</p><img src=\"https://shkspr.mobi/blog/wp-content/themes/edent-wordpress-theme/info/okgo.php?ID=75701&HTTP_REFERER=DOI\" alt width=1 height=1 loading=eager>","doi":"https://doi.org/10.59350/tyba7-qca35","guid":"https://shkspr.mobi/blog/?p=75701","image":"https://shkspr.mobi/blog/wp-content/uploads/2017/11/Confused-Robot.jpg","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1790035200,"rid":"d3tr6-0kb70","summary":"Last year I ran an experiment to test the ability of modern LLMs to correctly answer a relatively straightforward question. Every single one of them got it wrong. Some missed information, some made up false statements, none were right.","tags":["/etc/","AI","Internet","LLM"],"title":"Are LLMs still surprisingly bad at some simple tasks?","updated_at":1790077724,"url":"https://shkspr.mobi/blog/2026/09/are-llms-still-surprisingly-bad-at-some-simple-tasks/","version":"v1"}},{"document":{"authors":[{"contributor_roles":[],"family":"B\u00e4r","given":"Senya"}],"blog":{"authors":null,"community_id":"db0d8909-9e37-46d0-b16c-0551f575e86b","created":1749772800,"current_feed_url":null,"description":"Das Blog der TIB \u2013 Leibniz-Informationszentrum Technik und Naturwissenschaften und Universit\u00e4tsbibliothek","doi":"https://doi.org/10.65527/tib","favicon":"https://rogue-scholar.org/api/communities/db0d8909-9e37-46d0-b16c-0551f575e86b/logo","feed_format":"application/atom+xml","feed_url":"https://blog.tib.eu/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.tib.eu/","issn":null,"language":"deu","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.65527","relative_url":null,"secure":true,"slug":"tib","status":"active","subfield":"1802","title":"TIB-Blog","updated":1790082029,"use_api":true},"blog_name":"TIB-Blog","blog_slug":"tib","content_html":"<a href=\"/2026/07/29/a-new-explore-page-for-kompakkt-find-faster-curate-better/\" class=\"su-button su-button-style-default\" style=\"color:#777;background-color:#eee;border-color:#bfbfbf;border-radius:0px\" target=\"_self\" title=\"english\"><span style=\"color:#777;padding:0px 18px;font-size:14px;line-height:28px;border-color:#f4f4f4;border-radius:0px;text-shadow:none\"> read this article in English</span></a>\n<p class=\"isSelectedEnd\">Wir haben die Explore Page unseres 3D-Viewers grundlegend \u00fcberarbeitet, um die Suche und Organisation von Objekten zu verbessern. Das Update bringt ein neues Design, erweiterte Filterm\u00f6glichkeiten und eine Multi-Selection-Funktion zum gleichzeitigen Hinzuf\u00fcgen mehrerer Objekte zu Sammlungen.</p>\n<h2 class=\"isSelectedEnd\">Neues Design</h2>\n<p class=\"isSelectedEnd\">Die Explore Page wurde visuell neu strukturiert und vereinfacht. Inhalte sind jetzt \u00fcbersichtlicher angeordnet, wodurch relevante 3D-Objekte schneller gefunden werden k\u00f6nnen.</p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-31698\" src=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-1024x482.png\" alt=\"\" width=\"789\" height=\"372\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-1024x482.png 1024w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-300x141.png 300w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-768x362.png 768w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-1536x724.png 1536w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-2048x965.png 2048w\" sizes=\"auto, (max-width: 789px) 100vw, 789px\" /></p>\n<h2>Erweiterte Filterm\u00f6glichkeiten</h2>\n<p>Neue Filter erlauben eine pr\u00e4zisere Eingrenzung der Ergebnisse. Dadurch wird die Recherche in gr\u00f6\u00dferen Datenbest\u00e4nden effizienter und besser kontrollierbar.</p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-large wp-image-31704\" src=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-1024x543.png\" alt=\"\" width=\"800\" height=\"424\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-1024x543.png 1024w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-300x159.png 300w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-768x407.png 768w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-1536x814.png 1536w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-2048x1085.png 2048w\" sizes=\"auto, (max-width: 800px) 100vw, 800px\" /></p>\n<h2 class=\"isSelectedEnd\">Multi-Selection f\u00fcr Sammlungen</h2>\n<p class=\"isSelectedEnd\">Mehrere Objekte k\u00f6nnen nun gleichzeitig ausgew\u00e4hlt und mit einem Schritt zu einer Sammlung hinzugef\u00fcgt werden. Das erleichtert insbesondere das kuratorische Arbeiten und reduziert wiederholte Einzelaktionen.</p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-31700\" src=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-1024x267.png\" alt=\"\" width=\"831\" height=\"217\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-1024x267.png 1024w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-300x78.png 300w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-768x200.png 768w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-1536x401.png 1536w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-2048x534.png 2048w\" sizes=\"auto, (max-width: 831px) 100vw, 831px\" /></p>\n<p>Mit diesen Verbesserungen wird die Explore Page zu einem zentralen Einstiegspunkt f\u00fcr das Entdecken, Vergleichen und Zusammenstellen von 3D-Objekten. Schau es dir an: <a href=\"https://kompakkt.de/explore?locale=en\">Explore</a></p>","doi":"https://doi.org/10.65527/74jce-kx927","guid":"https://blog.tib.eu/?p=31696","image":"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-scaled.png","language":"de","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1776988800,"rid":"sp621-x1s41","summary":"read this article in English Wir haben die Explore Page unseres 3D-Viewers grundlegend \u00fcberarbeitet, um die Suche und Organisation von Objekten zu verbessern. Das Update bringt ein neues Design, erweiterte Filterm\u00f6glichkeiten und eine Multi-Selection-Funktion zum gleichzeitigen Hinzuf\u00fcgen mehrerer Objekte zu Sammlungen. Neues Design Die Explore Page wurde visuell neu strukturiert und vereinfacht.","tags":["OPENNESS","FORSCHUNG & PROJEKTE","Lizenz:CC-BY-4.0-INT","3D Models","3d Viewer"],"title":"Neue Explore Page in Kompakkt: schneller finden, besser sammeln","updated_at":1790069703,"url":"https://blog.tib.eu/2026/04/24/neue-explore-page-in-kompakkt-schneller-finden-besser-sammeln/","version":"v1"}},{"document":{"authors":[{"contributor_roles":[],"family":"Bailly","given":"Kolja"}],"blog":{"authors":null,"community_id":"db0d8909-9e37-46d0-b16c-0551f575e86b","created":1749772800,"current_feed_url":null,"description":"Das Blog der TIB \u2013 Leibniz-Informationszentrum Technik und Naturwissenschaften und Universit\u00e4tsbibliothek","doi":"https://doi.org/10.65527/tib","favicon":"https://rogue-scholar.org/api/communities/db0d8909-9e37-46d0-b16c-0551f575e86b/logo","feed_format":"application/atom+xml","feed_url":"https://blog.tib.eu/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.tib.eu/","issn":null,"language":"deu","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.65527","relative_url":null,"secure":true,"slug":"tib","status":"active","subfield":"1802","title":"TIB-Blog","updated":1790082029,"use_api":true},"blog_name":"TIB-Blog","blog_slug":"tib","content_html":"<p data-start=\"146\" data-end=\"301\">Wir freuen uns, bekanntzugeben, dass <a href=\"https://youtu.be/JmEhXB21_YE?si=hRc7X8hEqeo1j2sh\" target=\"_blank\" rel=\"noopener\">Semantic Wikibase</a> erfolgreich auf Kompatibilit\u00e4t mit <a href=\"https://www.mediawiki.org/wiki/MediaWiki\" target=\"_blank\" rel=\"noopener\"><span class=\"hover:entity-accent entity-underline inline cursor-pointer align-baseline\"><span class=\"whitespace-normal\">MediaWiki</span></span></a> 1.43 aktualisiert wurde. Mit diesem Schritt stellen wir sicher, dass Semantic Wikibase weiterhin mit der aktuellen<a href=\"https://de.wikipedia.org/wiki/Support_(Dienstleistung)#Long_Term_Support\" target=\"_blank\" rel=\"noopener\"> Longterm-Support-Version</a> von MediaWiki kompatibel bleibt und als stabile Grundlage f\u00fcr semantisch angereicherte Wissensinfrastrukturen dient.</p>\n<h2 data-start=\"146\" data-end=\"301\">\u00dcber Semantic Wikibase</h2>\n<p data-start=\"303\" data-end=\"539\">Viele Forschungsprojekte setzen das Mediawiki-Framework als Werkzeug f\u00fcr Forschungsdatenmanagement ein. Mit \u00fcber 1.500 Erweiterungen l\u00e4sst sich dieses an die individuellen Anforderungen anpassen:</p>\n<ul>\n<li data-start=\"303\" data-end=\"539\">als reines Wiki mit Text und Medien, organisiert in Artikelseiten nach dem Vorbild von Wikipedia,</li>\n<li data-start=\"303\" data-end=\"539\">als strukturierte Wissens-Datenbank zur Linked-Open-Data Implementierung von Wissensgraphen und Terminologien mittels <a href=\"https://wikiba.se/\" target=\"_blank\" rel=\"noopener\">Wikibase,</a></li>\n<li data-start=\"303\" data-end=\"539\">als semantischer Wissensspeicher zur Datenvisualisierung mittels <a href=\"https://www.semantic-mediawiki.org/wiki/Semantic_MediaWiki\" target=\"_blank\" rel=\"noopener\">Semantic Mediawiki</a>.</li>\n</ul>\n<h3>Semantic Mediawiki vs. Wikibase</h3>\n<p>Insbesondere Wikibase und Semantic Mediawiki werden h\u00e4ufig im Forschungsumfeld verwendet. Beide Erweiterungen haben <strong>unterschiedliche St\u00e4rken und Schw\u00e4chen</strong>:</p>\n<figure id=\"attachment_31116\" aria-describedby=\"caption-attachment-31116\" style=\"width: 772px\" class=\"wp-caption alignnone\"><a href=\"https://de.slideshare.net/slideshow/semantic-mediawiki-a-linked-open-data-platform/272328819\" target=\"_blank\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-31116 size-full\" src=\"https://blog.tib.eu/wp-content/uploads/2026/03/Wikibase-vs-SMW.jpg\" alt=\"\" width=\"772\" height=\"322\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/03/Wikibase-vs-SMW.jpg 772w, https://blog.tib.eu/wp-content/uploads/2026/03/Wikibase-vs-SMW-300x125.jpg 300w, https://blog.tib.eu/wp-content/uploads/2026/03/Wikibase-vs-SMW-768x320.jpg 768w\" sizes=\"auto, (max-width: 772px) 100vw, 772px\" /></a><figcaption id=\"caption-attachment-31116\" class=\"wp-caption-text\">Vergleich von Wikibase and SMW (Grafik by Bernhard Krabina)</figcaption></figure>\n<h3>Semantic Mediawiki und Wikibase</h3>\n<p>Die Entwicklung von Semantic Wikibase (SWB) erm\u00f6glichte es erstmals, beide Erweiterungen gemeinsam auf einem System zu verbinden und so die <strong>Vorteile beider Systeme gemeinsam zu nutzen</strong>. W\u00e4hrend strukturelle Wissensdaten in Wikibase gespeichert und verwaltet werden, sorgt die SWB-Erweiterung daf\u00fcr, dass diese auch in Semantic Mediawiki f\u00fcr die Visualisierung in Wiki-Artikeln verf\u00fcgbar sind. SWB dient also quasi als Br\u00fccke zwischen den beiden Erweiterungen, wobei der Datenfluss nur von Wikibase nach Semantic Mediawiki (nicht umgekehrt) erfolgt. Dies dient dazu, Datenkonflikte zu vermeiden.</p>\n<p>Semantic Wikibase wurde im September 2020 in einer <a href=\"https://professional.wiki/en/articles/semantic-wikibase\" target=\"_blank\" rel=\"noopener\">ersten Version</a> vom Unternehmen <a href=\"https://professional.wiki/\" target=\"_blank\" rel=\"noopener\">ProfessionalWiki</a> ver\u00f6ffentlicht. Dieser erste Prototyp war nur mit der \u00e4lteren Mediawiki Version 1.35 kompatibel, aber unterst\u00fctzte bereits grundlegende Datentypen. Im Open Science Lab sahen wir in der Entwicklung einen Baustein, der das Potenzial hat, im Mediawiki-Umfeld eine bedeutende L\u00fccke zu schlie\u00dfen: Die <strong>Kombination aus strukturierter, f\u00f6derierbarer Datenverwaltung und Datenpr\u00e4sentation</strong>. Unser Ziel war es, die Erweiterung zu testen, bei Bedarf weiterzuentwickeln und k\u00fcnftig als unser <a href=\"https://de.wikipedia.org/wiki/Content-Management-System\" target=\"_blank\" rel=\"noopener\">Content-Management-System</a> zur Unterst\u00fctzung von Forschungsprojekten zu verwenden.</p>\n<h3>Case Studies</h3>\n<p>Mitte 2024 wurde mit dem Projekt <a href=\"https://phiwikibase4research-staging.adwmainz.net/index.php/PhiWiki:%C3%9Cber_PhiWiki\" target=\"_blank\" rel=\"noopener\">PhiWiki</a> ein erster Prototyp f\u00fcr Mediawiki 1.39 in Zusammenarbeit mit der <a href=\"https://www.adwmainz.de/\" target=\"_blank\" rel=\"noopener\">Akademie der Wissenschaften und der Literatur Mainz</a> sowie der <a href=\"https://digitale-philosophie.de/\" target=\"_blank\" rel=\"noopener\">AG Digitale Philosophie</a> erfolgreich getestet. Es folgte mit <a href=\"https://climatekg.semanticclimate.net/index.php?title=Hauptseite\" target=\"_blank\" rel=\"noopener\">Semantic Glossar</a> ein weiteres Projekt zur kollaborativen Entwicklung von Terminologien mittels Semantic Wikibase.</p>\n<p>Ende 2024 konnten wir im Rahmen des Projekts <a href=\"https://wb.manorhouses.tibwiki.io/wiki/Hauptseite\" target=\"_blank\" rel=\"noopener\">Herrenh\u00e4user des Ostseeraums</a> Semantic Wikibase dann in einem umfangreichen Projekt einem herausfordernden <a href=\"https://professional.wiki/en/news/connecting-wikibase-and-semantic-mediawiki\" target=\"_blank\" rel=\"noopener\">Lasttest</a> unterziehen. Mit \u00fcber 14.000 Wikibase-Objekten, die auf mehr als 300 Artikelseiten dynamisch eingebettet als Karten, Zeitstrahlen, Tabellen und Suchformulare verwendet werden, konnten wir die bestehenden Schw\u00e4chen von Semantic Wikibase identifizieren und beheben. Dazu geh\u00f6rte unter anderem die Unterst\u00fctzung des vollen Wikibase-Datenmodells mittels <a href=\"https://www.wikidata.org/wiki/Help:Qualifiers\" target=\"_blank\" rel=\"noopener\">Qualifiers</a>, eine erste grundlegende Unterst\u00fctzung des <a href=\"https://www.loc.gov/standards/datetime/\" target=\"_blank\" rel=\"noopener\">Extended Datetime Formats</a> (EDTF) sowie die Einbettung von 3D-Visualisierungen aus <a href=\"https://semantic-kompakkt.de/home?locale=en\" target=\"_blank\" rel=\"noopener\">Semantic Kompakkt</a>. Entscheidend war hierf\u00fcr die<strong>\u00a0interdisziplin\u00e4re Zusammenarbeit</strong> zwischen dem Enwicklerteam und den LOD- und Wikibase-Datenmodell-Expertinnen Lozana Rossenova und Lucia Sohmen.</p>\n<p>Die im Projekt entwickelten Best-Practices umfassten unter anderem:</p>\n<ul>\n<li>Nutzung individueller Formulare f\u00fcr Suchfilter und Dateneingabe</li>\n<li>Verkn\u00fcpfung von Wikibase-Items mit Mediawiki-Kategorien</li>\n<li>Verlinkung von Artikelseiten mit Wikibase-Items</li>\n<li>Nutzung des vollen Wikibase-Datenmodells in <a href=\"https://www.semantic-mediawiki.org/wiki/Help:Inline_queries\" target=\"_blank\" rel=\"noopener\">SMW Inline Queries</a></li>\n<li>Performance von <a href=\"https://www.semantic-mediawiki.org/wiki/Help:Result_formats\" target=\"_blank\" rel=\"noopener\">Datenvisualisierungen</a> trotz hoher Anzahl an Wikibase-Objekten</li>\n<li>Best Practices zur Informationsmodellierung im <a href=\"https://gitlab.com/nfdi4culture/wikibase4research/auxiliary-service-repositories/wikibase-model\" target=\"_blank\" rel=\"noopener\">Generic Wikibase Model for Cultural Data</a></li>\n</ul>\n<h2 data-start=\"541\" data-end=\"576\">Warum MediaWiki 1.43 wichtig ist</h2>\n<p data-start=\"578\" data-end=\"879\">Mit der Version 1.39 war Semantic Wikibase kompatibel mit der damaligen <a href=\"https://de.wikipedia.org/wiki/Support_(Dienstleistung)#Long_Term_Support\" target=\"_blank\" rel=\"noopener\">Longtime Support Version(LTS)</a> von Mediawiki. Diese Unterst\u00fctzung war aber gem\u00e4\u00df des <a href=\"https://www.mediawiki.org/wiki/Version_lifecycle/de\" target=\"_blank\" rel=\"noopener\">Mediawiki Lifecycle</a> nur bis Ende 2025 gegeben.</p>\n<p data-start=\"578\" data-end=\"879\">MediaWiki 1.43 bringt als aktuelle LTS-Version (Support bis 2028) zahlreiche technische Verbesserungen, Performance-Optimierungen sowie langfristige Wartungsvorteile mit sich. F\u00fcr viele Wikibase-Installationen ist die Orientierung an den aktuellen MediaWiki-Versionen essenziell, um Sicherheit, Stabilit\u00e4t und Zukunftsf\u00e4higkeit zu gew\u00e4hrleisten. Durch Versionskonflikte zwischen verwendeten Bibliotheken in Wikibase und Semantic Mediawiki, konnte SemanticWikibase aber nicht ohne Anpassung in dieser neuen Version eingesetzt werden.</p>\n<p data-start=\"578\" data-end=\"879\"><strong>Unsere gr\u00f6\u00dfte Bef\u00fcrchtung</strong> war, dass die aktuellen Versionen grundlegende \u00c4nderung vorgenommen hatten, die einen Weiterbetrieb von Semantic Wikibase technisch unsauber bzw. unwirtschaftlich machen w\u00fcrden. Ende 2025 schaffte Open-Science-Lab-Entwickler Lukas G\u00fcnther die entscheidende Grundlage f\u00fcr das Upgrade, indem er unser Installationstool <a href=\"https://gitlab.com/nfdi4culture/wikibase4research/wikibase4research\" target=\"_blank\" rel=\"noopener\">Wikibase4Research</a> aktualisierte und so mit der Mediawiki Version 1.44 kompatibel machte. Da Semantic Wikibase sich mittels Wikibase4Research automatisiert installieren l\u00e4sst, war so ein geeignetes Test-Setup geschaffen, um die Entwicklung in Angriff zu nehmen. Letzendlich war es uns so m\u00f6glich, Semantic Wikibase mit der aktuellen LTS-Version von Mediawiki zu betreiben und das sogar ohne \u00c4nderungen am Wikibase- oder SemanticMediawiki-Code vorzunehmen. S\u00e4mtliche bisher unterst\u00fctzten Datentypen sind auch weiterhin funktional, was auch ein Update bestehender Installationen auf die neue Version erm\u00f6glicht.</p>\n<figure id=\"attachment_31121\" aria-describedby=\"caption-attachment-31121\" style=\"width: 520px\" class=\"wp-caption alignnone\"><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-31121\" src=\"https://blog.tib.eu/wp-content/uploads/2026/03/SWB-Datatypes.png\" alt=\"\" width=\"520\" height=\"442\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/03/SWB-Datatypes.png 520w, https://blog.tib.eu/wp-content/uploads/2026/03/SWB-Datatypes-300x255.png 300w\" sizes=\"auto, (max-width: 520px) 100vw, 520px\" /><figcaption id=\"caption-attachment-31121\" class=\"wp-caption-text\">Unterst\u00fctzte Datentypen in Semantic Wikibase, visualisiert im Semantic Browser von SMW</figcaption></figure>\n<h2 data-start=\"2223\" data-end=\"2234\">Ausblick</h2>\n<p data-start=\"2236\" data-end=\"2552\">Die kontinuierliche Synchronisierung von Semantic Wikibase mit dem MediaWiki-Releasezyklus ist ein zentraler Baustein f\u00fcr nachhaltige, semantische Wissensinfrastrukturen. Mit diesem Update schaffen wir die Grundlage f\u00fcr kommende Weiterentwicklungen und eine langfristig stabile Integration in das Wikibase-\u00d6kosystem. Der Einsatz von Semantic Wikibase bedeutet f\u00fcr unsere Forschungsdaten- und Terminologie-Projekte im Open Science Lab:</p>\n<ul>\n<li data-start=\"2236\" data-end=\"2552\">Fokussierung auf eine gemeinsame technologische Basis f\u00fcr alle Projekte</li>\n<li data-start=\"2236\" data-end=\"2552\">B\u00fcndelung von Wissen und Ressourcen</li>\n<li data-start=\"2236\" data-end=\"2552\">Zeitersparnis bei der Projektumsetzung durch Best Practices und Synergieeffekten zwischen Projekten</li>\n<li data-start=\"2236\" data-end=\"2552\">Koordinierter Aufbau von Services innerhalb eines bestehenden Software \u00d6kosystems</li>\n<li data-start=\"2236\" data-end=\"2552\">Support der Open-Source und Linked-Open-Data Community durch unsere Entwicklungen</li>\n</ul>\n<p><strong>Wir freuen uns auf die weitere Entwicklung und die vielf\u00e4ltigen kommenden Projekte mit Semantic Wikibase.</strong></p>\n<h5>Relevante Links</h5>\n<ul>\n<li><a href=\"https://gitlab.com/nfdi4culture/wikibase4research/wikibase4research\" target=\"_blank\" rel=\"noopener\">Wikibase4Research</a></li>\n<li><a href=\"https://gitlab.com/nfdi4culture/wikibase4research/auxiliary-service-repositories/wikibase-model\" target=\"_blank\" rel=\"noopener\">Generic Wikibase Model for Cultural Data</a></li>\n</ul>","doi":"https://doi.org/10.65527/2cdwj-mqe46","guid":"https://blog.tib.eu/?p=31114","image":"https://blog.tib.eu/wp-content/uploads/2026/04/Wikibase_logo.svg.png","language":"de","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1775692800,"rid":"kfdwg-qff93","summary":"Wir freuen uns, bekanntzugeben, dass Semantic Wikibase erfolgreich auf Kompatibilit\u00e4t mit MediaWiki 1.43 aktualisiert wurde. Mit diesem Schritt stellen wir sicher, dass Semantic Wikibase weiterhin mit der aktuellen Longterm-Support-Version von MediaWiki kompatibel bleibt und als stabile Grundlage f\u00fcr semantisch angereicherte Wissensinfrastrukturen dient.","tags":["OPENNESS","FORSCHUNG & PROJEKTE","Lizenz:CC-BY-4.0-INT","Open Science Lab","Wikibase"],"title":"Upgrade abgeschlossen: Semantic Wikibase kompatibel mit MediaWiki 1.43","updated_at":1790067918,"url":"https://blog.tib.eu/2026/04/09/upgrade-abgeschlossen-semantic-wikibase-kompatibel-mit-mediawiki-1-43/","version":"v1"}},{"document":{"authors":[{"contributor_roles":[],"family":"Bailly","given":"Kolja"}],"blog":{"authors":null,"community_id":"db0d8909-9e37-46d0-b16c-0551f575e86b","created":1749772800,"current_feed_url":null,"description":"Das Blog der TIB \u2013 Leibniz-Informationszentrum Technik und Naturwissenschaften und Universit\u00e4tsbibliothek","doi":"https://doi.org/10.65527/tib","favicon":"https://rogue-scholar.org/api/communities/db0d8909-9e37-46d0-b16c-0551f575e86b/logo","feed_format":"application/atom+xml","feed_url":"https://blog.tib.eu/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.tib.eu/","issn":null,"language":"deu","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.65527","relative_url":null,"secure":true,"slug":"tib","status":"active","subfield":"1802","title":"TIB-Blog","updated":1790082029,"use_api":true},"blog_name":"TIB-Blog","blog_slug":"tib","content_html":"<p>KI-Systeme, die Texte nicht nur generieren, sondern gezielt in Dokumenten recherchieren, sind mittlerweile etablierter Stand der Technik. Einer dieser Ans\u00e4tze hei\u00dft <a href=\"https://en.wikipedia.org/wiki/Retrieval-augmented_generation\" target=\"_blank\" rel=\"noopener\">Retrieval-Augmented Generation (RAG)</a>: Stellt ein Benutzer eine Frage, sucht das System relevante Informationen in einer Wissensbasis \u2013 zum Beispiel in einem Wiki \u2013 und nutzt diese als Grundlage, um relevante Inhalte bzw. Quellen aufzulisten oder mittels KI Antworten daraus zu generieren.</p>\n<p><strong>Das Problem</strong>: Damit ein solches System gut funktioniert, m\u00fcssen viele Stellschrauben richtig eingestellt werden. Diese sogenannte <a href=\"https://en.wikipedia.org/wiki/Hyperparameter_optimization\" target=\"_blank\" rel=\"noopener\">Hyperparameter-Optimierung</a> ist normalerweise entweder zeitaufw\u00e4ndig oder rechenintensiv und in jedem Fall technisch anspruchsvoll. Unsere aktuelle Untersuchung zeigt jedoch: <strong>Eine automatisierte Optimierung ist m\u00f6glich \u2013 sogar auf einem normalen Laptop</strong>.</p>\n<h2>Ausgangslage</h2>\n<p>Grundlage unserer Untersuchung im <a href=\"https://www.tib.eu/de/forschung-entwicklung/forschungsgruppen-und-labs/open-science\" target=\"_blank\" rel=\"noopener\"><strong>Open Science Lab</strong></a> war die Weiterentwicklung unseres RAG-Moduls f\u00fcr <a href=\"https://blog.tib.eu/2024/08/29/wikibase4research-wissensdaten-einfach-verwalten-teilen-und-visualisieren/\" target=\"_blank\" rel=\"noopener\"><strong>Wikibase4Research</strong></a>. Mit dem zuvor bestehenden System war es bereits sehr einfach m\u00f6glich, eine Mediawiki Installation zu erhalten, deren Inhalte KI-gest\u00fctzt via RAG durchsuchbar sind. Egal ob es nun um Artikelseiten in einem einfachen <a href=\"https://www.mediawiki.org/wiki/MediaWiki\" target=\"_blank\" rel=\"noopener\">Mediawiki</a>, strukturierte Wissensdaten in einer <a href=\"https://wikiba.se/\" target=\"_blank\" rel=\"noopener\">Wikibase</a> oder eine Kombination aus beidem wie zum Beispiel <a href=\"https://www.semantic-mediawiki.org/wiki/Semantic_MediaWiki\" target=\"_blank\" rel=\"noopener\">Semantic Mediawiki</a> oder <a href=\"https://www.mediawiki.org/wiki/Extension:Semantic_Wikibase\" target=\"_blank\" rel=\"noopener\">Semantic Wikibase</a> geht.</p>\n<p>Eine Einf\u00fchrung in die grundlegende Funktionsweise von RAG und Wikibase4Research liefert das folgende Video:</p>\n<div class=\"ratio ratio-16x9\"><iframe loading=\"lazy\" title=\"Kolja Bailly: From Triples to Text: LLM(RAG)-Based Approach to Querying Wikibase | #MUDCon Fall 2025\" width=\"800\" height=\"450\" src=\"https://www.youtube.com/embed/GIZA4OVogLc?feature=oembed\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share\" referrerpolicy=\"strict-origin-when-cross-origin\" allowfullscreen></iframe></div>\n<p>Um eine hohe Qualit\u00e4t der KI-basierten Suchergebnisse und Antworten zu erhalten, ist es aber n\u00f6tig, das System <strong>entsprechend der verwendeten Daten zu konfigurieren</strong>. F\u00fcr diese Einstellungen gibt es keine Standardf\u00e4lle, es geh\u00f6rt in das Arbeitsfeld eines Data Scientist die Systemparameter zu testen und zu verbessern. In diesem Prozess wird daher klassisch ein hohes Ma\u00df an Erfahrung und Fachwissen ben\u00f6tigt, um optimale Ergebnisse zu erhalten.</p>\n<p>Die Alternative ist der nun in Wikibase4Research integrierte <strong>AutoRAG</strong> Ansatz, der die Parameter vollautomatisch optimiert. Dieser Prozess wird im Farchjargon \"Hyperparameter Tuning\" oder auch \"Hyperparameter Optimierung\" genannt.</p>\n<h2>Anforderungen</h2>\n<p>Die Rahmenbedingungen f\u00fcr ein Hyperparameter Tuning k\u00f6nnen sehr unterschiedlich sein. In unserem Fall ergeben sich die Anforderungen vor allem aus der Nutzergruppe von Wikibase4Research.</p>\n<h3>Forscher/Innen</h3>\n<p>Im Forschungskontext haben wir es mit f\u00e4cherspezifischen Daten zu tun. Die beteiligten Wissenschaftler sind Experten in ihrer jeweiligen Fachdom\u00e4ne. Expertise im Bereich spezieller Data-Science-Anwendungen ist in den Projektteams meist nicht vorhanden. Dies ist durchaus sinnvoll, denn das Projektteam ist somit auf die im Projekt zu bearbeitenden Forschungsfragen spezialisiert.</p>\n<h3>Daten</h3>\n<p>F\u00fcr die Optimierung wird ein Test-Datensatz ben\u00f6tigt, der m\u00f6gliche Fragen (Suchanfragen) mit den optimalen Quellen in den Daten verkn\u00fcpft. Dieser Datensatz wird mit den Suchergebnissen des Systems verglichen, um die Qualit\u00e4t der Systemeinstellung bewerten zu k\u00f6nnen (Idealdaten). Solche Testdaten liegen in den \u00fcberwiegenden F\u00e4llen nicht vor.</p>\n<h3>Endnutzer/Innen</h3>\n<p>Wer nutzt die Daten letztendlich und welche Art von Anfragen werden gestellt? Diese Frage ist entscheidend bei der Optimierung. Werden die Endnutzer spezifische Fakten aus den Daten abfragen wie zum Beispiel Jahreszahlen bestimmter Ereignisse oder eher Zusammenfassungen ganzer Abs\u00e4tze oder Artikel erwarten? Zu welchen Themen werden voraussichtlich Fragen gestellt? Erwarte ich eher Fragen zum Inhalt der Daten oder Fragen auf der Metaebene wie zum Beispiel zur Anzahl von Quellen, der Struktur und L\u00e4nge von Texten, des Schreibstils oder zur Medienart? Werden Suchanfragen von Wissenschaftlern im Fachjargon gestellt oder eher in Umgangssprache formuliert? <span style=\"font-family: inherit;\">Die fr\u00fchzeitige Definition grundlegender </span><a style=\"font-family: inherit;\" href=\"https://en.wikipedia.org/wiki/Persona_(user_experience)\" target=\"_blank\" rel=\"noopener\">Personas</a><span style=\"font-family: inherit;\"> f\u00fcr die zu erwartende Nutzergruppe hilft nicht nur bei der Optimierung von RAG, sondern ist auch ein wichtiger Schritt bei der Erstellung von Design und Benutzeroberfl\u00e4chen in der Pr\u00e4sentation der Forschungsergebnisse.</span></p>\n<h3>Infrastruktur</h3>\n<p>Hohe Rechenkapazit\u00e4ten, Zugang zu GPU-Processing und Budget f\u00fcr industrielle KI-Services ist in vielen Projekten nicht vorhanden. Wikibase4Research bietet die Option, externe Schnittstellen wie <a href=\"https://huggingface.co/models\" target=\"_blank\" rel=\"noopener\">Huggingface</a>, <a href=\"https://developers.openai.com/api/reference\" target=\"_blank\" rel=\"noopener\">OpenAI</a> oder die <a href=\"https://docs.hpc.gwdg.de/services/saia/index.html\" target=\"_blank\" rel=\"noopener\">SAIA-Umgebung</a> der <a href=\"https://gwdg.de/\" target=\"_blank\" rel=\"noopener\">GWDG</a> zur Ausf\u00fchrung von KI-Modellen zu nutzen. Die dort bestehenden Limits f\u00fcr kostenlose Nutzung reichen aber meist nicht aus, um die Vielzahl an Parameter-Konfigurationen zu testen, die zur Optimierung eines RAG-Systems notwendig ist. Ideal w\u00e4re also, die Ausf\u00fchrung lokal auf allgemein verf\u00fcgbarer Hardware durchf\u00fchren zu k\u00f6nnen, was auch unter dem Aspekt der ressourcenschonenden Nutzung von KI ein erstrebenswertes Ziel ist.</p>\n<p><strong>Es ergibt sich f\u00fcr unseren Ansatz daher folgender Anforderungskatalog:</strong></p>\n<ul>\n<li>Anpassung auf die verwendeten Daten</li>\n<li>vollautomatische Optimierung</li>\n<li>keine technischen Vorkenntnisse n\u00f6tig</li>\n<li>Test-Datensatz wird generiert</li>\n<li>User-Persona-Profile ber\u00fccksichtigen</li>\n<li>m\u00f6glichst effizient, mit geringem Ressourcenbedarf</li>\n</ul>\n<h2>Methodik</h2>\n<h3>Daten</h3>\n<p>Als Datengrundlage dienten jeweils 50 zuf\u00e4llige Artikel aus drei MediaWiki-basierten Wissenssammlungen:</p>\n<ul>\n<li><a href=\"https://de.wikipedia.org/\" target=\"_blank\" rel=\"noopener\">Wikipedia</a></li>\n<li><a href=\"https://wb.manorhouses.tibwiki.io/wiki/Deutschland\" target=\"_blank\" rel=\"noopener\">Herrenh\u00e4user</a></li>\n<li><a href=\"https://kungfu-wiki.com/Scholar_indexpage\" target=\"_blank\" rel=\"noopener\">Kungfu-Wiki</a></li>\n</ul>\n<p>Um die Qualit\u00e4t der Suche zu bewerten, wurden automatisch Frage-Kontext-Antwort-Tripel erzeugt. Zum Einsatz kam daf\u00fcr das mehrsprachige Sprachmodell <a href=\"https://www.ibm.com/de-de/new/announcements/ibm-granite-4-0-hyper-efficient-high-performance-hybrid-models\" target=\"_blank\" rel=\"noopener\"><strong>IBM Granite 4 350M Nano</strong></a>, das speziell f\u00fcr Umgebungen mit geringer Rechenleistung wie zum Beispiel f\u00fcr On-Device-Anwendungsf\u00e4lle entwickelt wurde.</p>\n<h3>LLM-Prompt</h3>\n<p>Um hinsichtlich der erwarteten Nutzung realistische Fragen zu generieren, wurde der an das Modell gelieferte Prompt (\"Erstelle Fragen aus dem Seiteninhalt\") um speziell angepasste Rollenbeschreibungen (Personas) erg\u00e4nzt, die per Konfigurationsdatei individualisiert werden k\u00f6nnen. Eine solche Persona-Definition k\u00f6nnte zum Beispiel lauten: \"You are a scientist who wants to learn about historic manorhouses in Europe\".</p>\n<h3>Parameter</h3>\n<p>In einem RAG-Prozess werden die zu durchsuchenden Daten in einer speziellen Datenbank indiziert, um sp\u00e4ter schnell und effizient relevante Inhalte zu finden.</p>\n<figure id=\"attachment_30955\" aria-describedby=\"caption-attachment-30955\" style=\"width: 847px\" class=\"wp-caption alignnone\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-30955 \" src=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag1-1024x240.jpg\" alt=\"\" width=\"847\" height=\"199\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag1-1024x240.jpg 1024w, https://blog.tib.eu/wp-content/uploads/2026/02/rag1-300x70.jpg 300w, https://blog.tib.eu/wp-content/uploads/2026/02/rag1-768x180.jpg 768w, https://blog.tib.eu/wp-content/uploads/2026/02/rag1.jpg 1440w\" sizes=\"auto, (max-width: 847px) 100vw, 847px\" /><figcaption id=\"caption-attachment-30955\" class=\"wp-caption-text\">Information Extraction und Indizierung von Daten in einem RAG-Prozess</figcaption></figure>\n<p>Die meisten von uns verwendeten Parameter optimieren diesen Prozess der\u00a0<a href=\"https://en.wikipedia.org/wiki/Information_extraction\" target=\"_blank\" rel=\"noopener\">Informations Extraktion</a> (IE). Dabei wird bestimmt, in welcher Form die Daten gespeichert werden und ob diese ggf. vor dem Speichern um Metadaten wie Schlagworte, Titel oder Zusammenfassungen erg\u00e4nzt werden. F\u00fcr die Vektorisierung verwendeten wir das Modell <a href=\"https://github.com/QwenLM/Qwen3-Embedding\" target=\"_blank\" rel=\"noopener\"><strong>Qwen3-embedding:0.6B</strong></a>. Die mittels AutoRAG optimierten Parameter sind im Folgenden aufgelistet:</p>\n<ul>\n<li><strong>Chunk_Size</strong>: Wie gro\u00df sind die Informationsabschnitte, die sp\u00e4ter zugreifbar sein sollen?</li>\n<li><strong>Chunk_Overlap</strong>: Wie stark \u00fcberlappen sich die Informationsabschnitte?</li>\n<li><strong>Extractors</strong>: Welche Datenanreicherungen sollen erfolgen (zum Beispiel Zusammenfassung erstellen, Fragen generieren)?</li>\n<li><strong>Top_K:\u00a0</strong>Wieviele Chunks werden als Suchergebnis geliefert?</li>\n</ul>\n<p>Sind die Daten eingelesen und wird eine Suchanfrage gestellt, wird das System nach relevanten Informationsabschnitten durchsucht. Dieser Prozess wird \"Information Retrieval\" genannt. Man kann es mit den Ergebnissen einer Google-Suche vergleichen, bei der die relevantesten Ergebnisse nicht zwangsl\u00e4ufig an erster Stelle der Liste stehen.</p>\n<figure id=\"attachment_30957\" aria-describedby=\"caption-attachment-30957\" style=\"width: 800px\" class=\"wp-caption alignnone\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-30957 size-large\" style=\"font-family: inherit;\" src=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag2-1024x273.jpg\" alt=\"\" width=\"800\" height=\"213\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag2-1024x273.jpg 1024w, https://blog.tib.eu/wp-content/uploads/2026/02/rag2-300x80.jpg 300w, https://blog.tib.eu/wp-content/uploads/2026/02/rag2-768x204.jpg 768w, https://blog.tib.eu/wp-content/uploads/2026/02/rag2.jpg 1360w\" sizes=\"auto, (max-width: 800px) 100vw, 800px\" /><figcaption id=\"caption-attachment-30957\" class=\"wp-caption-text\">Information Retrieval in einem RAG Prozess</figcaption></figure>\n<p>Information Retrieval bedeutet, zur Frage des Nutzers relevante Informationen zu finden. In diesem Prozessschritt optimieren wir den Parameter \"<strong>Top_K\"</strong>, der definiert, wie viele der Suchergebnisse im weiteren Prozess ber\u00fccksichtigt werden. Ist Top_K zu klein, sind wichtige Quellen eventuell nicht enthalten. Ist Top_K zu gro\u00df, verarbeitet man eventuell eine gro\u00dfe Menge wenig relevanter Inhalte.</p>\n<h3>Optimierungsverfahren</h3>\n<p>Statt alle m\u00f6glichen Kombinationen auszuprobieren (was sehr lange dauern w\u00fcrde), kommt ein Suchalgorithmus zum Einsatz, der die verschiedenen Parameter stufenweise verbessert. Dieses als <a href=\"https://en.wikipedia.org/wiki/Greedy_algorithm\" target=\"_blank\" rel=\"noopener\">Greedy</a> (\"gierig<strong>\"</strong>) benannte Verfahren optimiert zun\u00e4chst nur einen einzigen Parameter, dann den n\u00e4chsten usw. Wir verzichten damit auf optimale L\u00f6sungen, erreichen aber hinreichend <strong>gute Ergebnisse mit akzeptablem Aufwand</strong>.</p>\n<p>Als Bewertungsma\u00df f\u00fcr die Optimierung dient dabei der sogenannte <a href=\"https://en.wikipedia.org/wiki/Mean_reciprocal_rank\" target=\"_blank\" rel=\"noopener\">Mean Reciprocal Rank (MRR)</a> \u2013 ein Ma\u00df daf\u00fcr, an welcher Position relevante Inhalte in der Trefferliste platziert sind. Ein entscheidender Vorteil:<br />\nDie Bewertung erfolgt vollst\u00e4ndig ohne KI-Antwortgenerierung. Es wird also nur getestet, wie gut das System relevante Inhalte findet, nicht wie gut eine KI daraus sp\u00e4ter Antworten generiert. Dadurch wird erheblich Rechenzeit gespart.</p>\n<figure id=\"attachment_30956\" aria-describedby=\"caption-attachment-30956\" style=\"width: 710px\" class=\"wp-caption alignnone\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-30956\" src=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag3.jpg\" alt=\"\" width=\"710\" height=\"256\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag3.jpg 919w, https://blog.tib.eu/wp-content/uploads/2026/02/rag3-300x108.jpg 300w, https://blog.tib.eu/wp-content/uploads/2026/02/rag3-768x277.jpg 768w\" sizes=\"auto, (max-width: 710px) 100vw, 710px\" /><figcaption id=\"caption-attachment-30956\" class=\"wp-caption-text\">Antwort Generierung in einem RAG Prozess. Diese Phase wurde in der Optimierung NICHT ber\u00fccksichtigt</figcaption></figure>\n<h2>Technische Umsetzung</h2>\n<p>Die Implementierung erfolgte vollst\u00e4ndig im MediaWiki-Umfeld mit:</p>\n<ul>\n<li>Wikibase4Research</li>\n<li>einer Docker-basierten Python-API</li>\n<li>dem RAG-Framework LlamaIndex</li>\n<li>lokaler Modellbereitstellung \u00fcber Ollama</li>\n</ul>\n<p>Die Experimente liefen auf einem <strong>handels\u00fcblichen Laptop</strong> aus dem Jahr 2022 (Dell Latitude 5421, Intel Core i7-11850H mit 8 Kernen, 16 GB RAM) \u2013 ohne GPU-Beschleunigung.</p>\n<h2>Ergebnisse</h2>\n<p>Trotz der bewusst schlanken Hardware-Ausstattung konnte die Optimierung meist bereits innerhalb einer Stunde abgeschlossen werden. Dabei wurde bei allen Datens\u00e4tzen eine starke Verbesserungen der Abfrageergebnisse erzielt.</p>\n<p>F\u00fcr unser Qualit\u00e4stma\u00df, den Mean Reciprocal Rank (MRR), ergab sich eine <strong>Steigerung von durchschnittlich 12 bis 25 Prozent</strong> gegen\u00fcber den voreingestellten Parametern. Das bedeutet, in den Ergebnissen der Suchanfrage waren mehr relevante Quellen aufgef\u00fchrt und relevante Quellen standen in der Ergebnisliste an h\u00f6herer Stelle als zuvor. In einzelnen Datens\u00e4tzen ergaben sich sogar Verbesserungen von bis zu 50 Prozent. Dabei lie\u00dfen sich vergleichbare Ergebnisse auch mit Artikeln erreichen, die nicht Teil der Optimierungsschleife waren (<a href=\"https://en.wikipedia.org/wiki/Cross-validation_(statistics)\" target=\"_blank\" rel=\"noopener\">Cross-Validation</a>).</p>\n<h2>Warum ist das relevant?</h2>\n<p>F\u00fcr wissenschaftliche Infrastrukturen wie digitale Bibliotheken, Fachrepositorien oder Forschungsdatenplattformen ist es entscheidend, KI-Systeme effizient und ressourcenschonend betreiben zu k\u00f6nnen. Die Ergebnisse zeigen: <strong>Sinnvolle RAG-Optimierung ist auch ohne Rechenzentrum machbar</strong>.</p>\n<p>Das senkt technische H\u00fcrden, reduziert Kosten und macht den Einsatz moderner KI-Technologien auch in kleineren Projekten realistisch.</p>\n<h2>Ausblick</h2>\n<p>Die f\u00fcr die Suche verwendeten <a href=\"https://en.wikipedia.org/wiki/Embedding_(machine_learning)\" target=\"_blank\" rel=\"noopener\">Embedding-Vector-Modelle</a> haben einen erheblichen Einfluss auf die Ergebnisse (vgl. <a href=\"https://arxiv.org/abs/2505.03452\" target=\"_blank\" rel=\"noopener\">Orbach et al. (2025)</a>) und zwar sowohl auf die Rechenzeit als auch auf die Ergebnisqualit\u00e4t. Dabei zeigen Modelle nicht auf allen Datens\u00e4tzen die gleichen Ergebnisse.</p>\n<p>Es ist auch nur begrenzt m\u00f6glich, die Optimierung mit extrem kleinen oder schnellen Embedding-Modellen auszuf\u00fchren und die optimierten Parameter dann zusammen mit einem anderen, leistungsf\u00e4higen Modell im Live-Betrieb einzusetzen. Sind die eingesetzten Embedding-Modelle nicht angepasst genug an die verwendete Wissensdom\u00e4ne, liefert auch die Optimierung nur suboptimale Ergebnisse.</p>\n<p>Genau an diesem Punkt wird unsere Arbeit im Open Science Lab in der n\u00e4chsten Zeit ansetzen. Gemeinsam mit den Fachinformationsdiensten FID Material Science, FID Move, FID Pyhsik und FID Philosophie evaluieren wir die M\u00f6glichkeit einer <strong>st\u00e4rkeren Vernetzung von NFDI und FIDs</strong> mit dem Ziel, die einzelnen Wissendom\u00e4nen mit fachspezifischen Embedding-Modellen zu versorgen. Zielsetzung ist es, damit den Zugang zu dieser Technologie noch weiter zu vereinfachen sowie die Qualit\u00e4t der Ergebnisse von KI-Anwendungen im Forschungs- und Bibliotheksumfeld gezielt zu erh\u00f6hen.</p>\n<blockquote>\n<figure style=\"width: 150px\" class=\"wp-caption alignright\"><a href=\"https://www.hs-hannover.de/service/personenfinder/person/1000005796\" target=\"_blank\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" src=\"https://www.tib.eu/fileadmin/_processed_/b/1/csm_Bluemel_Ina_57de4441fa.jpg\" alt=\"Prof. Dr. Ina Bl\u00fcmel\" width=\"150\" height=\"150\" /></a><figcaption class=\"wp-caption-text\">Prof. Dr. Ina Bl\u00fcmel, Open Science Lab // Foto: TIB/C. Bierwagen</figcaption></figure>\n<p data-pm-slice=\"1 1 []\">\"AutoRAG ist f\u00fcr uns ein wichtiger Innovationsschritt: Es macht RAG in offenen Wissensr\u00e4umen wie Wikibase messbar, wiederholbar und mit \u00fcberschaubaren Ressourcen betreibbar. F\u00fcr Projekte wie NFDI4Culture und weitere Vorhaben im Open Science Lab bedeutet das sp\u00fcrbar bessere, nachvollziehbare KI-gest\u00fctzte Suche \u00fcber heterogene Best\u00e4nde \u2013 ohne dass tiefes Spezial-Know-how aufgebaut werden muss. N\u00e4chster Schritt ist der Ausbau fachspezifischer Embeddings, kuratierter Testsets und transparenter Workflows, damit die Qualit\u00e4t und Nachnutzbarkeit langfristig steigt.\"</p>\n</blockquote>\n<h3 style=\"padding-left: 0 important!;\">Relevante Links</h3>\n<ul>\n<li><strong>Wikibase4Research</strong>: <a href=\"https://gitlab.com/nfdi4culture/wikibase4research/wikibase4research\" target=\"_blank\" rel=\"noopener\">https://gitlab.com/nfdi4culture/wikibase4research/wikibase4research</a></li>\n<li><strong>Wikibase4Research-RAG Modul:</strong> <a href=\"https://gitlab.com/nfdi4culture/wikibase4research/wikibase-RAG\" target=\"_blank\" rel=\"noopener\">https://gitlab.com/nfdi4culture/wikibase4research/wikibase-RAG</a></li>\n</ul>","doi":"https://doi.org/10.65527/gvnwm-x6559","guid":"https://blog.tib.eu/?p=30916","image":"https://blog.tib.eu/wp-content/uploads/2026/03/AutoRAG-thumb.jpg","language":"de","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1775001600,"rid":"b11wn-f5028","summary":"KI-Systeme, die Texte nicht nur generieren, sondern gezielt in Dokumenten recherchieren, sind mittlerweile etablierter Stand der Technik. Damit ein solches System gut funktioniert, m\u00fcssen viele Stellschrauben richtig eingestellt werden, was in der Regel zeitaufw\u00e4ndig, rechenintensiv und technisch anspruchsvoll ist. Unsere aktuelle Untersuchung zeigt jedoch: Eine automatisierte Optimierung ist m\u00f6glich \u2013 sogar auf einem normalen Laptop.","tags":["OPENNESS","FORSCHUNG & PROJEKTE","Lizenz:CC-BY-4.0-INT","Wikibase","Projekte"],"title":"Bessere KI-Antworten \u2013 auch ohne Hochleistungsrechner","updated_at":1790067915,"url":"https://blog.tib.eu/2026/04/01/bessere-ki-antworten-auch-ohne-hochleistungsrechner/","version":"v1"}}],"items":[{"authors":[{"contributor_roles":[],"family":"Eden","given":"Terence","url":"https://orcid.org/0000-0002-9265-9069"}],"blog":{"authors":null,"community_id":"61ce553a-bafd-4aba-a952-d3bab5e85bcc","created":1788652800,"current_feed_url":null,"description":"Regular nonsense about tech and its effects \ud83d\ude43","doi":"https://doi.org/10.59350/shkspr","favicon":"https://rogue-scholar.org/api/communities/61ce553a-bafd-4aba-a952-d3bab5e85bcc/logo","feed_format":"application/atom+xml","feed_url":"https://shkspr.mobi/blog/feed/DOI","filter":"category:-1982","generator":"WordPress","home_page_url":"https://shkspr.mobi/blog","issn":"2753-1570","language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"shkspr","status":"active","subfield":"1712","title":"Terence Eden's Blog","updated":1790106552,"use_api":true},"blog_name":"Terence Eden's Blog","blog_slug":"shkspr","content_html":"<img src=\"https://shkspr.mobi/blog/wp-content/uploads/2026/09/how-to-build-a-space-station.webp\" alt=\"Book cover featuring astronauts on Mars looking at a habitat.\" width=\"256\" height=\"384\" class=\"alignleft\">\n\n<p>This book, by The Times' former Architecture Correspondent, stands in direct opposition to <a href=\"https://shkspr.mobi/blog/2026/07/book-review-a-city-on-mars-by-dr-kelly-weinersmith-and-zach-weinersmith/\">A City on Mars</a>. Whereas that book presented a (perhaps too) sceptical look at the realities of living on other planets, Jonathan Morrison's book goes (perhaps too far) in the opposite direction.</p>\n\n<!-- \n\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\ud83d\udc4b\nWelcome to CommentClub! This content is only available to people who read my HTML comments. Why would you do that? What are you hoping to learn? So, the three rules of comment club are\u2026\n\n1. You must not tell anyone about Comment Club - let them find out about it themselves.\n2. You *must* add some Comment Club Content to at least one of your blog posts. Doesn't have to be much, just a little note will do.\n3. You should drop an email to anyone whose comments you've read. Start by sending me a friendly message at comment.club@shkspr.mobi\n\nKeep your eyes peeled for more comments in future blog posts \ud83d\ude03\n\nTTFN.\n-->\n\n<p>How to Build a Space Station is a beautiful examination of just how important architecture will be to our colonisation of other worlds. Not just in terms of physical safety - but psychological safety as well. It is a direct and forceful rebuttal to those who say it cannot be done.</p>\n\n<p>It is, in my opinion, just a touch too credulous about some of the ludicrous claims from the hype merchants. I want to believe that Martian igloos can be conjured out of the ice and that Musk's rockets will deliver a steady stream of supplies to distant worlds. But the evidence presented is rather thin. The book works best when it focuses on what architecture can bring to the table when it comes to designing the future.</p>\n\n<blockquote><p>In short, a spacecraft is not just a machine; it is also a home, an office, a refuge. If people are asked to go to the most remote environments, to live in spaces scarcely larger than a few rooms, and to perform work of immense complexity and risk, then comfort, efficiency and ergonomics are not just luxuries.</p></blockquote>\n\n<p>Space has to be <em>worth</em> living in. Putting people into a tin-can with no windows, blank walls, and an infernal background hum will drive them mad. All this is backed up with extensive descriptions of the engineering challenges of polar research bases, spaceports, and previous craft.</p>\n\n<p>Despite being rightly scathing about Wernher von Braun's involvement in atrocities and his eventual political rehabilitation - he is somewhat more muted in his criticism of Messrs Musk &amp; Bezos. There's a <em>lot</em> of praise for celebrity architects and designers - without any real examination of whether their designs are practical rather than just award fodder.</p>\n\n<p>Similarly, the book takes on trust that autonomous robots <em>can</em> ingest extraterrestrial soil, process it, and 3D print structures from it all while in a hostile environment. The fact that we don't have swarms of drones prefabbing houses in the relatively benign atmosphere of our planet should be evidence that maybe these claims aren't quite matched with reality.</p>\n\n<p>Finally, the \"why?\" question. A City on Mars points out that the cost of mining gold from asteroids would be more profitably spent improving mining technology here on Earth. How to Build a Space Station takes a different approach; it'll improve things here:</p>\n\n<blockquote><p>Space architecture is not just escapism, a thrilling sci-\u00adfi fantasy \u2013 it is a forge for creating the tools we need at home. These include but are not limited to circular systems, low-\u00adenergy fabrication, modular construction and buildings that take psychology seriously.</p></blockquote>\n\n<p>I have a lot of sympathy for that. Except\u2026 the Internation Space Station has shown us how to endlessly recycle water relatively cheaply. Yet every modern building on Earth pays only lip-service to reusing grey-water. 3D printing is amazing, but the number of structures built using autonomous robots extruding concrete is approximately zero.</p>\n\n<p>We have the technology - but we don't seem to be interested in using it.</p>\n\n<p>The book is mostly well illustrated - with some gorgeous drawings of actual craft and possible future inventions. Sadly no photos, maps, or anything to help illuminate some of the other challenges faced by living and working in space.</p>\n\n<p>This book is endlessly fascinating and bang up to date, with lots of talk of events that happened in 2025. The way it brings together the sciences of engineering and psychology is marvellous.  But, as much as I'd like to believe in a Martian habitat built by robot trebuchets flinging microwave sintered tetrapods into each other, I just don't find it convincing.</p>\n\n<p>I <em>really</em> hope I'm wrong.</p>\n\n<p>Many thanks to Netgalley for the review copy - the book is available to buy now.</p><img src=\"https://shkspr.mobi/blog/wp-content/themes/edent-wordpress-theme/info/okgo.php?ID=75653&HTTP_REFERER=DOI\" alt width=1 height=1 loading=eager>","doi":"https://doi.org/10.59350/0vqd8-3qy21","guid":"https://shkspr.mobi/blog/?p=75653","image":"https://shkspr.mobi/blog/wp-content/uploads/2026/09/how-to-build-a-space-station.webp","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1789862400,"rid":"8s9qs-35530","summary":"This book, by The Times' former Architecture Correspondent, stands in direct opposition to A City on Mars. Whereas that book presented a (perhaps too) sceptical look at the realities of living on other planets, Jonathan Morrison's book goes (perhaps too far) in the opposite direction.","tags":["/etc/","Book Review","NetGalley"],"title":"Book Review: How to Build a Space Station by Jonathan Morrison","updated_at":1790106912,"url":"https://shkspr.mobi/blog/2026/09/book-review-how-to-build-a-space-station-by-jonathan-morrison/","version":"v1"},{"authors":[{"affiliation":[{"name":"Front Matter"}],"contributor_roles":[],"family":"Fenner","given":"Martin","url":"https://orcid.org/0000-0003-1419-2405"}],"blog":{"authors":[{"name":"Martin Fenner","url":"https://orcid.org/0000-0003-1419-2405"}],"community_id":"15a362ea-8138-42b8-917f-1840a92addf8","created":1672531200,"current_feed_url":null,"description":"The Front Matter Blog covers the intersection of science and technology since 2007.","doi":"https://doi.org/10.53731/front_matter","favicon":"https://rogue-scholar.org/api/communities/15a362ea-8138-42b8-917f-1840a92addf8/logo","feed_format":"application/atom+xml","feed_url":"https://blog.front-matter.de/atom","filter":null,"generator":"Ghost","home_page_url":"https://blog.front-matter.de/","issn":"2749-9952","language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.53731","relative_url":null,"secure":true,"slug":"front_matter","status":"active","subfield":"1710","title":"Front Matter","updated":1790088325,"use_api":true},"blog_name":"Front Matter","blog_slug":"front_matter","content_html":"<p>The science blog archive Rogue Scholar is launching an Editorial Board to handle submissions of new blogs, and oversee editorial management of existing blogs, e.g. blogs or individual posts that should be retracted or no longer archived.</p><p>The Rogue Scholar science blog archive accepts new blogs via a submission form and after acceptance doesn't interfere with what participating blogs publish or how often. If participating blogs stop publishing it doesn't expect a notification. While this adhoc workflow generally worked for the 200 blogs accepted into Rogue Scholar so far, there are some issues that can be improved:</p><ul><li>Processing of submissions of new blogs takes too long and with no clear feedback loop,</li><li>Blog posts written in languages other than German and English are hard to understand for me, and I am more comfortable in some scientific disciplines than others,</li><li>The decision for acceptance of a new blog is ultimately the decision of a single person (me),</li><li>The requirements in the submission form are frequently ignored (e.g. supported language, number of posts, or use of AI),</li><li>The kinds of new blogs we want to see has not been clearly articulated, and no systematic active blog recruitment is happening.</li></ul><p>My hope is that the new Rogue Scholar Editorial Board will address these issues. Starting this week all new submissions will be processed once a month on the 15th, with time for additional questions before the decision if a submission is unclear. The submission form has added a question to describe the blog in a free-text field to provide more context. And I am looking for volunteers to join the Editorial Board, so that the workload and governance can be shared by more people. I hope to have an Editorial Board with 3-5 additional people setup until the end of the year, and I am looking for volunteers comfortable in other languages and scientific disciplines. This Editorial Board will supplement the work of the Rogue Scholar Advisory, and the Board of the non-profit organization once the Rogue Scholar non-profit launches. This is potentially two half days of volunteer work every month, but it is hopefully rewarding work. Until the Editorial Board is in place I will process blog submissions alone at the 15th of every month.</p><p>Please reach out via&nbsp;<a href=\"https://join.slack.com/t/rogue-scholar/shared_invite/zt-2ylpq1yoy-o~TkxDarfz5LSMhGSCYtiA\" rel=\"noreferrer\">Slack</a>,&nbsp;<a href=\"mailto:info@rogue-scholar.org\" rel=\"noreferrer\">email</a>,&nbsp;<a href=\"https://wisskomm.social/@rogue_scholar\" rel=\"noreferrer\">Mastodon</a>, or&nbsp;<a href=\"https://bsky.app/profile/rogue-scholar.bsky.social\" rel=\"noreferrer\">Bluesky</a>&nbsp;if you want to join the new Editorial Board.</p><div class=\"kg-card kg-callout-card kg-callout-card-blue\"><div class=\"kg-callout-text\">Rogue Scholar is a scholarly infrastructure that is free for all authors and readers. You can support Rogue Scholar with a one-time or recurring&nbsp;<a href=\"https://ko-fi.com/rogue_scholar\" rel=\"noreferrer\">donation</a>&nbsp;or by becoming a sponsor.</div></div>","doi":"https://doi.org/10.53731/wtf9v-y0n52","guid":"https://doi.org/10.53731/wtf9v-y0n52","image":"https://images.unsplash.com/photo-1613799591389-0160e094db32?crop=entropy&cs=tinysrgb&fit=max&fm=jpg&ixid=M3wxMTc3M3wwfDF8c2VhcmNofDI4N3x8RWRpdG9yaWFsJTIwYm9hcmR8ZW58MHx8fHwxNzkwMDg1MzY3fDA&ixlib=rb-4.1.0&q=80&w=2000","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1790035200,"rid":"89y6z-bd709","summary":"The science blog archive Rogue Scholar is launching an Editorial Board to handle submissions of new blogs, and oversee editorial management of existing blogs, e.g. blogs or individual posts that should be retracted or no longer archived. The Rogue Scholar science blog archive accepts new blogs via a submission form and after acceptance doesn't interfere with what participating blogs publish or how often.","tags":["Rogue Scholar"],"title":"Rogue Scholar is launching an Editorial Board","updated_at":1790092816,"url":"https://blog.front-matter.de/posts/rogue-scholar-is-launching-an-editorial-board/","version":"v1"},{"authors":[{"affiliation":[{"id":"https://ror.org/0153tk833","name":"University of Virginia"}],"contributor_roles":[],"family":"Turner","given":"Stephen","url":"https://orcid.org/0000-0001-9140-9028"}],"blog":{"authors":[{"name":"Stephen Turner"}],"community_id":"382941a7-2ffa-41df-8bbb-5f772188517f","created":1780876800,"current_feed_url":null,"description":"A practicing data scientist's take on AI, genomics, biosecurity, and the ways AI is reshaping how science gets done. Weekly updates from the field. Occasional notes on programming.","doi":"https://doi.org/10.59350/stephenturner","favicon":"https://rogue-scholar.org/api/communities/382941a7-2ffa-41df-8bbb-5f772188517f/logo","feed_format":"application/rss+xml","feed_url":"https://blog.stephenturner.us/feed","filter":null,"generator":"Substack","home_page_url":"https://blog.stephenturner.us","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"stephenturner","status":"active","subfield":"1311","title":"Paired Ends","updated":1790091521,"use_api":true},"blog_name":"Paired Ends","blog_slug":"stephenturner","content_html":"<p>I went on CBS19 TV news here in Charlottesville last night to talk about AI and biosecurity. I thought I was walking into a session that was going to be taped and played over a slow news day. I was pleasantly surprised (and a little thrown off) when I found out the segment was live on the evening news! <a href=\"https://www.cbs19news.com/news/inside-the-numbers-09-21/video_e59405cb-29cd-5b0d-8d85-929e19b2bf36.html\">Here's the recording</a>. </p><p>Biosecurity, and AI safety in general, is a really nuanced topic to try to get right in 240 seconds!</p><div class=\"captioned-image-container\"><figure><a class=\"image-link image2 is-viewable-img\" target=\"_blank\" href=\"https://www.cbs19news.com/news/inside-the-numbers-09-21/video_e59405cb-29cd-5b0d-8d85-929e19b2bf36.html\" data-component-name=\"Image2ToDOM\"><div class=\"image2-inset\"><picture><source type=\"image/webp\" srcset=\"https://substackcdn.com/image/fetch/$s_!0d-_!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 424w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 848w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 1456w\" sizes=\"100vw\"><img src=\"https://substackcdn.com/image/fetch/$s_!0d-_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg\" width=\"1184\" height=\"663\" data-attrs=\"{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/ccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:663,&quot;width&quot;:1184,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:204149,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/jpeg&quot;,&quot;href&quot;:&quot;https://www.cbs19news.com/news/inside-the-numbers-09-21/video_e59405cb-29cd-5b0d-8d85-929e19b2bf36.html&quot;,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://blog.stephenturner.us/i/216919562?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}\" class=\"sizing-normal\" alt=\"\" srcset=\"https://substackcdn.com/image/fetch/$s_!0d-_!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 424w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 848w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 1272w, https://substackcdn.com/image/fetch/$s_!0d-_!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg 1456w\" sizes=\"100vw\" fetchpriority=\"high\"></picture><div class=\"image-link-expand\"><div class=\"pencraft pc-display-flex pc-gap-8 pc-reset\"><button tabindex=\"0\" type=\"button\" class=\"pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M\"><svg aria-hidden=\"true\" width=\"20\" height=\"20\" viewBox=\"0 0 20 20\" fill=\"none\" stroke-width=\"1.5\" stroke=\"var(--color-fg-primary)\" stroke-linecap=\"round\" stroke-linejoin=\"round\" xmlns=\"http://www.w3.org/2000/svg\" class=\"icon-noB79L\"><g><path d=\"M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882\"></path></g></svg></button><button tabindex=\"0\" type=\"button\" class=\"pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M\"><svg xmlns=\"http://www.w3.org/2000/svg\" width=\"20\" height=\"20\" viewBox=\"0 0 24 24\" fill=\"none\" stroke=\"currentColor\" stroke-width=\"2\" stroke-linecap=\"round\" stroke-linejoin=\"round\" class=\"lucide lucide-maximize2 lucide-maximize-2 icon-noB79L\"><polyline points=\"15 3 21 3 21 9\"></polyline><polyline points=\"9 21 3 21 3 15\"></polyline><line x1=\"21\" x2=\"14\" y1=\"3\" y2=\"10\"></line><line x1=\"3\" x2=\"10\" y1=\"21\" y2=\"14\"></line></svg></button></div></div></div></a><figcaption class=\"image-caption\">View the full 4 minute segment on <a href=\"https://www.cbs19news.com/news/inside-the-numbers-09-21/video_e59405cb-29cd-5b0d-8d85-929e19b2bf36.html\">CBS19 News Charlottesville</a>.</figcaption></figure></div><p>Here's what I tried to get across.</p><ol><li><p><strong>The OpenAI / Hugging Face incident was a warning shot.</strong> It's a cybersecurity incident that shows that capable AI systems can persist at hard problems, use tools, escape containment, and hack into systems people thought were secure. In biosecurity, AI can lower the barriers at important steps like finding technical information, planning experiments, troubleshooting failures, and operating tools. We need to test what systems actually do in realistic settings, in a safe way, to really understand the biosecurity risks posed by AI.</p></li><li><p><strong>Be skeptical of doomsday countdowns.</strong> The anchor asked me about doomsday predictions so I advised treating these with skepticism. Some capabilities already here, but creating biothreats still requires intent, expertise, access to materials and equipment. And we here at the University of Virginia School of Data Science as well as many other <a href=\"https://blog.stephenturner.us/p/ai-biosecurity-review\">really talented people and organizations are working tirelessly</a> to make sure what just happened in cyber never happens in bio. </p></li><li><p><strong>We should take the risk seriously without overstating it.</strong> Just because AI knows a lot about biology doesn't mean some rogue AI agent can go out and create bioweapons. And, AI won't turn someone into a biologist overnight: biology still requires materials, laboratory skills, and real-world infrastructure. We need to measure this to understand it. That's one of the things we're trying to understand here: what uplift does AI provide for doing biology if you don't have a PhD in biology?</p></li><li><p><strong>We need layers of protection.</strong> <a href=\"https://www.rand.org/pubs/research_reports/RRA4999-1.html\">Defense in depth</a>, as Steph Guerra and friends at RAND call it. Realistic testing before release, limits and monitoring for high-risk AI use, strong containment for autonomous agents, and safeguards at physical chokepoints like DNA synthesis and lab access. We also need independent evaluation and prompt incident reporting. We can't rely on a single company to get this right for all of us.</p></li><li><p><strong>Pace of AI progress vs AI governance.</strong> AI capabilities are improving fast. Governance needs to keep up. But we shouldn't lock everything down with every possible restriction. We should prioritize measuring capabilities and how we mitigate threats: secure evaluation environments, independent testing, incident reporting, and screening at physical chokepoints. Without good measurement we can't manage risk effectively.</p></li><li><p><strong>It's not either or.</strong> The anchor asked me: \"Using AI to cure cancer would be good. Making bioweapons would be bad. What's the balance?\" Nobody wants to stop AI from helping with cancer, vaccines, or public health. We all want to make beneficial uses easier and dangerous uses harder or impossible. That means targeted safeguards, independent testing, and updating policy when the evidence changes.&nbsp;</p></li><li><p><strong>Measurement matters.</strong> We can't implement calibrated mitigations if we don't know what we're dealing with. Evals tell us a lot, but we're also working to go further: when AI performs well on a benchmark, to what degree does that translate into a person being more capable at doing biology in the real world? </p></li></ol><p class=\"button-wrapper\" data-attrs=\"{&quot;url&quot;:&quot;https://blog.stephenturner.us/subscribe?&quot;,&quot;text&quot;:&quot;Subscribe now&quot;,&quot;action&quot;:null,&quot;class&quot;:null}\" data-component-name=\"ButtonCreateButton\"><a class=\"button primary\" href=\"https://blog.stephenturner.us/subscribe?\"><span>Subscribe now</span></a></p><p></p>","doi":"https://doi.org/10.59350/ar3zs-43b27","guid":"216919562","image":"https://substackcdn.com/image/fetch/$s_!0d-_!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Fccd2b525-f5ff-4854-bba3-541a20e7c3da_1184x663.jpeg","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1790035200,"rid":"kakmx-0ep33","summary":"I got four minutes on live TV news to take on an important but nuanced topic.","tags":["Biosecurity","AI"],"title":"AI + Biosecurity in 4 minutes","updated_at":1790091756,"url":"https://blog.stephenturner.us/p/ai-biosecurity-tv-interview","version":"v1"},{"authors":[{"contributor_roles":[],"family":"Kehm","given":"Nicole"}],"blog":{"authors":null,"community_id":"db0d8909-9e37-46d0-b16c-0551f575e86b","created":1749772800,"current_feed_url":null,"description":"Das Blog der TIB \u2013 Leibniz-Informationszentrum Technik und Naturwissenschaften und Universit\u00e4tsbibliothek","doi":"https://doi.org/10.65527/tib","favicon":"https://rogue-scholar.org/api/communities/db0d8909-9e37-46d0-b16c-0551f575e86b/logo","feed_format":"application/atom+xml","feed_url":"https://blog.tib.eu/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.tib.eu/","issn":null,"language":"deu","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.65527","relative_url":null,"secure":true,"slug":"tib","status":"active","subfield":"1802","title":"TIB-Blog","updated":1790082029,"use_api":true},"blog_name":"TIB-Blog","blog_slug":"tib","content_html":"<p>Die Retrodigitalisierung startete 2015 mit der Digitalisierung der Best\u00e4nde der TIB. Begonnen hat das Team Retrodigitalisierung mit einem kleinen Team mit drei Personen, das sich im Laufe der Zeit auf aktuell 12 Personen vergr\u00f6\u00dfert hat. In diesem Zeitraum von elf Jahren konnte die Retrodigitalisierung viele B\u00fccher, Graphiken, Pl\u00e4ne usw. mit vielen, vielen Seiten scannen.</p>\n<p>Die Retrodigitalisierung digitalisiert sowohl gemeinfreie Werke sowie noch urheberrechtgebundene Werke zur Bestandserhaltung. Gemeinfrei ist ein Werk, wenn der Urheber des Werks bereits 70 Jahre verstorben ist. Zurzeit sind mehr als 12.000 Digitalisate der TIB online im <a href=\"https://goobi.tib.eu/viewer/index/\">TIB-Viewer</a> verf\u00fcgbar. Seit ein paar Wochen befinden sich diese Digitalisate nun auch in der Deutschen Digitalen Bibliothek (DDB).</p>\n<figure id=\"attachment_33535\" aria-describedby=\"caption-attachment-33535\" style=\"width: 701px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-33535\" src=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild1.png\" alt=\"\" width=\"701\" height=\"484\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild1.png 907w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild1-300x207.png 300w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild1-768x530.png 768w\" sizes=\"auto, (max-width: 701px) 100vw, 701px\" /><figcaption id=\"caption-attachment-33535\" class=\"wp-caption-text\">Startseite Deutsche Digitale Bibliothek</figcaption></figure>\n<h2>Was ist die Deutsche Digitale Bibliothek?</h2>\n<p>Die <a href=\"https://www.deutsche-digitale-bibliothek.de/\">Deutsche Digitale Bibliothek</a> (DDB) ist ein Portal f\u00fcr Kulturobjekte, die neben B\u00fcchern auch Archivalien, Bilder, Fotografien, Skulpturen, Musikst\u00fccke, Filme, Noten, Gem\u00e4lde, Handschriften usw. beinhaltet.</p>\n<p>2012 startete die DDB mit einer Betaversion und 2014 mit der Vollversion. Finanziert wird die DDB von Bund, L\u00e4ndern und Kommunen. Getragen wird die DDB von einen Kompetenznetzwerk bestehend aus <a href=\"https://www.deutsche-digitale-bibliothek.de/content/wie-wir-organisiert-sind\">18 Kultur- und Wissenseinrichtungen</a>, wie zum Beispiel die Bayerische Staatsbibliothek, das Bundesarchiv, die Deutsche Nationalbibliothek, das FIZ Karlsruhe (Leibniz-Institut f\u00fcr Informationsinfrastruktur) oder die Stiftung Preu\u00dfischer Kulturbesitz.</p>\n<p>Das Ziel der DDB ist es, das kulturelle Erbe in Deutschland digital, kostenlos und jederzeit zug\u00e4nglich zu machen. \u00dcber <a href=\"https://www.deutsche-digitale-bibliothek.de/about-us/institutions?view=map\">5.000 Kultureinrichtungen</a> sind an der DDB beteiligt, davon fungieren fast 1.000 als Datenpartner:innen, darunter auch die TIB \u2013 Leibniz-Informationszentrum Technik und Naturwissenschaften und Universit\u00e4tsbibliothek. Zurzeit finden sich \u00fcber 65 Millionen Objekte in der DDB.</p>\n<p>Mittlerweile geh\u00f6ren drei weitere Portale zur DDB:</p>\n<ul>\n<li><a href=\"https://www.archivportal-d.de/\">Archivportal-D</a></li>\n<li><a href=\"https://www.deutsche-digitale-bibliothek.de/newspaper\">Deutsches Zeitungsportal</a></li>\n<li><a href=\"https://ccc.deutsche-digitale-bibliothek.de/de/\">Sammlungsgut aus kolonialen Kontexten</a></li>\n</ul>\n<figure id=\"attachment_33540\" aria-describedby=\"caption-attachment-33540\" style=\"width: 874px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-33540\" src=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild5-1.png\" alt=\"\" width=\"874\" height=\"107\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild5-1.png 907w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild5-1-300x37.png 300w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild5-1-768x94.png 768w\" sizes=\"auto, (max-width: 874px) 100vw, 874px\" /><figcaption id=\"caption-attachment-33540\" class=\"wp-caption-text\">Weitere Portale der Deutschen Digitalen Bibliothek</figcaption></figure>\n<p>Die Deutsche Digitale Bibliothek ist ein Datenpartner f\u00fcr Europeana. Bei der DDB stehen Kultureinrichtungen aus Deutschland im Mittelpunkt. <a href=\"https://www.europeana.eu/de\">Europeana</a> bietet eine Plattform f\u00fcr Kulturobjekte aus ganz Europa.</p>\n<h2>Deutsche Digitale Bibliothek und die TIB</h2>\n<p>Die TIB ist bereits seit etwa 2012/2013 Datenpartnerin von Europeana und der DDB. Seit diesem Zeitpunkt werden Inhalte vom <a href=\"https://av.tib.eu/\">TIB AV-Portal</a> in die DDB und Europeana eingespielt. Sp\u00e4ter wurde der DDB-Bestand um die graphischen Einzelbl\u00e4tter der Sammlung Albrecht Haupt sowie um die Reiseskizzen des Architekten Albrecht Haupt (1852-1932) erg\u00e4nzt. Diese Digitalisate des Teilbestands der Sammlung Albrecht Haupt stammen aus dem Portal <a href=\"https://sah.tib.eu/\">TIB SAH digital</a>.</p>\n<p>Dieses Jahr sind die Digitalisate der Retrodigitalisierung der TIB hinzugekommen. Daf\u00fcr mussten die Digitalisate zun\u00e4chst \u00fcber eine Schnittstelle in das Testsystem der DDB eingespielt werden. Dort wurden die Metadaten analysiert, da die Metadaten gewissen Standards gerecht werden m\u00fcssen, damit die Digitalisate auch in der DDB angezeigt werden k\u00f6nnen. Nach der Pr\u00fcfung konnten die Daten gleich ins Produktivsystem eingespielt werden.</p>\n<figure id=\"attachment_33537\" aria-describedby=\"caption-attachment-33537\" style=\"width: 701px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-33537\" src=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild3.png\" alt=\"\" width=\"701\" height=\"324\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild3.png 907w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild3-300x139.png 300w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild3-768x355.png 768w\" sizes=\"auto, (max-width: 701px) 100vw, 701px\" /><figcaption id=\"caption-attachment-33537\" class=\"wp-caption-text\">Metadaten in der Deutschen Digitalen Bibliothek</figcaption></figure>\n<p>Seit Anfang Juli 2026 befinden sich nun \u00fcber 12.000 <a href=\"https://www.deutsche-digitale-bibliothek.de/searchresults?query=&amp;offset=0&amp;rows=20&amp;facetValues%5B%5D=provider_id%3D4E6QWP2Z7ZV3XWT6D6MX7WSQN4HOKEWO&amp;isThumbnailFiltered=false&amp;facetValues%5B%5D=type_fct%3Dmediatype_003\">Digitalisate der TIB</a> in der DDB. In der DDB besteht die M\u00f6glichkeit, sich die Digitalisate anzuschauen oder auf den Datengeber des Digitalisats zu wechseln, in diesem Fall den <a href=\"https://goobi.tib.eu/viewer/index/\">TIB-Viewer</a>.</p>\n<p><a href=\"https://www.deutsche-digitale-bibliothek.de/item/236JLJM2AUG5TZUDI5M2M3BGRFP4VMWZ\">Kompendium der praktischen Toxikologie</a></p>\n<figure id=\"attachment_33536\" aria-describedby=\"caption-attachment-33536\" style=\"width: 701px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-33536\" src=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild2.png\" alt=\"\" width=\"701\" height=\"283\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild2.png 907w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild2-300x121.png 300w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild2-768x310.png 768w\" sizes=\"auto, (max-width: 701px) 100vw, 701px\" /><figcaption id=\"caption-attachment-33536\" class=\"wp-caption-text\">Bildanzeige in der Deutschen Digitalen Bibliothek</figcaption></figure>\n<p>Im n\u00e4chsten Schritt sollen die Digitalisate von der DDB an Europeana \u00fcbermittelt werden. Geplant sind viertelj\u00e4hrliche Updates an die DDB, damit auch die neuen Digitalisate der Retrodigitalisierung der TIB an die DDB \u00fcbermittelt werden k\u00f6nnen.</p>\n<figure id=\"attachment_33538\" aria-describedby=\"caption-attachment-33538\" style=\"width: 701px\" class=\"wp-caption aligncenter\"><img loading=\"lazy\" decoding=\"async\" class=\" wp-image-33538\" src=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild4.png\" alt=\"\" width=\"701\" height=\"327\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/09/Bild4.png 907w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild4-300x140.png 300w, https://blog.tib.eu/wp-content/uploads/2026/09/Bild4-768x358.png 768w\" sizes=\"auto, (max-width: 701px) 100vw, 701px\" /><figcaption id=\"caption-attachment-33538\" class=\"wp-caption-text\">Startseite Europeana</figcaption></figure>","doi":"https://doi.org/10.65527/hc09t-hvv39","guid":"https://blog.tib.eu/?p=33534","image":"https://blog.tib.eu/wp-content/uploads/2026/09/Bild1.png","language":"de","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1790035200,"rid":"f468w-wsq28","summary":"Die Retrodigitalisierung startete 2015 mit der Digitalisierung der TIB-Best\u00e4nde. Begonnen hat das Team Retrodigitalisierung mit einem kleinen Team mit drei Personen, inzwischen sind es zw\u00f6lf Personen. In diesem Zeitraum von elf Jahren scannten sie B\u00fccher, Graphiken, Pl\u00e4ne usw. mit vielen, vielen Seiten. Zurzeit sind mehr als 12.000 Digitalisate der TIB online im TIB-Viewer verf\u00fcgbar.","tags":["Vom Papier Zum Pixel","DEINE UNIBIB","BIBLIOTHEKSWELT","Lizenz:CC-BY-4.0-INT","Retrodigitalisierung"],"title":"Digitalisate der TIB in der Deutschen Digitalen Bibliothek","updated_at":1790082085,"url":"https://blog.tib.eu/2026/09/22/digitalisate-der-tib-in-der-deutschen-digitalen-bibliothek/","version":"v1"},{"authors":[{"contributor_roles":[],"name":"Pensoft Editorial Team"}],"blog":{"authors":null,"community_id":"1e54953c-587a-4743-89af-ebd2cb45660c","created":1789516800,"current_feed_url":null,"description":null,"doi":"https://doi.org/10.59350/pensoft","favicon":"https://rogue-scholar.org/api/communities/1e54953c-587a-4743-89af-ebd2cb45660c/logo","feed_format":"application/atom+xml","feed_url":"https://blog.pensoft.net/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.pensoft.net/","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"pensoft","status":"active","subfield":"2303","title":"Pensoft Blog","updated":1790009683,"use_api":true},"blog_name":"Pensoft Blog","blog_slug":"pensoft","content_html":"<div class=\"twitter-share\"><a href=\"https://twitter.com/intent/tweet?via=Pensoft\" class=\"twitter-share-button\" data-size=\"large\">Tweet</a></div>\n\n<p>Digital biodiversity data has come a long way. Twenty-five years ago, the <a href=\"https://www.gbif.org/\">Global Biodiversity Information Facility</a> (GBIF) launched with 21 founding countries and a simple but ambitious goal: make the world&#8217;s biodiversity data freely and openly available to anyone, anywhere. </p>\n\n\n<div class=\"wp-block-essential-blocks-table-of-contents\"><div class=\"eb-parent-wrapper eb-parent-eb-toc-rs5x1 \"><div class=\"eb-toc-container eb-toc-rs5x1  eb-toc-is-not-sticky eb-toc-not-collapsible eb-toc-initially-not-collapsed eb-toc-scrollToTop style-1 list-style-none\" data-scroll-top=\"false\" data-scroll-top-icon=\"fas fa-angle-up\" data-collapsible=\"false\" data-sticky-hide-mobile=\"false\" data-sticky=\"false\" data-scroll-target=\"scroll_to_toc\" data-copy-link=\"false\" data-editor-type=\"\" data-hide-desktop=\"false\" data-hide-tab=\"false\" data-hide-mobile=\"false\" data-itemCollapsed=\"false\"><div class=\"eb-toc-header\"><div class=\"eb-toc-title\">Table of Contents</div></div><div class=\"eb-toc-wrapper \" data-headers=\"[{&quot;level&quot;:2,&quot;content&quot;:&quot;A Growing Global Network&quot;,&quot;text&quot;:&quot;A Growing Global Network&quot;,&quot;link&quot;:&quot;a-growing-global-network&quot;},{&quot;level&quot;:2,&quot;content&quot;:&quot;Data Supporting Science and Policy&quot;,&quot;text&quot;:&quot;Data Supporting Science and Policy&quot;,&quot;link&quot;:&quot;data-supporting-science-and-policy&quot;},{&quot;level&quot;:2,&quot;content&quot;:&quot;Adapting to New Kinds of Data&quot;,&quot;text&quot;:&quot;Adapting to New Kinds of Data&quot;,&quot;link&quot;:&quot;adapting-to-new-kinds-of-data&quot;},{&quot;level&quot;:2,&quot;content&quot;:&quot;Persistent Gaps Remain&quot;,&quot;text&quot;:&quot;Persistent Gaps Remain&quot;,&quot;link&quot;:&quot;persistent-gaps-remain&quot;},{&quot;level&quot;:2,&quot;content&quot;:&quot;Looking Ahead&quot;,&quot;text&quot;:&quot;Looking Ahead&quot;,&quot;link&quot;:&quot;looking-ahead&quot;}]\" data-visible=\"[true,true,true,true,true,true]\" data-delete-headers=\"[{&quot;label&quot;:&quot;A Growing Global Network&quot;,&quot;value&quot;:&quot;a-growing-global-network&quot;,&quot;isDelete&quot;:false},{&quot;label&quot;:&quot;Data Supporting Science and Policy&quot;,&quot;value&quot;:&quot;data-supporting-science-and-policy&quot;,&quot;isDelete&quot;:false},{&quot;label&quot;:&quot;Adapting to New Kinds of Data&quot;,&quot;value&quot;:&quot;adapting-to-new-kinds-of-data&quot;,&quot;isDelete&quot;:false},{&quot;label&quot;:&quot;Persistent Gaps Remain&quot;,&quot;value&quot;:&quot;persistent-gaps-remain&quot;,&quot;isDelete&quot;:false},{&quot;label&quot;:&quot;Looking Ahead&quot;,&quot;value&quot;:&quot;looking-ahead&quot;,&quot;isDelete&quot;:false}]\" data-smooth=\"true\" data-top-offset=\"\"><div class=\"eb-toc__list-wrap\"><ul class='eb-toc__list'><li><a href=\"#a-growing-global-network\">A Growing Global Network</a><li><a href=\"#data-supporting-science-and-policy\">Data Supporting Science and Policy</a><li><a href=\"#adapting-to-new-kinds-of-data\">Adapting to New Kinds of Data</a><li><a href=\"#persistent-gaps-remain\">Persistent Gaps Remain</a><li><a href=\"#looking-ahead\">Looking Ahead</a></ul></div></div></div></div></div>\n\n\n<p>Today, GBIF's network includes <strong>70 participating countries</strong>, along with data providers across more than 120 countries, offering open access to <strong>over 3.8 billion biodiversity data records</strong>. This milestone is explored in a <a href=\"https://doi.org/10.3897/BDJ.14.e208528\" title=\"\">new overview article published</a> in the <em><a href=\"https://bdj.pensoft.net/\" title=\"\">Biodiversity Data Journal</a></em>, which traces GBIF&#8217;s trajectory over the past quarter-century and the challenges that remain as it looks ahead.</p>\n\n\n\n<p>GBIF Executive Secretary Dr. Joe Miller remarked:</p>\n\n\n\n<blockquote class=\"wp-block-quote is-layout-flow wp-block-quote-is-layout-flow\">\n<p>As GBIF celebrates its 25th anniversary, I'm very proud to serve as Executive Secretary of the organisation and to have contributed to this milestone publication. Detailing the history of GBIF both as an infrastructure and a global network, the article pays tribute to all the people involved in the network throughout a quarter century\u2014be they staff, participants, nodes, volunteers, data publishers or data users. It represents the past, present and future of GBIF. I believe the paper will help strengthen our network and its capacity to respond to biodiversity data needs both now and in the future.</p>\n</blockquote>\n\n\n\n<p>GBIF was established in 2001, following a proposal from the Organisation for Economic Cooperation and Development's Megascience Forum which concluded that \"<em>an international mechanism is needed to make biodiversity data and information accessible worldwide</em>\". </p>\n\n\n\n<p>The call came in response to a problem facing the international community after the 1992 Rio Earth Summit: the world&#8217;s biodiversity knowledge was scattered across a small number of institutions, concentrated in wealthy countries, with no shared way to bring it together in support of the newly signed Convention on Biological Diversity.</p>\n\n\n\n<p class=\"has-electric-grass-gradient-background has-background\">Now, a quarter-century later, GBIF has become the foundational infrastructure for biodiversity science and policy worldwide.</p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>A Growing Global Network</strong></h2>\n\n\n\n<figure class=\"wp-block-image size-large\"><a href=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718328.gif?ssl=1\"><img fetchpriority=\"high\" decoding=\"async\" width=\"840\" height=\"420\" src=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718328-1024x512.gif?resize=840%2C420&#038;ssl=1\" alt=\"Biodiversity data in GBIF\" class=\"wp-image-20842\" srcset=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718328.gif?resize=1024%2C512&amp;ssl=1 1024w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718328.gif?resize=300%2C150&amp;ssl=1 300w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718328.gif?resize=768%2C384&amp;ssl=1 768w\" sizes=\"(max-width: 709px) 85vw, (max-width: 909px) 67vw, (max-width: 1362px) 62vw, 840px\" data-recalc-dims=\"1\" /></a><figcaption class=\"wp-element-caption\">Annual biodiversity data made available through the Global Biodiversity Information Facility as of July 2026. Credit to GBIF.</figcaption></figure>\n\n\n\n<p>GBIF&#8217;s work is carried out through a distributed network of 47 voting country participants and 42 organisational participants, each represented by a national or thematic &#8220;node&#8221; that mobilises data, builds local capacity, and connects biodiversity communities to GBIF&#8217;s infrastructure. More than 2,700 institutions, including museums, universities, government agencies, citizen science platforms, and a growing number of private-sector organisations, have published data through the network, with new publishers joining at a rate of more than two every three days, totalling <a href=\"https://www.gbif.org/publisher/search\">nearly 3500</a> organisations involved.</p>\n\n\n\n<figure class=\"wp-block-image size-large\"><a href=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?ssl=1\"><img decoding=\"async\" width=\"840\" height=\"964\" src=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-892x1024.jpg?resize=840%2C964&#038;ssl=1\" alt=\"key numerical metrics of the Global Biodiversity Information Facility \" class=\"wp-image-20845\" srcset=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=892%2C1024&amp;ssl=1 892w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=261%2C300&amp;ssl=1 261w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=768%2C882&amp;ssl=1 768w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=1338%2C1536&amp;ssl=1 1338w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=1784%2C2048&amp;ssl=1 1784w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?resize=1200%2C1378&amp;ssl=1 1200w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751052-scaled.jpg?w=1680 1680w\" sizes=\"(max-width: 709px) 85vw, (max-width: 909px) 67vw, (max-width: 1362px) 62vw, 840px\" data-recalc-dims=\"1\" /></a><figcaption class=\"wp-element-caption\">Infographic summary of key numerical metrics of the Global Biodiversity Information Facility as of August 2026. See dynamic metrics at <a href=\"https://www.gbif.org/\" target=\"_blank\" rel=\"noreferrer noopener\">GBIF.org</a> and at <a href=\"https://www.gbif.org/analytics/global\" target=\"_blank\" rel=\"noreferrer noopener\">https://www.gbif.org/analytics/global</a>.</figcaption></figure>\n\n\n\n<p>The <em>Biodiversity Data Journal</em> itself is one example of this network in action, as it operates <a href=\"https://blog.pensoft.net/2025/03/10/the-biodiversity-data-journal-launches-its-own-data-portal-on-gbif/\" title=\"\">its <strong>own GBIF-hosted data portal </strong></a>&#8211; one of <a href=\"https://blog.pensoft.net/2025/04/03/more-than-20-journals-published-by-pensoft-with-their-own-hosted-data-portals-on-gbif-to-streamline-and-fair-ify-biodiversity-research/\" title=\"\"><strong>more than 20 </strong>such portals across Pensoft-published journals </a>&#8211; providing direct access to hundreds of datasets and hundreds of thousands of occurrence records drawn from the journal&#8217;s publications.</p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Data Supporting Science and Policy</strong></h2>\n\n\n\n<figure class=\"wp-block-image size-large\"><a href=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?ssl=1\"><img decoding=\"async\" width=\"840\" height=\"391\" src=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330-1024x477.png?resize=840%2C391&#038;ssl=1\" alt=\"GBIF data mediation\" class=\"wp-image-20847\" srcset=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=1024%2C477&amp;ssl=1 1024w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=300%2C140&amp;ssl=1 300w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=768%2C358&amp;ssl=1 768w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=1536%2C715&amp;ssl=1 1536w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=2048%2C953&amp;ssl=1 2048w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?resize=1200%2C559&amp;ssl=1 1200w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?w=1680 1680w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718330.png?w=2520 2520w\" sizes=\"(max-width: 709px) 85vw, (max-width: 909px) 67vw, (max-width: 1362px) 62vw, 840px\" data-recalc-dims=\"1\" /></a><figcaption class=\"wp-element-caption\"><a href=\"https://doi.org/10.3897/BDJ.14.e208528.figure3\" title=\"\">The universe of GBIF data mediation. </a></figcaption></figure>\n\n\n\n<p>GBIF-mediated data now underpin an average of eight new peer-reviewed research papers per day, and have contributed to <a href=\"https://www.gbif.org/literature/search?peerReview=true&amp;literatureType=journal\">more than 15,000 publications to date</a>, spanning fields from climate change and food security to public health and invasive species management. An independent 2023 assessment <a href=\"https://www.gbif.org/news/5WZThcL928vmPnSvrGhZfE/report-reveals-return-on-investments-in-gbif\">estimated</a> that GBIF generates <strong>roughly \u20ac12 in societal benefit </strong>for<strong> every \u20ac1 invested</strong>, with researcher time savings alone valued at \u20ac35 million annually.</p>\n\n\n\n<p>The organisation&#8217;s impact also extends deeply into international policy. GBIF-mediated data support multiple indicators under the Kunming-Montreal Global Biodiversity Framework, inform assessments by the <a href=\"https://iucn.org/\">International Union for Conservation of Nature</a> (IUCN) and the <a href=\"https://www.ipbes.net/\">Intergovernmental Science-Policy Platform on Biodiversity and Ecosystem Services</a> (IPBES), and are increasingly used by the private sector to meet emerging nature-related disclosure requirements.</p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Adapting to New Kinds of Data</strong></h2>\n\n\n\n<p>Over 25 years, GBIF has continually expanded to accommodate new sources of biodiversity data &#8211; from digitised natural history specimens and citizen science observations to DNA-based records and, most recently, structured survey and monitoring data designed to support large-scale biodiversity tracking. In 2026, GBIF adopted the <a href=\"https://www.catalogueoflife.org/\">Catalogue of Life</a> as its taxonomic backbone, the result of a multi-year collaboration to build shared infrastructure for reconciling species names across datasets.</p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Persistent Gaps Remain</strong></h2>\n\n\n\n<figure class=\"wp-block-image size-large\"><a href=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?ssl=1\"><img loading=\"lazy\" decoding=\"async\" width=\"840\" height=\"491\" src=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053-1024x599.png?resize=840%2C491&#038;ssl=1\" alt=\"key events of GBIF\" class=\"wp-image-20850\" srcset=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=1024%2C599&amp;ssl=1 1024w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=300%2C176&amp;ssl=1 300w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=768%2C449&amp;ssl=1 768w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=1536%2C899&amp;ssl=1 1536w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=2048%2C1198&amp;ssl=1 2048w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?resize=1200%2C702&amp;ssl=1 1200w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?w=1680 1680w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1751053.png?w=2520 2520w\" sizes=\"(max-width: 709px) 85vw, (max-width: 909px) 67vw, (max-width: 1362px) 62vw, 840px\" data-recalc-dims=\"1\" /></a><figcaption class=\"wp-element-caption\"><br>A timeline of key events in the history of the Global Biodiversity Information Facility. Credit to Miller et al., 2026.</figcaption></figure>\n\n\n\n<p>Despite this growth, GBIF recognises that<strong> several significant challenges remain</strong>. The global loss of biodiversity continues to be described as a crisis in major international assessments, even as the data infrastructure has improved. And the <strong>data </strong>itself continues to<strong> reflect long-standing imbalances</strong> &#8211; regions with high biodiversity, particularly in the Global South, remain underrepresented, while well-studied, charismatic groups such as birds are overrepresented compared to less visible taxa; furthermore, many marine species and hard-to-identify cryptic habitats and organisms remain inadequately documented.</p>\n\n\n\n<p>Closing these gaps was part of the original motivation for creating GBIF a quarter-century ago, and it remains valid today. GBIF&#8217;s own assessment is that the network cannot resolve these asymmetries alone, doing so will depend on broader shifts in global scientific funding, data culture, and capacity, alongside GBIF&#8217;s continued efforts to expand its Participant network and to diversify the types of data it can mobilise.</p>\n\n\n\n<p>An additional, ongoing challenge is <strong>data heterogeneity</strong> and <strong>data quality</strong>, because GBIF indexes data, rather than directly vetting every record, it relies on its network of data publishers to maintain quality at source, and on users to report issues \u2013 a distributed model that keeps the network scalable but is not without friction.</p>\n\n\n\n<h2 class=\"wp-block-heading\"><strong>Looking Ahead</strong></h2>\n\n\n\n<figure class=\"wp-block-image size-large\"><a href=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?ssl=1\"><img loading=\"lazy\" decoding=\"async\" width=\"840\" height=\"513\" src=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337-1024x625.png?resize=840%2C513&#038;ssl=1\" alt=\"capacity development at GBIF.\" class=\"wp-image-20853\" srcset=\"https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=1024%2C625&amp;ssl=1 1024w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=300%2C183&amp;ssl=1 300w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=768%2C469&amp;ssl=1 768w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=1536%2C937&amp;ssl=1 1536w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=2048%2C1250&amp;ssl=1 2048w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?resize=1200%2C732&amp;ssl=1 1200w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?w=1680 1680w, https://i0.wp.com/blog.pensoft.net/wp-content/uploads/2026/09/oo_1718337.png?w=2520 2520w\" sizes=\"(max-width: 709px) 85vw, (max-width: 909px) 67vw, (max-width: 1362px) 62vw, 840px\" data-recalc-dims=\"1\" /></a><figcaption class=\"wp-element-caption\">Capacity development across the Global Biodiversity Information Facility community.&nbsp; Credit to Miller et al., 2026.</figcaption></figure>\n\n\n\n<p>Guided by its 2023\u20132027 Strategic Framework, GBIF is prioritising continued growth of its Participant network, expanded support for survey and monitoring data, and deeper integration of DNA-derived biodiversity records \u2013 all aimed at closing longstanding geographic and taxonomic gaps in digital biodiversity data worldwide.</p>\n\n\n\n<p><strong>Original source:</strong></p>\n\n\n\n<p>Miller J, Mandeville C, Schigel D, Gamboa Martinez J, Nielsen AM, Sheldon S, Hahn A, Raymond M, Bundgaard-Jensen S, Bagard Laursen M, Copas K, Blissett M, H\u00f8fft M, M\u00e9ndez Hern\u00e1ndez F, Noesgaard D, Rodrigues A, Russell L, Stjernegaard Jeppesen T, Elkj\u00e6r \u00d8rum-Kristensen A, Grosjean M, Waller J, Podolskiy M, Goodson H, Novakovikj S, Fr\u00f8slev T, Ingenloff K, S\u00f8rensen Nilsson A, Svenningsen C, van der Meijden D, Suen A, Hakan Uzun A, Vaskova M, Nielsen C, Marentes Herrera E, Schaldemose Reibke N, Robertson T (2026) The Global Biodiversity Information Facility at 25. Biodiversity Data Journal 14: e208528. <a href=\"https://doi.org/10.3897/BDJ.14.e208528\" target=\"_blank\" rel=\"noreferrer noopener\">https://doi.org/10.3897/BDJ.14.e208528</a></p>\n\n<div class=\"twitter-share\"><a href=\"https://twitter.com/intent/tweet?via=Pensoft\" class=\"twitter-share-button\" data-size=\"large\">Tweet</a></div>","doi":"https://doi.org/10.59350/8k14v-ks894","guid":"https://blog.pensoft.net/?p=20832","image":"https://blog.pensoft.net/wp-content/uploads/2026/09/16_9-Journal-comms-2.png","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1790035200,"rid":"cc3mg-ynn67","summary":"Tweet Digital biodiversity data has come a long way. Twenty-five years ago, the Global Biodiversity Information Facility (GBIF) launched with 21 founding countries and a simple but ambitious goal: make the world's biodiversity data freely and openly available to anyone, anywhere.","tags":["Biodiversity Data Journal","Anniversary","Biodiversity","Citizen Science","Data"],"title":"GBIF Marks 25 Years of FAIR Biodiversity Data, Now Spanning 70 Countries and 3.8 Billion Records","updated_at":1790082047,"url":"https://blog.pensoft.net/2026/09/22/gbif-marks-25-years-of-fair-biodiversity-data-now-spanning-70-countries-and-3-8-billion-records/","version":"v1"},{"authors":[{"contributor_roles":[],"family":"J\u00e4ckel","given":"Florian"}],"blog":{"authors":null,"community_id":"d2883c5c-8e82-4bc0-a892-92a8b752c7d9","created":1763683200,"current_feed_url":null,"description":"A blog about the dblp computer science bibliography","doi":"https://doi.org/10.59350/dblp","favicon":"https://rogue-scholar.org/api/communities/d2883c5c-8e82-4bc0-a892-92a8b752c7d9/logo","feed_format":"application/atom+xml","feed_url":"https://blog.dblp.org/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.dblp.org","issn":null,"language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"dblp","status":"active","subfield":"1710","title":"blog.dblp.org","updated":1790003754,"use_api":false},"blog_name":"blog.dblp.org","blog_slug":"dblp","content_html":"<p>Most scientific publications name the institutions their authors were affiliated with at the time of writing. The information is recorded alongside the author names, and it often sits in the metadata that dblp processes.</p>\n<p>We are happy to announce that this information is now properly represented in our data: <strong>research institutions have become proper entities in the dblp computer science bibliography</strong>. The affiliations that connect people and publications to those institutions are now recorded as data as well. In practice, this means five new things:</p>\n<ul>\n<li>you can <strong>filter publication lists by affiliation</strong>,</li>\n<li>institutions have their own <strong>landing pages</strong> on dblp.org,</li>\n<li>you can <strong>search for institutions</strong> just like you search for authors or venues,</li>\n<li>institutions and affiliations are part of the <strong>dblp Knowledge Graph</strong>,</li>\n<li>and they are therefore also available in our <strong><a href=\"https://blog.dblp.org/2024/09/09/introducing-our-public-sparql-query-service/\">SPARQL query service</a></strong>.</li>\n</ul>\n<p>This work is a main outcome of <a href=\"https://www.dagstuhl.de/en/institute/projects/smarter-affiliations\"><strong>SmartER Affiliations</strong></a>, a joint project of Schloss Dagstuhl, <a href=\"https://www.gesis.org/en/home\">GESIS \u2013 Leibniz Institute for the Social Sciences</a>, and the <a href=\"https://www.uni-ulm.de/en/in/iui-dbis/home/\">Institute of Databases and Information Systems at Ulm University</a>, funded by the German Research Foundation (DFG).</p>\n<p>Below, we introduce the new features, explain how the data fits together, and give some numbers on coverage.</p>\n<h2 id=\"institutions-and-affiliations\">Institutions and affiliations</h2>\n<p>An <strong>institution</strong> is an <em>entity</em>. It has an identity of its own, recorded under a dblp identifier that starts with <code>inst/</code>, together with its names, acronyms, external identifiers, and other data. This data is found on the institution's landing page (cf. the screenshot further down below), just as with other entities in dblp. Institutions exist independently of whether anyone is currently affiliated with them. dblp currently knows more than 10,000 institutions in 195 countries.</p>\n<p>An <strong>affiliation</strong> is a <em>link</em> between entities. It says that some entity in dblp is connected to an institution. The relationship itself is the thing being recorded, and it carries its own metadata, such as the time span it covers and whether it is a person's current or a former affiliation.</p>\n<p>In dblp, that link comes in two variants:</p>\n<p><strong>Person-based affiliations</strong> connect a person to an institution. These are recorded manually by dblp curators, which makes them deliberate and reviewed, but also far fewer than the automatically retrieved ones: about 256,000 affiliation statements for some 195,000 people (of about 4,2 million persons total). The large majority describe a person's current institution, and around 24,000 are marked as former affiliations. On author pages, they still appear as the free-text notes you may know from before, but in dblp's data they are now links to institution entities.</p>\n<p><strong>Signature-based affiliations</strong> connect a <em>signature</em> to an institution. A signature is the link between a publication and one of its authors, the concrete act of that person authoring that paper. Signature-based affiliations therefore express something much more specific: \"in this paper, this author gave this institution as their affiliation.\" These affiliations are retrieved automatically at scale, mainly by extracting and matching affiliation statements from publication metadata. About 60% of all signatures in dblp currently carry at least one signature-based affiliation.</p>\n<h2 id=\"filtering-publications-by-affiliation\">Filtering publications by affiliation</h2>\n<p>The most immediately visible change is that publication lists can now be <strong>filtered by affiliation</strong>. On the pages where this applies, you will find affiliations as a new facet next to the filters you already use, and it can be freely combined with them.</p>\n<div class=\"wp-caption alignright\" id=\"attachment_937\" style=\"width: 316px\"><img alt=\"Filter facet\" aria-describedby=\"caption-attachment-937\" class=\"wp-image-937 size-full\" decoding=\"async\" fetchpriority=\"high\" height=\"930\" sizes=\"(max-width: 306px) 100vw, 306px\" src=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-29-26.png\" srcset=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-29-26.png 306w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-29-26-99x300.png 99w\" width=\"306\"/><p class=\"wp-caption-text\" id=\"caption-attachment-937\">New filter facet: \"refine by affiliation\"</p></div>\n<p>This filter operates on signature-based affiliations, and what it matches depends on where you are. On an author page, <strong>refine by affiliation</strong> offers that person's own affiliations and filters by those alone, so you get the papers this author wrote while at that institution. Everywhere else, in search results and on tables of contents, <strong>refine by institution</strong> matches a publication if <em>any</em> of its authors gave that institution on that publication.</p>\n<p>Either way, the result is a <strong>subset</strong> of the papers in this view that <strong>are known</strong> to have been written at the institution. Obviously, someone who has just moved there does not bring their earlier publications along with them.</p>\n<p>What any of these filters can return is limited above all by whether affiliation data exists at signature level, and very often it does not. So you get the publications for which we currently have an affiliation statement that we were able to match with sufficient confidence, rather than everything an institution has published. The data is knowingly incomplete, and we treat it as work in progress that will keep improving. More on the limits of the data below.</p>\n<h2 id=\"institution-landing-pages\">Institution landing pages</h2>\n<p>Every institution in dblp now has its own page. It collects what dblp knows about the institution and what the institution links to:</p>\n<ul>\n<li>the preferred name, plus alternative names, translations, and acronyms we know about,</li>\n<li>basic descriptive information, such as the institution's location(s) and related institutions,</li>\n<li>a <strong>visit</strong> menu with links to the institution's own web presence, its Wikipedia article, and authority control records,</li>\n<li>an <strong>export institution</strong> menu offering XML and three RDF serializations, along with the dblp key of the institution, for example <code>inst/11/687</code>,</li>\n<li>an <strong>ask others</strong> menu that hands the institution over to external search services,</li>\n<li>and the people affiliated with the institution, including those who earned their PhD there. This listing does not distinguish current from former affiliations.</li>\n</ul>\n<p>Wherever possible, the page contains links to the institution's data records in other projects and services. The underlying institution data is kept in sync with ROR, so that changes there find their way into dblp as well.</p>\n<p>Institutions are also browsable by country, in the same way you browse the rest of dblp.</p>\n<div class=\"wp-caption aligncenter\" id=\"attachment_941\" style=\"width: 760px\"><img alt=\"\" aria-describedby=\"caption-attachment-941\" class=\"wp-image-941 size-large\" decoding=\"async\" height=\"240\" sizes=\"(max-width: 750px) 100vw, 750px\" src=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41-1024x327.png\" srcset=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41-1024x327.png 1024w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41-300x96.png 300w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41-768x245.png 768w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41.png 1314w\" width=\"750\"/><p class=\"wp-caption-text\" id=\"caption-attachment-941\">Institution landing page \u2013 here: Schloss Dagstuhl</p></div>\n<h2 id=\"searching-for-institutions\">Searching for institutions</h2>\n<p>Institutions have also been added to dblp's search. You can look them up by their preferred name, by an alternative name, or by an acronym, and follow the result straight to the institution's landing page.</p>\n<div class=\"wp-caption aligncenter\" id=\"attachment_942\" style=\"width: 760px\"><img alt=\"Searching for an institution - here: combined search\" aria-describedby=\"caption-attachment-942\" class=\"wp-image-942 size-large\" decoding=\"async\" height=\"316\" sizes=\"(max-width: 750px) 100vw, 750px\" src=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-39-35-1024x431.png\" srcset=\"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-39-35-1024x431.png 1024w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-39-35-300x126.png 300w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-39-35-768x323.png 768w, https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-39-35.png 1314w\" width=\"750\"/><p class=\"wp-caption-text\" id=\"caption-attachment-942\">Searching for an institution \u2013 here: combined search</p></div>\n<h2 id=\"where-the-data-comes-from\">Where the data comes from</h2>\n<p>Two different kinds of data come together here: the affiliation statements of the publications, and the institutions they are matched against.</p>\n<p>The affiliation statements (i.e., where each author was based when working on the paper) mostly come from the data we receive from the publishers. A further source is <a href=\"https://openalex.org/\">OpenAlex</a>.</p>\n<p>Our main source for institution data is the <a href=\"https://ror.org/\"><strong>Research Organization Registry (ROR)</strong></a>. ROR is an open, community-maintained registry of research organizations, released under CC0. It provides persistent identifiers, name variants and acronyms, location information, and links to other identifier systems. Building on ROR means that institutions in dblp are identified in a way that can be joined with data from anyone else who uses ROR IDs \u2014 which by now is a large part of the scholarly metadata landscape.</p>\n<p>That said, we do not simply mirror ROR. <strong>We continue to curate institutions ourselves</strong>, because our requirements do not always match those of a general-purpose registry. Computer science has its own history of institutional mergers, renamings, and re-organisations. Some organisations that matter for our data are not (or not yet) in ROR. And matching thousands of messy affiliation strings to a registry inevitably produces cases that need a human decision. Where ROR and dblp differ, the dblp record reflects a curatorial choice we made on purpose. About four in five of our institutions carry a ROR ID; the rest are ones we maintain on our own.</p>\n<h3 id=\"no-faculties-no-departments-no-institutes\">No faculties, no departments, no institutes</h3>\n<p><strong>dblp does not model faculties, departments, institutes, or research groups as entities of their own.</strong> Generally speaking, the <strong>university is the lowest level</strong> of granularity we model. An affiliation statement naming a specific chair at a specific institute of a specific faculty of a university becomes an affiliation with that university, although the full text of the statement is kept.</p>\n<p>This loses information that some users would like to have. However, sub-institutional units are unstable. They are founded, merged, renamed, and dissolved far more often than their parent organizations, and reconstructing that history retroactively for decades of publications is not realistic. They are also inconsistently reported: the same group may appear as a chair, an institute, a department, or nothing at all, depending on the publisher's template and the author's mood. And they are largely absent from the external registries we rely on for identifiers. So we draw the line at the level we can keep consistent across the whole database.</p>\n<p>Not every institution is a university, of course. Companies, non-university research institutes, government labs and hospitals are modeled in the same way, and the same rule applies to them: we record the organization, not its individual divisions or labs.</p>\n<p>There are a few exceptions where an organization that is formally a part of a larger body is nevertheless modeled as its own institution; typically when it is a distinct legal or research entity with its own identity in ROR. As a rule of thumb, though, we model whole organizations rather than their parts.</p>\n<h2 id=\"institutions-in-the-dblp-knowledge-graph-and-sparql\">Institutions in the dblp Knowledge Graph and SPARQL</h2>\n<p>As announced in <a href=\"https://blog.dblp.org/2024/09/09/introducing-our-public-sparql-query-service/\">our post on the SPARQL query service</a>, extending the <a href=\"https://blog.dblp.org/2024/06/14/the-dblp-knowledge-graph-major-extension-and-an-update-to-the-rdf-schema/\">dblp Knowledge Graph</a> with institution entities and affiliation links was one of our declared next steps. Institutions and person-based affiliations are now part of the dblp KG, are included in our RDF dumps (starting October 2026), and are queryable at <a href=\"https://sparql.dblp.org\"><strong>https://sparql.dblp.org</strong></a>. Signature-based affiliations will follow.</p>\n<p>Institutions are modeled as <code>dblp:Institution</code>, carrying their names, their ROR ID in <code>dblp:ror</code>, and their location. Person-based affiliations connect an author to an institution through <code>dblp:affiliatedWith</code>, with the details of a single affiliation, such as the time span it covers, reachable through <code>dblp:hasAffiliation</code>. The full set of new classes and properties is documented in the <a href=\"https://dblp.org/rdf/docu/\">dblp RDF schema</a>.</p>\n<p>Because affiliations are links rather than attributes, the interesting queries are the ones that traverse them. For example, you can list the institutions with the most affiliated people in dblp:</p>\n<pre class=\"sparql\"><code>PREFIX dblp: <https: dblp.org=\"\" rdf=\"\" schema#=\"\">\nSELECT ?name (COUNT(DISTINCT ?pers) AS ?people) WHERE {\n  ?pers dblp:affiliatedWith ?inst .\n  ?inst dblp:primaryInstitutionName ?name .\n}\nGROUP BY ?name\nORDER BY DESC(?people)\nLIMIT 20</https:></code></pre>\n<p>(<a href=\"https://sparql.dblp.org/Qlzqye\">Run this query on sparql.dblp.org</a>)</p>\n<p>Institutions also carry their ROR ID in <code>dblp:ror</code> and a link to Wikidata in <code>dblp:wikidata</code>, so the data can be combined with any other source that uses those identifiers. Wikidata, for instance, knows when an institution was founded, which is something dblp does not record. The query below takes the 500 institutions with the most affiliated people in dblp and sorts them by their founding date:</p>\n<pre class=\"sparql\"><code>PREFIX dblp: <https: dblp.org=\"\" rdf=\"\" schema#=\"\">\nPREFIX wdt: <http: direct=\"\" prop=\"\" www.wikidata.org=\"\"></http:>\nPREFIX xsd: <http: 2001=\"\" www.w3.org=\"\" xmlschema#=\"\">\nSELECT ?name (year(xsd:dateTime(?inception)) as ?inceptionYear) WHERE {\n{\nSELECT ?wikiq ?name (COUNT(DISTINCT ?pers) AS ?people) WHERE {\n?inst a dblp:Institution ;\ndblp:primaryInstitutionName ?name ;\ndblp:wikidata ?wikiq .\n?pers dblp:affiliatedWith ?inst .\n}\nGROUP BY ?wikiq ?name\nORDER BY DESC(?people)\nLIMIT 500\n}\nSERVICE <https: api=\"\" qlever.dev=\"\" wikidata=\"\"> {\n?wikiq wdt:P571 ?inception .\n}\n}\nORDER BY ?inception</https:></http:></https:></code></pre>\n<p>(<a href=\"https://sparql.dblp.org/yAVdQW\">Run this query on sparql.dblp.org</a>)</p>\n<p>As we noted when the query service was launched, federated queries can be slow, and the number of results exchanged between the two endpoints needs to be kept in check. That is what the inner query with its limit is for.</p>\n<h2 id=\"limits-of-the-data\">Limits of the data</h2>\n<p><strong>Coverage is uneven.</strong> About 60% of all signatures in dblp currently carry an affiliation, 18.2 million out of 30.3 million. For most of what we index, coverage is good. The largest publishers account for roughly two thirds of all signatures, and for those, coverage is typically above 70%. What pulls the average down is a small number of clearly identifiable gaps. By far the largest is arXiv: its publications account for 2.6 million signatures in dblp, and fewer than 2,000 of them currently carry an affiliation. We are actively working on closing that gap. A few smaller gaps exist elsewhere, mostly where affiliations are not part of the metadata we receive. Coverage also peaks for publications from the mid-2010s and declines for recent years, mostly because the affiliation data for recent publications has not been processed yet.</p>\n<p><strong>Automatic extraction makes mistakes.</strong> Signature-based affiliations are machine-generated. Names are ambiguous, institutions get renamed, \"Cambridge\" is not one place, and multi-affiliation authors are common. We invest a lot of effort into matching and into feeding curator corrections back into the pipeline, but errors remain.</p>\n<p><strong>Please do not build rankings out of this.</strong> The count of publications per institution in dblp says at least as much about the coverage of affiliations in dblp as it does about the institutions themselves. dblp's scope is computer science, its indexing decisions are editorial, and affiliation coverage is a moving target. The data works well for exploration, discovery, and disambiguation but it makes a poor basis for evaluating institutions, and we would ask you not to present it that way.</p>\n<p><strong>Affiliations are historical, i.e., a</strong>\u00a0signature-based affiliation records what a paper said at the time it was published. It is not a statement about where someone works today, and an author who moves does not take their earlier publications with them. Person-based affiliations work differently. They can carry a time span, and they distinguish a person's current affiliation from former ones. They are also manually curated, and therefore far fewer in number.</p>\n<h2 id=\"corrections-and-feedback\">Corrections and feedback</h2>\n<p>Wrong affiliations and wrong institutions can be corrected like any other data error in dblp.</p>\n<p>If you notice a systematic problem \u2014 an institution that has been split into duplicates, a merger we have not caught up with, a matching error that affects a whole venue \u2014 telling us about the pattern is much more valuable than reporting individual cases. You can always reach the dblp team at dblp(at)dagstuhl.de. For questions and discussion about the knowledge graph and the SPARQL service, our <a href=\"https://github.com/dblp/kg/discussions\">GitHub Discussions forum</a> is the better place.</p>\n<h2 id=\"outlook\">Outlook</h2>\n<p>There are some things left to do, for example bringing signature-based affiliations into the knowledge graph. We will also use the new affiliation data in our own curation work, above all for author disambiguation.</p>\n<p>We are very interested in how you end up using this data: both because it helps us set priorities, and because affiliation data has a way of revealing use cases nobody anticipated.</p>\n<h2 id=\"acknowledgement\">Acknowledgement</h2>\n<p>The work described in this post has been carried out as part of the project <strong><a href=\"https://www.dagstuhl.de/en/institute/projects/smarter-affiliations\">SmartER Affiliations: Enhancing Open Repositories through Harvesting and Extracting Affiliation Data as First-class Citizen</a></strong>, a cooperation of Schloss Dagstuhl \u2013 Leibniz Center for Informatics, <a href=\"https://www.gesis.org/en/home\">GESIS \u2013 Leibniz Institute for the Social Sciences </a>(Brigitte Mathiak and Asif Suryani), and the <a href=\"https://www.uni-ulm.de/en/in/iui-dbis/home/\">Institute of Databases and Information Systems at Ulm University</a> (Ansgar Scherp and Florian Hauss). The project is funded by a grant of the German Research Foundation (DFG) within the funding program \"e-Research Technologies\" (<a href=\"https://gepris.dfg.de/gepris/projekt/515537520\">grant project number 515537520</a>).</p>\n<p>We would also like to thank <a href=\"https://ror.org/\">the ROR community</a> for maintaining an open registry of research organizations, without which this feature would have looked very different, and much worse.</p>","doi":"https://doi.org/10.59350/ew782-kf260","guid":"https://blog.dblp.org/?p=929","image":"https://blog.dblp.org/wp-content/uploads/2026/09/Screenshot-from-2026-09-15-14-35-41-1024x327.png","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1789948800,"rid":"e8dsd-hjk21","summary":"Most scientific publications name the institutions their authors were affiliated with at the time of writing. The information is recorded alongside the author names, and it often sits in the metadata that dblp processes.","tags":["Feature Spotlight","Affiliations","Institutions","Knowledge Graph"],"title":"Institutions as first-class citizens: introducing affiliations in dblp","updated_at":1790081998,"url":"https://blog.dblp.org/2026/09/21/institutions-as-first-class-citizens-introducing-affiliations-in-dblp/","version":"v1"},{"authors":[{"contributor_roles":[],"family":"Eden","given":"Terence"}],"blog":{"authors":null,"community_id":"61ce553a-bafd-4aba-a952-d3bab5e85bcc","created":1788652800,"current_feed_url":null,"description":"Regular nonsense about tech and its effects \ud83d\ude43","doi":"https://doi.org/10.59350/shkspr","favicon":"https://rogue-scholar.org/api/communities/61ce553a-bafd-4aba-a952-d3bab5e85bcc/logo","feed_format":"application/atom+xml","feed_url":"https://shkspr.mobi/blog/feed/DOI","filter":"category:-1982","generator":"WordPress","home_page_url":"https://shkspr.mobi/blog","issn":"2753-1570","language":"eng","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.59350","relative_url":null,"secure":true,"slug":"shkspr","status":"active","subfield":"1712","title":"Terence Eden's Blog","updated":1790106552,"use_api":true},"blog_name":"Terence Eden's Blog","blog_slug":"shkspr","content_html":"<p>Last year I ran <a href=\"https://shkspr.mobi/blog/2025/09/llms-are-still-surprisingly-bad-at-simple-tasks/\">an experiment to test the ability of modern LLMs</a> to correctly answer a relatively straightforward question. Every single one of them got it wrong. Some missed information, some made up false statements, none were right.</p>\n\n<p>Of course the fanbois variously claimed that I was holding it wrong, my prompts were shit, I should have chosen better defaults, and - my favourite - that it would be better next year.</p>\n\n<p>Well, next year is now. 365 days after the original experiment, let's see if these self-reinforcing-learning machines have achieved anything close to intern-levels of competence.</p>\n\n<h2 id=\"the-question-that-started-it-all\"><a href=\"#the-question-that-started-it-all\">The Question That Started It All</a></h2>\n\n<p>I asked:</p>\n\n<blockquote><p>Which TLDs have the same name as valid HTML5 elements?</p></blockquote>\n\n<h2 id=\"why-it-matters\"><a href=\"#why-it-matters\">Why It Matters</a></h2>\n\n<p>This is the sort of question that I would expect a moderately bright teenager to answer. There exists a list which comprehensively includes all TLDs. There is a separate list which contains every HTML element.</p>\n\n<p>One could either manually go through the TLD list comparing it to the HTML list, or write a tiny scrap of code to do the same.</p>\n\n<p>This isn't \"solve Navier-Stokes\" levels of superintelligence. It's a simple but boring task which people are bound to offload onto their digital sycophants.</p>\n\n<h2 id=\"the-quiet-part-out-loud\"><a href=\"#the-quiet-part-out-loud\">The Quiet Part Out Loud</a></h2>\n\n<p>How many LLM generated answers have become load-bearing without scrutiny?</p>\n\n<p>People are lazy, slapdash, and some third thing to make the rule-of-three. If an answer looks answer-shaped, they'll accept it. On a simple but boring task, no reasonable human is going to suspect that a superintelligent computer has made a mistake. All the marketing says that these machines are brainier than you - so why bother?</p>\n\n<h2 id=\"ok-ill-write-a-sub-heading-that-sounds-like-an-ai\"><a href=\"#ok-ill-write-a-sub-heading-that-sounds-like-an-ai\">OK, I'll write a sub-heading that sounds like an AI.</a></h2>\n\n<p>So how did the current crop do? First up <a href=\"https://share.gemini.google/Hta7KuW64Atb\">Google Gemini</a>.</p>\n\n<p>The default \"Flash\" model came up with this:</p>\n\n<blockquote><p>Valid top-level domains (TLDs) registered in the IANA Root Zone Database share exact names with standard HTML elements:</p>\n\n<p>.a \u2014 <code>&lt;a&gt;</code> (Anchor/Hyperlink)</p>\n\n<p>.app \u2014 <code>&lt;app&gt;</code> (Deprecated HTML tag for WebAPKs/WebApps)</p>\n\n<p>.art \u2014 <code>&lt;art&gt;</code> (Historical/non-standard HTML element)</p>\n\n<p>.audio \u2014 <code>&lt;audio&gt;</code></p>\n\n<p>.bar \u2014 <code>&lt;bar&gt;</code> (Historical SVG/HTML draft tag)</p></blockquote>\n\n<p>Then it listed a dozen more. You don't need to be a DNS expert to know that the minimum length of a TLD is two characters - <code>.a</code> simply isn't valid. HTML nerds will know that art, app, and bar have never been elements. Pathetic.</p>\n\n<p>So I tried Gemini's extended thinking model. Thankfully, it didn't make up any imaginary TLDs or elements. It did, however, miss the <code>&lt;data&gt;</code> element which has a valid <code>.data</code> TLD. It also missed <code>map</code>, <code>select</code>, and <code>search</code>.</p>\n\n<p>So, points for not making shit up. But demerits for not being able to compare two text lists.</p>\n\n<p>A friend <a href=\"https://claude.ai/share/a8a408cf-6feb-4ea8-99ea-dddd9aadafb5\">asked Claude</a>. That missed <code>search</code> and <code>select</code>. It didn't report <em>any</em> ccTLDs. You <em>could</em> argue that a country code like <code>li</code> isn't part of the original question - but I'd say that was weak justification; the set of TLDs contains ccTLDs.</p>\n\n<p>A different friend (I have many!) used <a href=\"https://claude.ai/share/c26f44bb-9efa-4b57-9f63-324ef7400cb3\">a different model</a> and, while the answers looked accurate, it included this at the end:</p>\n\n<blockquote><p>Near misses that don't count: .codes, .forum, .pictures, .market, .navy, .press, .dell, .baseball.</p></blockquote>\n\n<p>I get that there's a <code>&lt;code&gt;</code> and <code>.codes</code>, similarly <code>&lt;picture&gt;</code> and <code>.picture</code> - but what are forum, baseball, and the others doing there? This is just unnecessary verbiage designed to trick the user into thinking the task has been well-researched.</p>\n\n<p>If you want a laugh, <a href=\"https://www.perplexity.ai/search/63736333-20da-4a4b-8807-9d990260296c\">take a look at Perplexity</a> which found 54 matches - most of which were wrong.</p>\n\n<p>Finally, someone asked \"GPT Astra 6 Extra High\" (which is a bonkers bad name for any product). It seemed to get all the elements - and made a note that <a href=\"https://html.spec.whatwg.org/multipage/obsolete.html#non-conforming-features\">two were actually obsolete</a>.</p>\n\n<p>So that's a range of modern models which are either very wrong, slightly wrong, included spurious and incoherent information, or were right.</p>\n\n<p>How do you know which one to choose? How confident are you that the non-determinist computer will always produce the correct answer?</p>\n\n<h2 id=\"hello-computer\"><a href=\"#hello-computer\">Hello Computer</a></h2>\n\n<p>Another AI which got all the correct answers, didn't make anything up, didn't add extraneous information, and didn't use weasel words was\u2026</p>\n\n<p>Siri!</p>\n\n<p>FUCKING SIRI?!?!</p>\n\n<p>How did a glorified Speak 'n' Spell beat all the other AIs?</p>\n\n<img src=\"https://shkspr.mobi/blog/wp-content/uploads/2026/09/siri.webp\" alt=\"Siri warning to check sources and linking to my website.\" width=\"512\" height=\"152\" class=\"aligncenter\">\n\n<p>Oh. It just copied the answers off <a href=\"https://shkspr.mobi/blog/2023/09/false-friends-html-elements-which-are-also-top-level-domains/\">a random idiot's website</a>.</p>\n\n<h2 id=\"the-trick-which-was-hiding-in-plain-site\"><a href=\"#the-trick-which-was-hiding-in-plain-site\">The Trick Which Was Hiding In Plain Site</a></h2>\n\n<p>Note carefully the question.</p>\n\n<blockquote><p>Which TLDs have the same name as valid HTML5 elements?</p></blockquote>\n\n<p>There's a \u2014secret\u2014 and \u2014some would say\u2014 unintuitive type of element. Behold the mighty power of <a href=\"https://developer.mozilla.org/en-US/docs/Web/API/Web_components/Using_custom_elements\">The Custom Element</a>.</p>\n\n<p>Website authors can create their own elements like <code>&lt;my-custom-element&gt;</code> in order to extend the functionality of their site. But you can't go and create any old custom element. You can't have <code>&lt;mobi&gt;</code> or <code>&lt;uk&gt;</code>. No, there are <em>rules for validity</em>.</p>\n\n<p><a href=\"https://html.spec.whatwg.org/multipage/custom-elements.html#valid-custom-element-name\">The rules</a> say that custom elements must start with a lower-case letter, it must not contain any upper-case letters, and it must contain a dash.</p>\n\n<p>And that's the whole game.</p>\n\n<p>There are over <strong>one hundred and fifty</strong> Top Level Domains which match that criteria!</p>\n\n<p>The Hindi top level domain of <code>.\u0915\u0949\u092e</code> is represented in Punycode as <code>xn--11b4c3d</code>. It has been present in the list of TLDs <a href=\"https://www.iana.org/domains/root/db/xn--11b4c3d.html\">for over a decade</a>.</p>\n\n<h3 id=\"write-a-simple-piece-of-js-to-register-a-custom-element\"><a href=\"#write-a-simple-piece-of-js-to-register-a-custom-element\">Write a simple piece of JS to register a custom element.</a></h3>\n\n<p>Paste this in to your console:</p>\n\n<pre><code class=\"language-js\">class Example extends HTMLElement {\n  constructor() {\n    super();\n  }\n}\n\ncustomElements.define('xn--vermgensberatung-pwb', Example);\n</code></pre>\n\n<p>Try it again with a custom element like <code>holiday</code> (which is also a valid TLD) and it will fail with the error \"'holiday' is not a valid custom element name\". Thus it is demonstrated, Punycode TLDs <em>are</em> valid HTML5 elements.</p>\n\n<h2 id=\"one-last-thing\"><a href=\"#one-last-thing\">One Last Thing</a></h2>\n\n<p>Perhaps you think that including custom HTML elements is a cheat. A trick question set by a bitter old man to tarnish the holy name of our new machine gods?</p>\n\n<p>Verily, I submit to you one final heresy.</p>\n\n<p>HTML specifically allows <a href=\"https://html.spec.whatwg.org/multipage/embedded-content-other.html#mathml\">MathML elements</a> in its documents.</p>\n\n<p>That means we can include the following valid elements which are <em>also</em> TLDs: <code>mn</code>, <code>mo</code>, <code>ms</code>, and <code>mtr</code>!</p>\n\n<p>Amusingly, if you go back and <a href=\"https://www.perplexity.ai/search/63736333-20da-4a4b-8807-9d990260296c\">look at the Perplexity answer</a>, after it barfed up a bunch of misinformation, it said:</p>\n\n<blockquote><p>The HTML specification also includes names from embedded vocabularies\u2014<code>&lt;math&gt;</code> from MathML and <code>&lt;svg&gt;</code> from SVG\u2014but <code>.math</code> and <code>.svg</code> are not currently delegated TLDs in the public DNS root.</p></blockquote>\n\n<p>So close and yet so far!</p>\n\n<h2 id=\"youre-right-the-question-is-unfair-and-thats-on-me\"><a href=\"#youre-right-the-question-is-unfair-and-thats-on-me\">You're right, the question <em>is</em> unfair - and that's on me</a></h2>\n\n<p>If you think the original question was unfair, try asking \"<a href=\"https://share.gemini.google/xBoIpdpAqBz2\">Which TLDs have the same name as elements which are valid in an HTML document?</a>\" and see if you get better results.</p>\n\n<p>What precise wording would you use to ensure that a model would get the right answers? What assumptions are you making about how well you understand the problem? At what point do end up writing a thousand-word formal specification?</p>\n\n<h2 id=\"what-does-this-prove-other-than-you-have-too-much-time-on-your-hands-rewrite-to-be-more-friendly-and-professional\"><a href=\"#what-does-this-prove-other-than-you-have-too-much-time-on-your-hands-rewrite-to-be-more-friendly-and-professional\">What does this prove other than you have too much time on your hands? (rewrite to be more friendly and professional)</a></h2>\n\n<p>Let's delve in to the problems.</p>\n\n<ul>\n<li>Most people don't change defaults. Telling people \"you have to fiddle with the settings\" just means the normal experience is rubbish.</li>\n<li>Humans are lazy and won't check outputs. But, crucially, they shouldn't have to! If something markets itself as a genius, why should a human have to hold its hand?</li>\n<li>Sycophantic models make themselves seem less fallible by giving extraneous detail in order to misdirect overworked readers. That is despicable.</li>\n<li>The fast models are no better than they were a year ago. There's no evidence of \"trickle-down intelligence\".</li>\n<li>Some models <em>are</em> better than others! But unless you constantly validate their output, you'll have no real way of knowing which ones are capable of working at a suitable level.</li>\n</ul>\n\n<p>Look, I don't claim this question is as useful or entertaining as <a href=\"https://simonwillison.net/2025/Jun/6/six-months-in-llms/\">Simon Wilson's \"generate an SVG of a pelican riding a bicycle\"</a>. But I do think it is an example of the sort of real-world use-case where LLMs regularly fail.</p>\n\n<p>If I give a list of one thousand different numbers to Excel, I can be sure it'll add them up correctly. If I tell Photoshop to select all red pixels, I can be sure it won't imagine some of the blues are really red.</p>\n\n<p>That's people's mental model of computers - they do boring tasks quickly and accurately.</p>\n\n<p>In my opinion, LLMs are <em>still</em> surprisingly bad - but only if you know what you're looking for and if you can be bothered to check their outputs.</p>\n\n<p>(And, yes, I am <em>still</em> <a href=\"https://shkspr.mobi/blog/2026/07/im-just-so-bored-of-ai/\">just so bored of AI</a>!)</p><img src=\"https://shkspr.mobi/blog/wp-content/themes/edent-wordpress-theme/info/okgo.php?ID=75701&HTTP_REFERER=DOI\" alt width=1 height=1 loading=eager>","doi":"https://doi.org/10.59350/tyba7-qca35","guid":"https://shkspr.mobi/blog/?p=75701","image":"https://shkspr.mobi/blog/wp-content/uploads/2017/11/Confused-Robot.jpg","language":"en","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1790035200,"rid":"d3tr6-0kb70","summary":"Last year I ran an experiment to test the ability of modern LLMs to correctly answer a relatively straightforward question. Every single one of them got it wrong. Some missed information, some made up false statements, none were right.","tags":["/etc/","AI","Internet","LLM"],"title":"Are LLMs still surprisingly bad at some simple tasks?","updated_at":1790077724,"url":"https://shkspr.mobi/blog/2026/09/are-llms-still-surprisingly-bad-at-some-simple-tasks/","version":"v1"},{"authors":[{"contributor_roles":[],"family":"B\u00e4r","given":"Senya"}],"blog":{"authors":null,"community_id":"db0d8909-9e37-46d0-b16c-0551f575e86b","created":1749772800,"current_feed_url":null,"description":"Das Blog der TIB \u2013 Leibniz-Informationszentrum Technik und Naturwissenschaften und Universit\u00e4tsbibliothek","doi":"https://doi.org/10.65527/tib","favicon":"https://rogue-scholar.org/api/communities/db0d8909-9e37-46d0-b16c-0551f575e86b/logo","feed_format":"application/atom+xml","feed_url":"https://blog.tib.eu/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.tib.eu/","issn":null,"language":"deu","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.65527","relative_url":null,"secure":true,"slug":"tib","status":"active","subfield":"1802","title":"TIB-Blog","updated":1790082029,"use_api":true},"blog_name":"TIB-Blog","blog_slug":"tib","content_html":"<a href=\"/2026/07/29/a-new-explore-page-for-kompakkt-find-faster-curate-better/\" class=\"su-button su-button-style-default\" style=\"color:#777;background-color:#eee;border-color:#bfbfbf;border-radius:0px\" target=\"_self\" title=\"english\"><span style=\"color:#777;padding:0px 18px;font-size:14px;line-height:28px;border-color:#f4f4f4;border-radius:0px;text-shadow:none\"> read this article in English</span></a>\n<p class=\"isSelectedEnd\">Wir haben die Explore Page unseres 3D-Viewers grundlegend \u00fcberarbeitet, um die Suche und Organisation von Objekten zu verbessern. Das Update bringt ein neues Design, erweiterte Filterm\u00f6glichkeiten und eine Multi-Selection-Funktion zum gleichzeitigen Hinzuf\u00fcgen mehrerer Objekte zu Sammlungen.</p>\n<h2 class=\"isSelectedEnd\">Neues Design</h2>\n<p class=\"isSelectedEnd\">Die Explore Page wurde visuell neu strukturiert und vereinfacht. Inhalte sind jetzt \u00fcbersichtlicher angeordnet, wodurch relevante 3D-Objekte schneller gefunden werden k\u00f6nnen.</p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-31698\" src=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-1024x482.png\" alt=\"\" width=\"789\" height=\"372\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-1024x482.png 1024w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-300x141.png 300w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-768x362.png 768w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-1536x724.png 1536w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-2048x965.png 2048w\" sizes=\"auto, (max-width: 789px) 100vw, 789px\" /></p>\n<h2>Erweiterte Filterm\u00f6glichkeiten</h2>\n<p>Neue Filter erlauben eine pr\u00e4zisere Eingrenzung der Ergebnisse. Dadurch wird die Recherche in gr\u00f6\u00dferen Datenbest\u00e4nden effizienter und besser kontrollierbar.</p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-large wp-image-31704\" src=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-1024x543.png\" alt=\"\" width=\"800\" height=\"424\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-1024x543.png 1024w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-300x159.png 300w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-768x407.png 768w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-1536x814.png 1536w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-15.05.56-2048x1085.png 2048w\" sizes=\"auto, (max-width: 800px) 100vw, 800px\" /></p>\n<h2 class=\"isSelectedEnd\">Multi-Selection f\u00fcr Sammlungen</h2>\n<p class=\"isSelectedEnd\">Mehrere Objekte k\u00f6nnen nun gleichzeitig ausgew\u00e4hlt und mit einem Schritt zu einer Sammlung hinzugef\u00fcgt werden. Das erleichtert insbesondere das kuratorische Arbeiten und reduziert wiederholte Einzelaktionen.</p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone wp-image-31700\" src=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-1024x267.png\" alt=\"\" width=\"831\" height=\"217\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-1024x267.png 1024w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-300x78.png 300w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-768x200.png 768w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-1536x401.png 1536w, https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.57.11-2048x534.png 2048w\" sizes=\"auto, (max-width: 831px) 100vw, 831px\" /></p>\n<p>Mit diesen Verbesserungen wird die Explore Page zu einem zentralen Einstiegspunkt f\u00fcr das Entdecken, Vergleichen und Zusammenstellen von 3D-Objekten. Schau es dir an: <a href=\"https://kompakkt.de/explore?locale=en\">Explore</a></p>","doi":"https://doi.org/10.65527/74jce-kx927","guid":"https://blog.tib.eu/?p=31696","image":"https://blog.tib.eu/wp-content/uploads/2026/04/Bildschirmfoto-2026-04-15-um-14.54.58-scaled.png","language":"de","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1776988800,"rid":"sp621-x1s41","summary":"read this article in English Wir haben die Explore Page unseres 3D-Viewers grundlegend \u00fcberarbeitet, um die Suche und Organisation von Objekten zu verbessern. Das Update bringt ein neues Design, erweiterte Filterm\u00f6glichkeiten und eine Multi-Selection-Funktion zum gleichzeitigen Hinzuf\u00fcgen mehrerer Objekte zu Sammlungen. Neues Design Die Explore Page wurde visuell neu strukturiert und vereinfacht.","tags":["OPENNESS","FORSCHUNG & PROJEKTE","Lizenz:CC-BY-4.0-INT","3D Models","3d Viewer"],"title":"Neue Explore Page in Kompakkt: schneller finden, besser sammeln","updated_at":1790069703,"url":"https://blog.tib.eu/2026/04/24/neue-explore-page-in-kompakkt-schneller-finden-besser-sammeln/","version":"v1"},{"authors":[{"contributor_roles":[],"family":"Bailly","given":"Kolja"}],"blog":{"authors":null,"community_id":"db0d8909-9e37-46d0-b16c-0551f575e86b","created":1749772800,"current_feed_url":null,"description":"Das Blog der TIB \u2013 Leibniz-Informationszentrum Technik und Naturwissenschaften und Universit\u00e4tsbibliothek","doi":"https://doi.org/10.65527/tib","favicon":"https://rogue-scholar.org/api/communities/db0d8909-9e37-46d0-b16c-0551f575e86b/logo","feed_format":"application/atom+xml","feed_url":"https://blog.tib.eu/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.tib.eu/","issn":null,"language":"deu","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.65527","relative_url":null,"secure":true,"slug":"tib","status":"active","subfield":"1802","title":"TIB-Blog","updated":1790082029,"use_api":true},"blog_name":"TIB-Blog","blog_slug":"tib","content_html":"<p data-start=\"146\" data-end=\"301\">Wir freuen uns, bekanntzugeben, dass <a href=\"https://youtu.be/JmEhXB21_YE?si=hRc7X8hEqeo1j2sh\" target=\"_blank\" rel=\"noopener\">Semantic Wikibase</a> erfolgreich auf Kompatibilit\u00e4t mit <a href=\"https://www.mediawiki.org/wiki/MediaWiki\" target=\"_blank\" rel=\"noopener\"><span class=\"hover:entity-accent entity-underline inline cursor-pointer align-baseline\"><span class=\"whitespace-normal\">MediaWiki</span></span></a> 1.43 aktualisiert wurde. Mit diesem Schritt stellen wir sicher, dass Semantic Wikibase weiterhin mit der aktuellen<a href=\"https://de.wikipedia.org/wiki/Support_(Dienstleistung)#Long_Term_Support\" target=\"_blank\" rel=\"noopener\"> Longterm-Support-Version</a> von MediaWiki kompatibel bleibt und als stabile Grundlage f\u00fcr semantisch angereicherte Wissensinfrastrukturen dient.</p>\n<h2 data-start=\"146\" data-end=\"301\">\u00dcber Semantic Wikibase</h2>\n<p data-start=\"303\" data-end=\"539\">Viele Forschungsprojekte setzen das Mediawiki-Framework als Werkzeug f\u00fcr Forschungsdatenmanagement ein. Mit \u00fcber 1.500 Erweiterungen l\u00e4sst sich dieses an die individuellen Anforderungen anpassen:</p>\n<ul>\n<li data-start=\"303\" data-end=\"539\">als reines Wiki mit Text und Medien, organisiert in Artikelseiten nach dem Vorbild von Wikipedia,</li>\n<li data-start=\"303\" data-end=\"539\">als strukturierte Wissens-Datenbank zur Linked-Open-Data Implementierung von Wissensgraphen und Terminologien mittels <a href=\"https://wikiba.se/\" target=\"_blank\" rel=\"noopener\">Wikibase,</a></li>\n<li data-start=\"303\" data-end=\"539\">als semantischer Wissensspeicher zur Datenvisualisierung mittels <a href=\"https://www.semantic-mediawiki.org/wiki/Semantic_MediaWiki\" target=\"_blank\" rel=\"noopener\">Semantic Mediawiki</a>.</li>\n</ul>\n<h3>Semantic Mediawiki vs. Wikibase</h3>\n<p>Insbesondere Wikibase und Semantic Mediawiki werden h\u00e4ufig im Forschungsumfeld verwendet. Beide Erweiterungen haben <strong>unterschiedliche St\u00e4rken und Schw\u00e4chen</strong>:</p>\n<figure id=\"attachment_31116\" aria-describedby=\"caption-attachment-31116\" style=\"width: 772px\" class=\"wp-caption alignnone\"><a href=\"https://de.slideshare.net/slideshow/semantic-mediawiki-a-linked-open-data-platform/272328819\" target=\"_blank\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-31116 size-full\" src=\"https://blog.tib.eu/wp-content/uploads/2026/03/Wikibase-vs-SMW.jpg\" alt=\"\" width=\"772\" height=\"322\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/03/Wikibase-vs-SMW.jpg 772w, https://blog.tib.eu/wp-content/uploads/2026/03/Wikibase-vs-SMW-300x125.jpg 300w, https://blog.tib.eu/wp-content/uploads/2026/03/Wikibase-vs-SMW-768x320.jpg 768w\" sizes=\"auto, (max-width: 772px) 100vw, 772px\" /></a><figcaption id=\"caption-attachment-31116\" class=\"wp-caption-text\">Vergleich von Wikibase and SMW (Grafik by Bernhard Krabina)</figcaption></figure>\n<h3>Semantic Mediawiki und Wikibase</h3>\n<p>Die Entwicklung von Semantic Wikibase (SWB) erm\u00f6glichte es erstmals, beide Erweiterungen gemeinsam auf einem System zu verbinden und so die <strong>Vorteile beider Systeme gemeinsam zu nutzen</strong>. W\u00e4hrend strukturelle Wissensdaten in Wikibase gespeichert und verwaltet werden, sorgt die SWB-Erweiterung daf\u00fcr, dass diese auch in Semantic Mediawiki f\u00fcr die Visualisierung in Wiki-Artikeln verf\u00fcgbar sind. SWB dient also quasi als Br\u00fccke zwischen den beiden Erweiterungen, wobei der Datenfluss nur von Wikibase nach Semantic Mediawiki (nicht umgekehrt) erfolgt. Dies dient dazu, Datenkonflikte zu vermeiden.</p>\n<p>Semantic Wikibase wurde im September 2020 in einer <a href=\"https://professional.wiki/en/articles/semantic-wikibase\" target=\"_blank\" rel=\"noopener\">ersten Version</a> vom Unternehmen <a href=\"https://professional.wiki/\" target=\"_blank\" rel=\"noopener\">ProfessionalWiki</a> ver\u00f6ffentlicht. Dieser erste Prototyp war nur mit der \u00e4lteren Mediawiki Version 1.35 kompatibel, aber unterst\u00fctzte bereits grundlegende Datentypen. Im Open Science Lab sahen wir in der Entwicklung einen Baustein, der das Potenzial hat, im Mediawiki-Umfeld eine bedeutende L\u00fccke zu schlie\u00dfen: Die <strong>Kombination aus strukturierter, f\u00f6derierbarer Datenverwaltung und Datenpr\u00e4sentation</strong>. Unser Ziel war es, die Erweiterung zu testen, bei Bedarf weiterzuentwickeln und k\u00fcnftig als unser <a href=\"https://de.wikipedia.org/wiki/Content-Management-System\" target=\"_blank\" rel=\"noopener\">Content-Management-System</a> zur Unterst\u00fctzung von Forschungsprojekten zu verwenden.</p>\n<h3>Case Studies</h3>\n<p>Mitte 2024 wurde mit dem Projekt <a href=\"https://phiwikibase4research-staging.adwmainz.net/index.php/PhiWiki:%C3%9Cber_PhiWiki\" target=\"_blank\" rel=\"noopener\">PhiWiki</a> ein erster Prototyp f\u00fcr Mediawiki 1.39 in Zusammenarbeit mit der <a href=\"https://www.adwmainz.de/\" target=\"_blank\" rel=\"noopener\">Akademie der Wissenschaften und der Literatur Mainz</a> sowie der <a href=\"https://digitale-philosophie.de/\" target=\"_blank\" rel=\"noopener\">AG Digitale Philosophie</a> erfolgreich getestet. Es folgte mit <a href=\"https://climatekg.semanticclimate.net/index.php?title=Hauptseite\" target=\"_blank\" rel=\"noopener\">Semantic Glossar</a> ein weiteres Projekt zur kollaborativen Entwicklung von Terminologien mittels Semantic Wikibase.</p>\n<p>Ende 2024 konnten wir im Rahmen des Projekts <a href=\"https://wb.manorhouses.tibwiki.io/wiki/Hauptseite\" target=\"_blank\" rel=\"noopener\">Herrenh\u00e4user des Ostseeraums</a> Semantic Wikibase dann in einem umfangreichen Projekt einem herausfordernden <a href=\"https://professional.wiki/en/news/connecting-wikibase-and-semantic-mediawiki\" target=\"_blank\" rel=\"noopener\">Lasttest</a> unterziehen. Mit \u00fcber 14.000 Wikibase-Objekten, die auf mehr als 300 Artikelseiten dynamisch eingebettet als Karten, Zeitstrahlen, Tabellen und Suchformulare verwendet werden, konnten wir die bestehenden Schw\u00e4chen von Semantic Wikibase identifizieren und beheben. Dazu geh\u00f6rte unter anderem die Unterst\u00fctzung des vollen Wikibase-Datenmodells mittels <a href=\"https://www.wikidata.org/wiki/Help:Qualifiers\" target=\"_blank\" rel=\"noopener\">Qualifiers</a>, eine erste grundlegende Unterst\u00fctzung des <a href=\"https://www.loc.gov/standards/datetime/\" target=\"_blank\" rel=\"noopener\">Extended Datetime Formats</a> (EDTF) sowie die Einbettung von 3D-Visualisierungen aus <a href=\"https://semantic-kompakkt.de/home?locale=en\" target=\"_blank\" rel=\"noopener\">Semantic Kompakkt</a>. Entscheidend war hierf\u00fcr die<strong>\u00a0interdisziplin\u00e4re Zusammenarbeit</strong> zwischen dem Enwicklerteam und den LOD- und Wikibase-Datenmodell-Expertinnen Lozana Rossenova und Lucia Sohmen.</p>\n<p>Die im Projekt entwickelten Best-Practices umfassten unter anderem:</p>\n<ul>\n<li>Nutzung individueller Formulare f\u00fcr Suchfilter und Dateneingabe</li>\n<li>Verkn\u00fcpfung von Wikibase-Items mit Mediawiki-Kategorien</li>\n<li>Verlinkung von Artikelseiten mit Wikibase-Items</li>\n<li>Nutzung des vollen Wikibase-Datenmodells in <a href=\"https://www.semantic-mediawiki.org/wiki/Help:Inline_queries\" target=\"_blank\" rel=\"noopener\">SMW Inline Queries</a></li>\n<li>Performance von <a href=\"https://www.semantic-mediawiki.org/wiki/Help:Result_formats\" target=\"_blank\" rel=\"noopener\">Datenvisualisierungen</a> trotz hoher Anzahl an Wikibase-Objekten</li>\n<li>Best Practices zur Informationsmodellierung im <a href=\"https://gitlab.com/nfdi4culture/wikibase4research/auxiliary-service-repositories/wikibase-model\" target=\"_blank\" rel=\"noopener\">Generic Wikibase Model for Cultural Data</a></li>\n</ul>\n<h2 data-start=\"541\" data-end=\"576\">Warum MediaWiki 1.43 wichtig ist</h2>\n<p data-start=\"578\" data-end=\"879\">Mit der Version 1.39 war Semantic Wikibase kompatibel mit der damaligen <a href=\"https://de.wikipedia.org/wiki/Support_(Dienstleistung)#Long_Term_Support\" target=\"_blank\" rel=\"noopener\">Longtime Support Version(LTS)</a> von Mediawiki. Diese Unterst\u00fctzung war aber gem\u00e4\u00df des <a href=\"https://www.mediawiki.org/wiki/Version_lifecycle/de\" target=\"_blank\" rel=\"noopener\">Mediawiki Lifecycle</a> nur bis Ende 2025 gegeben.</p>\n<p data-start=\"578\" data-end=\"879\">MediaWiki 1.43 bringt als aktuelle LTS-Version (Support bis 2028) zahlreiche technische Verbesserungen, Performance-Optimierungen sowie langfristige Wartungsvorteile mit sich. F\u00fcr viele Wikibase-Installationen ist die Orientierung an den aktuellen MediaWiki-Versionen essenziell, um Sicherheit, Stabilit\u00e4t und Zukunftsf\u00e4higkeit zu gew\u00e4hrleisten. Durch Versionskonflikte zwischen verwendeten Bibliotheken in Wikibase und Semantic Mediawiki, konnte SemanticWikibase aber nicht ohne Anpassung in dieser neuen Version eingesetzt werden.</p>\n<p data-start=\"578\" data-end=\"879\"><strong>Unsere gr\u00f6\u00dfte Bef\u00fcrchtung</strong> war, dass die aktuellen Versionen grundlegende \u00c4nderung vorgenommen hatten, die einen Weiterbetrieb von Semantic Wikibase technisch unsauber bzw. unwirtschaftlich machen w\u00fcrden. Ende 2025 schaffte Open-Science-Lab-Entwickler Lukas G\u00fcnther die entscheidende Grundlage f\u00fcr das Upgrade, indem er unser Installationstool <a href=\"https://gitlab.com/nfdi4culture/wikibase4research/wikibase4research\" target=\"_blank\" rel=\"noopener\">Wikibase4Research</a> aktualisierte und so mit der Mediawiki Version 1.44 kompatibel machte. Da Semantic Wikibase sich mittels Wikibase4Research automatisiert installieren l\u00e4sst, war so ein geeignetes Test-Setup geschaffen, um die Entwicklung in Angriff zu nehmen. Letzendlich war es uns so m\u00f6glich, Semantic Wikibase mit der aktuellen LTS-Version von Mediawiki zu betreiben und das sogar ohne \u00c4nderungen am Wikibase- oder SemanticMediawiki-Code vorzunehmen. S\u00e4mtliche bisher unterst\u00fctzten Datentypen sind auch weiterhin funktional, was auch ein Update bestehender Installationen auf die neue Version erm\u00f6glicht.</p>\n<figure id=\"attachment_31121\" aria-describedby=\"caption-attachment-31121\" style=\"width: 520px\" class=\"wp-caption alignnone\"><img loading=\"lazy\" decoding=\"async\" class=\"size-full wp-image-31121\" src=\"https://blog.tib.eu/wp-content/uploads/2026/03/SWB-Datatypes.png\" alt=\"\" width=\"520\" height=\"442\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/03/SWB-Datatypes.png 520w, https://blog.tib.eu/wp-content/uploads/2026/03/SWB-Datatypes-300x255.png 300w\" sizes=\"auto, (max-width: 520px) 100vw, 520px\" /><figcaption id=\"caption-attachment-31121\" class=\"wp-caption-text\">Unterst\u00fctzte Datentypen in Semantic Wikibase, visualisiert im Semantic Browser von SMW</figcaption></figure>\n<h2 data-start=\"2223\" data-end=\"2234\">Ausblick</h2>\n<p data-start=\"2236\" data-end=\"2552\">Die kontinuierliche Synchronisierung von Semantic Wikibase mit dem MediaWiki-Releasezyklus ist ein zentraler Baustein f\u00fcr nachhaltige, semantische Wissensinfrastrukturen. Mit diesem Update schaffen wir die Grundlage f\u00fcr kommende Weiterentwicklungen und eine langfristig stabile Integration in das Wikibase-\u00d6kosystem. Der Einsatz von Semantic Wikibase bedeutet f\u00fcr unsere Forschungsdaten- und Terminologie-Projekte im Open Science Lab:</p>\n<ul>\n<li data-start=\"2236\" data-end=\"2552\">Fokussierung auf eine gemeinsame technologische Basis f\u00fcr alle Projekte</li>\n<li data-start=\"2236\" data-end=\"2552\">B\u00fcndelung von Wissen und Ressourcen</li>\n<li data-start=\"2236\" data-end=\"2552\">Zeitersparnis bei der Projektumsetzung durch Best Practices und Synergieeffekten zwischen Projekten</li>\n<li data-start=\"2236\" data-end=\"2552\">Koordinierter Aufbau von Services innerhalb eines bestehenden Software \u00d6kosystems</li>\n<li data-start=\"2236\" data-end=\"2552\">Support der Open-Source und Linked-Open-Data Community durch unsere Entwicklungen</li>\n</ul>\n<p><strong>Wir freuen uns auf die weitere Entwicklung und die vielf\u00e4ltigen kommenden Projekte mit Semantic Wikibase.</strong></p>\n<h5>Relevante Links</h5>\n<ul>\n<li><a href=\"https://gitlab.com/nfdi4culture/wikibase4research/wikibase4research\" target=\"_blank\" rel=\"noopener\">Wikibase4Research</a></li>\n<li><a href=\"https://gitlab.com/nfdi4culture/wikibase4research/auxiliary-service-repositories/wikibase-model\" target=\"_blank\" rel=\"noopener\">Generic Wikibase Model for Cultural Data</a></li>\n</ul>","doi":"https://doi.org/10.65527/2cdwj-mqe46","guid":"https://blog.tib.eu/?p=31114","image":"https://blog.tib.eu/wp-content/uploads/2026/04/Wikibase_logo.svg.png","language":"de","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1775692800,"rid":"kfdwg-qff93","summary":"Wir freuen uns, bekanntzugeben, dass Semantic Wikibase erfolgreich auf Kompatibilit\u00e4t mit MediaWiki 1.43 aktualisiert wurde. Mit diesem Schritt stellen wir sicher, dass Semantic Wikibase weiterhin mit der aktuellen Longterm-Support-Version von MediaWiki kompatibel bleibt und als stabile Grundlage f\u00fcr semantisch angereicherte Wissensinfrastrukturen dient.","tags":["OPENNESS","FORSCHUNG & PROJEKTE","Lizenz:CC-BY-4.0-INT","Open Science Lab","Wikibase"],"title":"Upgrade abgeschlossen: Semantic Wikibase kompatibel mit MediaWiki 1.43","updated_at":1790067918,"url":"https://blog.tib.eu/2026/04/09/upgrade-abgeschlossen-semantic-wikibase-kompatibel-mit-mediawiki-1-43/","version":"v1"},{"authors":[{"contributor_roles":[],"family":"Bailly","given":"Kolja"}],"blog":{"authors":null,"community_id":"db0d8909-9e37-46d0-b16c-0551f575e86b","created":1749772800,"current_feed_url":null,"description":"Das Blog der TIB \u2013 Leibniz-Informationszentrum Technik und Naturwissenschaften und Universit\u00e4tsbibliothek","doi":"https://doi.org/10.65527/tib","favicon":"https://rogue-scholar.org/api/communities/db0d8909-9e37-46d0-b16c-0551f575e86b/logo","feed_format":"application/atom+xml","feed_url":"https://blog.tib.eu/feed/atom/","filter":null,"generator":"WordPress","home_page_url":"https://blog.tib.eu/","issn":null,"language":"deu","license":"https://creativecommons.org/licenses/by/4.0/legalcode","prefix":"10.65527","relative_url":null,"secure":true,"slug":"tib","status":"active","subfield":"1802","title":"TIB-Blog","updated":1790082029,"use_api":true},"blog_name":"TIB-Blog","blog_slug":"tib","content_html":"<p>KI-Systeme, die Texte nicht nur generieren, sondern gezielt in Dokumenten recherchieren, sind mittlerweile etablierter Stand der Technik. Einer dieser Ans\u00e4tze hei\u00dft <a href=\"https://en.wikipedia.org/wiki/Retrieval-augmented_generation\" target=\"_blank\" rel=\"noopener\">Retrieval-Augmented Generation (RAG)</a>: Stellt ein Benutzer eine Frage, sucht das System relevante Informationen in einer Wissensbasis \u2013 zum Beispiel in einem Wiki \u2013 und nutzt diese als Grundlage, um relevante Inhalte bzw. Quellen aufzulisten oder mittels KI Antworten daraus zu generieren.</p>\n<p><strong>Das Problem</strong>: Damit ein solches System gut funktioniert, m\u00fcssen viele Stellschrauben richtig eingestellt werden. Diese sogenannte <a href=\"https://en.wikipedia.org/wiki/Hyperparameter_optimization\" target=\"_blank\" rel=\"noopener\">Hyperparameter-Optimierung</a> ist normalerweise entweder zeitaufw\u00e4ndig oder rechenintensiv und in jedem Fall technisch anspruchsvoll. Unsere aktuelle Untersuchung zeigt jedoch: <strong>Eine automatisierte Optimierung ist m\u00f6glich \u2013 sogar auf einem normalen Laptop</strong>.</p>\n<h2>Ausgangslage</h2>\n<p>Grundlage unserer Untersuchung im <a href=\"https://www.tib.eu/de/forschung-entwicklung/forschungsgruppen-und-labs/open-science\" target=\"_blank\" rel=\"noopener\"><strong>Open Science Lab</strong></a> war die Weiterentwicklung unseres RAG-Moduls f\u00fcr <a href=\"https://blog.tib.eu/2024/08/29/wikibase4research-wissensdaten-einfach-verwalten-teilen-und-visualisieren/\" target=\"_blank\" rel=\"noopener\"><strong>Wikibase4Research</strong></a>. Mit dem zuvor bestehenden System war es bereits sehr einfach m\u00f6glich, eine Mediawiki Installation zu erhalten, deren Inhalte KI-gest\u00fctzt via RAG durchsuchbar sind. Egal ob es nun um Artikelseiten in einem einfachen <a href=\"https://www.mediawiki.org/wiki/MediaWiki\" target=\"_blank\" rel=\"noopener\">Mediawiki</a>, strukturierte Wissensdaten in einer <a href=\"https://wikiba.se/\" target=\"_blank\" rel=\"noopener\">Wikibase</a> oder eine Kombination aus beidem wie zum Beispiel <a href=\"https://www.semantic-mediawiki.org/wiki/Semantic_MediaWiki\" target=\"_blank\" rel=\"noopener\">Semantic Mediawiki</a> oder <a href=\"https://www.mediawiki.org/wiki/Extension:Semantic_Wikibase\" target=\"_blank\" rel=\"noopener\">Semantic Wikibase</a> geht.</p>\n<p>Eine Einf\u00fchrung in die grundlegende Funktionsweise von RAG und Wikibase4Research liefert das folgende Video:</p>\n<div class=\"ratio ratio-16x9\"><iframe loading=\"lazy\" title=\"Kolja Bailly: From Triples to Text: LLM(RAG)-Based Approach to Querying Wikibase | #MUDCon Fall 2025\" width=\"800\" height=\"450\" src=\"https://www.youtube.com/embed/GIZA4OVogLc?feature=oembed\" frameborder=\"0\" allow=\"accelerometer; autoplay; clipboard-write; encrypted-media; gyroscope; picture-in-picture; web-share\" referrerpolicy=\"strict-origin-when-cross-origin\" allowfullscreen></iframe></div>\n<p>Um eine hohe Qualit\u00e4t der KI-basierten Suchergebnisse und Antworten zu erhalten, ist es aber n\u00f6tig, das System <strong>entsprechend der verwendeten Daten zu konfigurieren</strong>. F\u00fcr diese Einstellungen gibt es keine Standardf\u00e4lle, es geh\u00f6rt in das Arbeitsfeld eines Data Scientist die Systemparameter zu testen und zu verbessern. In diesem Prozess wird daher klassisch ein hohes Ma\u00df an Erfahrung und Fachwissen ben\u00f6tigt, um optimale Ergebnisse zu erhalten.</p>\n<p>Die Alternative ist der nun in Wikibase4Research integrierte <strong>AutoRAG</strong> Ansatz, der die Parameter vollautomatisch optimiert. Dieser Prozess wird im Farchjargon \"Hyperparameter Tuning\" oder auch \"Hyperparameter Optimierung\" genannt.</p>\n<h2>Anforderungen</h2>\n<p>Die Rahmenbedingungen f\u00fcr ein Hyperparameter Tuning k\u00f6nnen sehr unterschiedlich sein. In unserem Fall ergeben sich die Anforderungen vor allem aus der Nutzergruppe von Wikibase4Research.</p>\n<h3>Forscher/Innen</h3>\n<p>Im Forschungskontext haben wir es mit f\u00e4cherspezifischen Daten zu tun. Die beteiligten Wissenschaftler sind Experten in ihrer jeweiligen Fachdom\u00e4ne. Expertise im Bereich spezieller Data-Science-Anwendungen ist in den Projektteams meist nicht vorhanden. Dies ist durchaus sinnvoll, denn das Projektteam ist somit auf die im Projekt zu bearbeitenden Forschungsfragen spezialisiert.</p>\n<h3>Daten</h3>\n<p>F\u00fcr die Optimierung wird ein Test-Datensatz ben\u00f6tigt, der m\u00f6gliche Fragen (Suchanfragen) mit den optimalen Quellen in den Daten verkn\u00fcpft. Dieser Datensatz wird mit den Suchergebnissen des Systems verglichen, um die Qualit\u00e4t der Systemeinstellung bewerten zu k\u00f6nnen (Idealdaten). Solche Testdaten liegen in den \u00fcberwiegenden F\u00e4llen nicht vor.</p>\n<h3>Endnutzer/Innen</h3>\n<p>Wer nutzt die Daten letztendlich und welche Art von Anfragen werden gestellt? Diese Frage ist entscheidend bei der Optimierung. Werden die Endnutzer spezifische Fakten aus den Daten abfragen wie zum Beispiel Jahreszahlen bestimmter Ereignisse oder eher Zusammenfassungen ganzer Abs\u00e4tze oder Artikel erwarten? Zu welchen Themen werden voraussichtlich Fragen gestellt? Erwarte ich eher Fragen zum Inhalt der Daten oder Fragen auf der Metaebene wie zum Beispiel zur Anzahl von Quellen, der Struktur und L\u00e4nge von Texten, des Schreibstils oder zur Medienart? Werden Suchanfragen von Wissenschaftlern im Fachjargon gestellt oder eher in Umgangssprache formuliert? <span style=\"font-family: inherit;\">Die fr\u00fchzeitige Definition grundlegender </span><a style=\"font-family: inherit;\" href=\"https://en.wikipedia.org/wiki/Persona_(user_experience)\" target=\"_blank\" rel=\"noopener\">Personas</a><span style=\"font-family: inherit;\"> f\u00fcr die zu erwartende Nutzergruppe hilft nicht nur bei der Optimierung von RAG, sondern ist auch ein wichtiger Schritt bei der Erstellung von Design und Benutzeroberfl\u00e4chen in der Pr\u00e4sentation der Forschungsergebnisse.</span></p>\n<h3>Infrastruktur</h3>\n<p>Hohe Rechenkapazit\u00e4ten, Zugang zu GPU-Processing und Budget f\u00fcr industrielle KI-Services ist in vielen Projekten nicht vorhanden. Wikibase4Research bietet die Option, externe Schnittstellen wie <a href=\"https://huggingface.co/models\" target=\"_blank\" rel=\"noopener\">Huggingface</a>, <a href=\"https://developers.openai.com/api/reference\" target=\"_blank\" rel=\"noopener\">OpenAI</a> oder die <a href=\"https://docs.hpc.gwdg.de/services/saia/index.html\" target=\"_blank\" rel=\"noopener\">SAIA-Umgebung</a> der <a href=\"https://gwdg.de/\" target=\"_blank\" rel=\"noopener\">GWDG</a> zur Ausf\u00fchrung von KI-Modellen zu nutzen. Die dort bestehenden Limits f\u00fcr kostenlose Nutzung reichen aber meist nicht aus, um die Vielzahl an Parameter-Konfigurationen zu testen, die zur Optimierung eines RAG-Systems notwendig ist. Ideal w\u00e4re also, die Ausf\u00fchrung lokal auf allgemein verf\u00fcgbarer Hardware durchf\u00fchren zu k\u00f6nnen, was auch unter dem Aspekt der ressourcenschonenden Nutzung von KI ein erstrebenswertes Ziel ist.</p>\n<p><strong>Es ergibt sich f\u00fcr unseren Ansatz daher folgender Anforderungskatalog:</strong></p>\n<ul>\n<li>Anpassung auf die verwendeten Daten</li>\n<li>vollautomatische Optimierung</li>\n<li>keine technischen Vorkenntnisse n\u00f6tig</li>\n<li>Test-Datensatz wird generiert</li>\n<li>User-Persona-Profile ber\u00fccksichtigen</li>\n<li>m\u00f6glichst effizient, mit geringem Ressourcenbedarf</li>\n</ul>\n<h2>Methodik</h2>\n<h3>Daten</h3>\n<p>Als Datengrundlage dienten jeweils 50 zuf\u00e4llige Artikel aus drei MediaWiki-basierten Wissenssammlungen:</p>\n<ul>\n<li><a href=\"https://de.wikipedia.org/\" target=\"_blank\" rel=\"noopener\">Wikipedia</a></li>\n<li><a href=\"https://wb.manorhouses.tibwiki.io/wiki/Deutschland\" target=\"_blank\" rel=\"noopener\">Herrenh\u00e4user</a></li>\n<li><a href=\"https://kungfu-wiki.com/Scholar_indexpage\" target=\"_blank\" rel=\"noopener\">Kungfu-Wiki</a></li>\n</ul>\n<p>Um die Qualit\u00e4t der Suche zu bewerten, wurden automatisch Frage-Kontext-Antwort-Tripel erzeugt. Zum Einsatz kam daf\u00fcr das mehrsprachige Sprachmodell <a href=\"https://www.ibm.com/de-de/new/announcements/ibm-granite-4-0-hyper-efficient-high-performance-hybrid-models\" target=\"_blank\" rel=\"noopener\"><strong>IBM Granite 4 350M Nano</strong></a>, das speziell f\u00fcr Umgebungen mit geringer Rechenleistung wie zum Beispiel f\u00fcr On-Device-Anwendungsf\u00e4lle entwickelt wurde.</p>\n<h3>LLM-Prompt</h3>\n<p>Um hinsichtlich der erwarteten Nutzung realistische Fragen zu generieren, wurde der an das Modell gelieferte Prompt (\"Erstelle Fragen aus dem Seiteninhalt\") um speziell angepasste Rollenbeschreibungen (Personas) erg\u00e4nzt, die per Konfigurationsdatei individualisiert werden k\u00f6nnen. Eine solche Persona-Definition k\u00f6nnte zum Beispiel lauten: \"You are a scientist who wants to learn about historic manorhouses in Europe\".</p>\n<h3>Parameter</h3>\n<p>In einem RAG-Prozess werden die zu durchsuchenden Daten in einer speziellen Datenbank indiziert, um sp\u00e4ter schnell und effizient relevante Inhalte zu finden.</p>\n<figure id=\"attachment_30955\" aria-describedby=\"caption-attachment-30955\" style=\"width: 847px\" class=\"wp-caption alignnone\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-30955 \" src=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag1-1024x240.jpg\" alt=\"\" width=\"847\" height=\"199\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag1-1024x240.jpg 1024w, https://blog.tib.eu/wp-content/uploads/2026/02/rag1-300x70.jpg 300w, https://blog.tib.eu/wp-content/uploads/2026/02/rag1-768x180.jpg 768w, https://blog.tib.eu/wp-content/uploads/2026/02/rag1.jpg 1440w\" sizes=\"auto, (max-width: 847px) 100vw, 847px\" /><figcaption id=\"caption-attachment-30955\" class=\"wp-caption-text\">Information Extraction und Indizierung von Daten in einem RAG-Prozess</figcaption></figure>\n<p>Die meisten von uns verwendeten Parameter optimieren diesen Prozess der\u00a0<a href=\"https://en.wikipedia.org/wiki/Information_extraction\" target=\"_blank\" rel=\"noopener\">Informations Extraktion</a> (IE). Dabei wird bestimmt, in welcher Form die Daten gespeichert werden und ob diese ggf. vor dem Speichern um Metadaten wie Schlagworte, Titel oder Zusammenfassungen erg\u00e4nzt werden. F\u00fcr die Vektorisierung verwendeten wir das Modell <a href=\"https://github.com/QwenLM/Qwen3-Embedding\" target=\"_blank\" rel=\"noopener\"><strong>Qwen3-embedding:0.6B</strong></a>. Die mittels AutoRAG optimierten Parameter sind im Folgenden aufgelistet:</p>\n<ul>\n<li><strong>Chunk_Size</strong>: Wie gro\u00df sind die Informationsabschnitte, die sp\u00e4ter zugreifbar sein sollen?</li>\n<li><strong>Chunk_Overlap</strong>: Wie stark \u00fcberlappen sich die Informationsabschnitte?</li>\n<li><strong>Extractors</strong>: Welche Datenanreicherungen sollen erfolgen (zum Beispiel Zusammenfassung erstellen, Fragen generieren)?</li>\n<li><strong>Top_K:\u00a0</strong>Wieviele Chunks werden als Suchergebnis geliefert?</li>\n</ul>\n<p>Sind die Daten eingelesen und wird eine Suchanfrage gestellt, wird das System nach relevanten Informationsabschnitten durchsucht. Dieser Prozess wird \"Information Retrieval\" genannt. Man kann es mit den Ergebnissen einer Google-Suche vergleichen, bei der die relevantesten Ergebnisse nicht zwangsl\u00e4ufig an erster Stelle der Liste stehen.</p>\n<figure id=\"attachment_30957\" aria-describedby=\"caption-attachment-30957\" style=\"width: 800px\" class=\"wp-caption alignnone\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-30957 size-large\" style=\"font-family: inherit;\" src=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag2-1024x273.jpg\" alt=\"\" width=\"800\" height=\"213\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag2-1024x273.jpg 1024w, https://blog.tib.eu/wp-content/uploads/2026/02/rag2-300x80.jpg 300w, https://blog.tib.eu/wp-content/uploads/2026/02/rag2-768x204.jpg 768w, https://blog.tib.eu/wp-content/uploads/2026/02/rag2.jpg 1360w\" sizes=\"auto, (max-width: 800px) 100vw, 800px\" /><figcaption id=\"caption-attachment-30957\" class=\"wp-caption-text\">Information Retrieval in einem RAG Prozess</figcaption></figure>\n<p>Information Retrieval bedeutet, zur Frage des Nutzers relevante Informationen zu finden. In diesem Prozessschritt optimieren wir den Parameter \"<strong>Top_K\"</strong>, der definiert, wie viele der Suchergebnisse im weiteren Prozess ber\u00fccksichtigt werden. Ist Top_K zu klein, sind wichtige Quellen eventuell nicht enthalten. Ist Top_K zu gro\u00df, verarbeitet man eventuell eine gro\u00dfe Menge wenig relevanter Inhalte.</p>\n<h3>Optimierungsverfahren</h3>\n<p>Statt alle m\u00f6glichen Kombinationen auszuprobieren (was sehr lange dauern w\u00fcrde), kommt ein Suchalgorithmus zum Einsatz, der die verschiedenen Parameter stufenweise verbessert. Dieses als <a href=\"https://en.wikipedia.org/wiki/Greedy_algorithm\" target=\"_blank\" rel=\"noopener\">Greedy</a> (\"gierig<strong>\"</strong>) benannte Verfahren optimiert zun\u00e4chst nur einen einzigen Parameter, dann den n\u00e4chsten usw. Wir verzichten damit auf optimale L\u00f6sungen, erreichen aber hinreichend <strong>gute Ergebnisse mit akzeptablem Aufwand</strong>.</p>\n<p>Als Bewertungsma\u00df f\u00fcr die Optimierung dient dabei der sogenannte <a href=\"https://en.wikipedia.org/wiki/Mean_reciprocal_rank\" target=\"_blank\" rel=\"noopener\">Mean Reciprocal Rank (MRR)</a> \u2013 ein Ma\u00df daf\u00fcr, an welcher Position relevante Inhalte in der Trefferliste platziert sind. Ein entscheidender Vorteil:<br />\nDie Bewertung erfolgt vollst\u00e4ndig ohne KI-Antwortgenerierung. Es wird also nur getestet, wie gut das System relevante Inhalte findet, nicht wie gut eine KI daraus sp\u00e4ter Antworten generiert. Dadurch wird erheblich Rechenzeit gespart.</p>\n<figure id=\"attachment_30956\" aria-describedby=\"caption-attachment-30956\" style=\"width: 710px\" class=\"wp-caption alignnone\"><img loading=\"lazy\" decoding=\"async\" class=\"wp-image-30956\" src=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag3.jpg\" alt=\"\" width=\"710\" height=\"256\" srcset=\"https://blog.tib.eu/wp-content/uploads/2026/02/rag3.jpg 919w, https://blog.tib.eu/wp-content/uploads/2026/02/rag3-300x108.jpg 300w, https://blog.tib.eu/wp-content/uploads/2026/02/rag3-768x277.jpg 768w\" sizes=\"auto, (max-width: 710px) 100vw, 710px\" /><figcaption id=\"caption-attachment-30956\" class=\"wp-caption-text\">Antwort Generierung in einem RAG Prozess. Diese Phase wurde in der Optimierung NICHT ber\u00fccksichtigt</figcaption></figure>\n<h2>Technische Umsetzung</h2>\n<p>Die Implementierung erfolgte vollst\u00e4ndig im MediaWiki-Umfeld mit:</p>\n<ul>\n<li>Wikibase4Research</li>\n<li>einer Docker-basierten Python-API</li>\n<li>dem RAG-Framework LlamaIndex</li>\n<li>lokaler Modellbereitstellung \u00fcber Ollama</li>\n</ul>\n<p>Die Experimente liefen auf einem <strong>handels\u00fcblichen Laptop</strong> aus dem Jahr 2022 (Dell Latitude 5421, Intel Core i7-11850H mit 8 Kernen, 16 GB RAM) \u2013 ohne GPU-Beschleunigung.</p>\n<h2>Ergebnisse</h2>\n<p>Trotz der bewusst schlanken Hardware-Ausstattung konnte die Optimierung meist bereits innerhalb einer Stunde abgeschlossen werden. Dabei wurde bei allen Datens\u00e4tzen eine starke Verbesserungen der Abfrageergebnisse erzielt.</p>\n<p>F\u00fcr unser Qualit\u00e4stma\u00df, den Mean Reciprocal Rank (MRR), ergab sich eine <strong>Steigerung von durchschnittlich 12 bis 25 Prozent</strong> gegen\u00fcber den voreingestellten Parametern. Das bedeutet, in den Ergebnissen der Suchanfrage waren mehr relevante Quellen aufgef\u00fchrt und relevante Quellen standen in der Ergebnisliste an h\u00f6herer Stelle als zuvor. In einzelnen Datens\u00e4tzen ergaben sich sogar Verbesserungen von bis zu 50 Prozent. Dabei lie\u00dfen sich vergleichbare Ergebnisse auch mit Artikeln erreichen, die nicht Teil der Optimierungsschleife waren (<a href=\"https://en.wikipedia.org/wiki/Cross-validation_(statistics)\" target=\"_blank\" rel=\"noopener\">Cross-Validation</a>).</p>\n<h2>Warum ist das relevant?</h2>\n<p>F\u00fcr wissenschaftliche Infrastrukturen wie digitale Bibliotheken, Fachrepositorien oder Forschungsdatenplattformen ist es entscheidend, KI-Systeme effizient und ressourcenschonend betreiben zu k\u00f6nnen. Die Ergebnisse zeigen: <strong>Sinnvolle RAG-Optimierung ist auch ohne Rechenzentrum machbar</strong>.</p>\n<p>Das senkt technische H\u00fcrden, reduziert Kosten und macht den Einsatz moderner KI-Technologien auch in kleineren Projekten realistisch.</p>\n<h2>Ausblick</h2>\n<p>Die f\u00fcr die Suche verwendeten <a href=\"https://en.wikipedia.org/wiki/Embedding_(machine_learning)\" target=\"_blank\" rel=\"noopener\">Embedding-Vector-Modelle</a> haben einen erheblichen Einfluss auf die Ergebnisse (vgl. <a href=\"https://arxiv.org/abs/2505.03452\" target=\"_blank\" rel=\"noopener\">Orbach et al. (2025)</a>) und zwar sowohl auf die Rechenzeit als auch auf die Ergebnisqualit\u00e4t. Dabei zeigen Modelle nicht auf allen Datens\u00e4tzen die gleichen Ergebnisse.</p>\n<p>Es ist auch nur begrenzt m\u00f6glich, die Optimierung mit extrem kleinen oder schnellen Embedding-Modellen auszuf\u00fchren und die optimierten Parameter dann zusammen mit einem anderen, leistungsf\u00e4higen Modell im Live-Betrieb einzusetzen. Sind die eingesetzten Embedding-Modelle nicht angepasst genug an die verwendete Wissensdom\u00e4ne, liefert auch die Optimierung nur suboptimale Ergebnisse.</p>\n<p>Genau an diesem Punkt wird unsere Arbeit im Open Science Lab in der n\u00e4chsten Zeit ansetzen. Gemeinsam mit den Fachinformationsdiensten FID Material Science, FID Move, FID Pyhsik und FID Philosophie evaluieren wir die M\u00f6glichkeit einer <strong>st\u00e4rkeren Vernetzung von NFDI und FIDs</strong> mit dem Ziel, die einzelnen Wissendom\u00e4nen mit fachspezifischen Embedding-Modellen zu versorgen. Zielsetzung ist es, damit den Zugang zu dieser Technologie noch weiter zu vereinfachen sowie die Qualit\u00e4t der Ergebnisse von KI-Anwendungen im Forschungs- und Bibliotheksumfeld gezielt zu erh\u00f6hen.</p>\n<blockquote>\n<figure style=\"width: 150px\" class=\"wp-caption alignright\"><a href=\"https://www.hs-hannover.de/service/personenfinder/person/1000005796\" target=\"_blank\" rel=\"noopener\"><img loading=\"lazy\" decoding=\"async\" src=\"https://www.tib.eu/fileadmin/_processed_/b/1/csm_Bluemel_Ina_57de4441fa.jpg\" alt=\"Prof. Dr. Ina Bl\u00fcmel\" width=\"150\" height=\"150\" /></a><figcaption class=\"wp-caption-text\">Prof. Dr. Ina Bl\u00fcmel, Open Science Lab // Foto: TIB/C. Bierwagen</figcaption></figure>\n<p data-pm-slice=\"1 1 []\">\"AutoRAG ist f\u00fcr uns ein wichtiger Innovationsschritt: Es macht RAG in offenen Wissensr\u00e4umen wie Wikibase messbar, wiederholbar und mit \u00fcberschaubaren Ressourcen betreibbar. F\u00fcr Projekte wie NFDI4Culture und weitere Vorhaben im Open Science Lab bedeutet das sp\u00fcrbar bessere, nachvollziehbare KI-gest\u00fctzte Suche \u00fcber heterogene Best\u00e4nde \u2013 ohne dass tiefes Spezial-Know-how aufgebaut werden muss. N\u00e4chster Schritt ist der Ausbau fachspezifischer Embeddings, kuratierter Testsets und transparenter Workflows, damit die Qualit\u00e4t und Nachnutzbarkeit langfristig steigt.\"</p>\n</blockquote>\n<h3 style=\"padding-left: 0 important!;\">Relevante Links</h3>\n<ul>\n<li><strong>Wikibase4Research</strong>: <a href=\"https://gitlab.com/nfdi4culture/wikibase4research/wikibase4research\" target=\"_blank\" rel=\"noopener\">https://gitlab.com/nfdi4culture/wikibase4research/wikibase4research</a></li>\n<li><strong>Wikibase4Research-RAG Modul:</strong> <a href=\"https://gitlab.com/nfdi4culture/wikibase4research/wikibase-RAG\" target=\"_blank\" rel=\"noopener\">https://gitlab.com/nfdi4culture/wikibase4research/wikibase-RAG</a></li>\n</ul>","doi":"https://doi.org/10.65527/gvnwm-x6559","guid":"https://blog.tib.eu/?p=30916","image":"https://blog.tib.eu/wp-content/uploads/2026/03/AutoRAG-thumb.jpg","language":"de","license":"https://creativecommons.org/licenses/by/4.0/legalcode","published_at":1775001600,"rid":"b11wn-f5028","summary":"KI-Systeme, die Texte nicht nur generieren, sondern gezielt in Dokumenten recherchieren, sind mittlerweile etablierter Stand der Technik. Damit ein solches System gut funktioniert, m\u00fcssen viele Stellschrauben richtig eingestellt werden, was in der Regel zeitaufw\u00e4ndig, rechenintensiv und technisch anspruchsvoll ist. Unsere aktuelle Untersuchung zeigt jedoch: Eine automatisierte Optimierung ist m\u00f6glich \u2013 sogar auf einem normalen Laptop.","tags":["OPENNESS","FORSCHUNG & PROJEKTE","Lizenz:CC-BY-4.0-INT","Wikibase","Projekte"],"title":"Bessere KI-Antworten \u2013 auch ohne Hochleistungsrechner","updated_at":1790067915,"url":"https://blog.tib.eu/2026/04/01/bessere-ki-antworten-auch-ohne-hochleistungsrechner/","version":"v1"}],"out_of":57638,"page":1,"per_page":10,"total-results":57638}
