Skip to main contentSkip to search
Episciences
Open Access Journals
Sign in(new window)
Transformations: A DARIAH Journal  logo
Transformations: A DARIAH Journal
Transformations: A DARIAH Journal  logo
Transformations: A DARIAH Journal
Sign in(new window)
Articles & Issues
All articlesAll accepted articlesAll volumesLast volumeSectionsAuthors
About
The journalNews
Boards
Publish
For authorsEthical charterProposing special issues
Submit
Transformations: A DARIAH Journal  logo
Journal's leaflet
|
Contact
|
Credits
eISSN 3095-3200
|
RSS
|
Atom
Episciences
Documentation
|
Acknowledgements
|
Publishing policy
Accessibility: non-compliant
|
Legal mentions
|
Privacy statement
|
Terms of use
  1. Home > Articles & Issues >
  2. Sections >
  3. Metadata-based workflows

Metadata-based workflows

section
5 articles
5 articles
Article
Genre Classification Workflow for the English Short Title Catalogue (ESTC)
Iiro Tiihonen, Kira Hinderks
Abstract
An article about genre classification of early modern British books published in Transformations: A DARIAH Journal, 1(1), 2025. Abstract below:This article introduces an open-box workflow for labelling 94 percent of the English Short Title Catalogue (ESTC) with a unified genre classification scheme, as well as an approach to evaluating the classifications. As the ESTC covers most of the surviving books published in the early modern Anglosphere, our categorisation offers new opportunities for large-scale quantitative research on the early modern book trade. Our evaluation process directly engages with the ambiguity of any genre labelling or annotation schemes of early modern books and highlights problematic boundaries between the categories. We also provide summary statistics about the genre composition of the ESTC, demonstrate how the new data can be used to detect biases in other datasets of early modern books and discuss further possibilities for genre-related computational work with the ESTC.
Published on June 11, 2025
PDF
Article
A reproducible framework to publish and reuse Collections as data: the case of the European Literary Bibliography
Gustavo Candela, Cezary Rosiński, Arkadiusz Margraf
Abstract
GLAM (Galleries, Libraries, Archives and Museums) institutions host rich content that is provided in the form of digital collections. Bibliographic databases are collections of references focused on a particular topic that can be used to apply Digital Humanities (DH) methods. Recent approaches such as Collections as Data and Labs promote the publication of digital collections supporting computational use.This work aims to provide a framework for publishing and reusing digital collections based on literary bibliographies published by GLAM institutions in order to make them suitable for computational use in the form of Collections as Data, in particular, in the context of the European Literary Bibliography. It also describes the infrastructure used and DH research scenarios to illustrate how the results can be reused for different goals. Digital curators and DH researchers interested in making their datasets available in the form of Collections as Data are the intended audience of this work.
Published on June 16, 2025
PDF
Article
A Workflow to publish Collections as Data: looking back at Europeana.eu and forward to the common European data space for cultural heritage
Gustavo Candela, Sally Chambers, Alba Irollo, Nuno Freire, Vicky Dritsou, Antoine Isaac, Agiatis Benardou, Vicky Garnett, Toma Tasovac
Abstract
For decades, cultural heritage (CH) institutions have been making their digital collections available for potential communities of users. Recent advances in technology, such as machine learning, have provided a new context in which digital collections are a rich resource that can be analysed and reused by means of computational methods, for example in the humanities. Initiatives such as Collections as Data, the FAIR (findable, accessible, interoperable, reusable and CARE (collective benefit, authority to control, responsibility, ethics) data principles, and experimental Labs provide best practices and guidelines for publishing digital collections suitable for responsible computational use. In addition, data spaces have recently emerged as a new concept to foster creation, access and reuse of heritage data in which CH institutions play a key role as data providers. In this work we present a workflow for adopting Collections as Data in CH data spaces. The workflow has been developed in the context of the common European data space for cultural heritage and published on the Social Sciences and Humanities (SSH) Open Marketplace. It aims to support CH institutions, humanities researchers and computer scientists interested in making CH data available in data spaces. The article illustrates how the workflow can be adopted by CH institutions or applied by researchers wishing to publish digital collections suitable for computational use. We also demonstrate how the workflow is being adopted in the common European data space for cultural heritage, and describe a selection of potential areas that we believe will most benefit from the wide application of the workflow in the CH domain.
Published on June 24, 2025
PDF
Article
Adding Every Arabic Periodical Published Before 1930 to Wikidata: Moving the Scholarly Crowd-Sourcing Project Jarāʾid to the Digital Commons
Till Grallert
Abstract
This paper documents the contribution of comprehensive bibliographic data on all Arabic periodicals published before 1930 to Wikidata, the largest public and open knowledge graph. The dataset originated with the scholarly crowdsourcing project Jarāʾid and comprises information on more than 3,000 periodicals, about 2,700 editors and almost 350 holding institutions. As a living union list of Arabic periodicals, the dataset alleviates the infrastructural weaknesses of library catalogues and discovery systems, as well as the epistemic violence of knowledge ecologies. The move to Wikidata addresses the socio-technical shortcomings of our original approach by making the dataset available in a FAIR (findable, accessible, interoperable, reusable) and five-star Linked Open Data environment. In addition, the platform provides multilingual interfaces and robust user management and version control, which significantly improves the usability and maintenance of evolving datasets. The paper details workflows and data models and demonstrates the reusability of our approach in other contexts with a second dataset of periodicals from the Ottoman Empire. Finally, the paper shows how the move to Wikidata generates continuous engagement with wider Wikimedia communities that significantly broaden our knowledge about periodicals and their holdings.
Published on July 23, 2025
PDF
Article
Supporting executable scientific workflows in a clustered Infrastructure: DARIAH-IT and H2IOSC
Emiliano Degl'Innocenti, Francesco Pinna, Alessia Spadi, Federica Spinelli
Abstract
Workflows have become essential in digital humanities, enabling the formalisation, automation and reproducibility of complex research processes. As the humanities increasingly adopt data-driven methodologies, workflows offer structured approaches to manage diverse data and tools while supporting transparency and collaboration. Recognising this need, DARIAH-IT has advanced research infrastructure development by focusing on workflow-based services within the H2IOSC (Humanities and Cultural Heritage Italian Open Science Cloud) project. DARIAH-IT leads the design and implementation of a national cloud system to support digital humanities research, ensuring interoperability, FAIR data practices and semantic integration across disciplines. Central to DARIAH-IT’s effort within H2IOSC is AEON (dAriah sErvice Oriented iNfrastructure), a platform that enables the creation, execution and management of scientific workflows. AEON integrates service provisioning, semantic validation and runtime orchestration, supporting reproducible research and collaborative practices.
Published on October 14, 2025
PDF