Skip to content

Provenance-Aware Spreadsheet FAIRification Using Layered Metadata Integration and RDM-based Storage

Database scientists have long criticized the use of spreadsheets for scientific data management because spreadsheets lack amongst others explicit schemas, transactions, and declarative query mechanisms [1-3]. Nevertheless, spreadsheets remain among the most widely used tools for research data collection, transformation, annotation, and exchange across many scientific disciplines. Since Lotus 1-2-3, the first spreadsheet system that also considered the spreadsheet as a data collection, continuous improvements in usability and flexibility have further increased their attractiveness as temporary and collaborative data stores.

A key strength of spreadsheets is the workbook concept. Workbooks are ordered collections of sheets that allow researchers to combine data, calculations, metadata, annotations, and workflow context within a single portable research object. However, while spreadsheet software simplifies data entry and ad hoc analysis, it does not solve the problem of data harmonization. Due to their flexibility, spreadsheets created by different researchers frequently differ substantially in structure and semantics unless harmonization efforts are applied explicitly.

This motivates the development of tools that facilitate spreadsheet FAIRification, transformation, annotation, and provenance tracking. Existing tools, however, focus on different aspects of spreadsheet workflows. Transformation-oriented tools such as OpenRefine ( https://openrefine.org/) and ReStoRunT [4] support reshaping and documenting processing workflows. Annotation-oriented tools such as RightField and (ARC-)SWATE (https://arc-rdm.org/details/tools-and-services/SWATE/) [5] support ontology-based semantic enrichment of spreadsheet data.

At first sight, these tools appear complementary. However, closer inspection reveals substantial interoperability problems. We argue that one major reason for these limitations is that different tools implicitly operate on different provenance and metadata layers without clearly distinguishing them.

We distinguish at least three provenance and metadata layers relevant for spreadsheet-centered FAIRification workflows:

  1. Operational provenance, describing transformation and processing steps applied during data cleaning and restructuring workflows. OpenRefine and ReStoRunT primarily support this layer by recording transformation histories or storing provenance information within spreadsheet formulas.

  2. Explicit metadata annotations, describing semantic enrichment intentionally added by researchers, including ontology annotations, units, protocols, variable descriptions, and experimental context. (ARC-)SWATE and RightField+FAIRDOM SEEK primarily support this layer through ontology-guided spreadsheet annotation workflows integrated or for integration in a data management solution.

  3. Structural spreadsheet context, describing organizational information encoded through workbook structures, sheet ordering, hidden sheets, formulas, cross-sheet references, and spreadsheet layout conventions that frequently contain implicit scientific meaning.

Current spreadsheet FAIRification tools preserve only subsets of these layers. As a consequence, interoperation between tools often leads to information loss. For example, OpenRefine lets the user operate on single-table projects and does not preserve workbook structures or hidden sheets during import and export. This prevents straightforward interoperability with RightField, which stores ontology-related information in hidden workbook sheets. Similarly, ReStoRunT stores provenance information in spreadsheet formulas, whereas OpenRefine primarily operates on values rather than formulas, resulting in provenance loss during transformation workflows.

These examples demonstrate that spreadsheet interoperability should not be treated solely as a table-conversion problem. Instead, it should be understood as a provenance-aware integration problem involving multiple distinct metadata layers.

To address this challenge, we propose a provenance-aware spreadsheet FAIRification architecture that explicitly separates operational provenance, semantic metadata annotations, and research data content while maintaining transparent spreadsheet-centered workflows for users. Our approach builds upon OpenRefine as an operational provenance backbone while extending spreadsheet workflows with explicit metadata integration, validation, and ontology-aware annotation support.

As underlying storage representation, we propose the use of ARC-based or similar RDM structures that allow transparent separation of metadata layers while remaining compatible with spreadsheet-oriented user interfaces. This separation enables machine-actionable metadata integration, facilitates validation workflows, and supports direct compatibility with RO-Crate-based research object representations.

Rather than replacing spreadsheets, our goal is to preserve the advantages of spreadsheet-centered research workflows while improving interoperability, provenance preservation, and FAIR compliance. Based on experiences from teaching, FAIRification workflows, and practical research data management scenarios, we discuss requirements and challenges for provenance-aware spreadsheet interoperability in integrative bioinformatics and related disciplines.

[1]Panko RR. What We Know About Spreadsheet Errors. Journal of Organizational and End User Computing (JOEUC) 1998;10:15--21.
[2]Karr JR, Liebermeister W, Goldberg AP, Sekar JAP, Shaikh B. ObjTables: structured spreadsheets that promote data quality, reuse, and integration. arXiv; 2020.
[3]Bendre M, Venkataraman V, Zhou X, Chang K, Parameswaran A. Towards a Holistic Integration of Spreadsheets with Databases: A Scalable Storage Engine for Presentational Data Management.
[4]Müller M, Lukrécia M. ReStoRunT: Simple Recording, Storing, Running and Tracing changes in Spreadsheets.
[5]Doniparthi G, Mühlhaus T, Deßloch S. Integrating FAIR Experimental Metadata for Multi-omics Data Analysis. Datenbank-Spektrum 2024;24:107--115.
WOL
Wolfgang Müller
Scientific Databases and Visualization, Heidelberg Institute for Theoretical Studies (HITS), Heidelberg, Germany
TIM
Timo Mühlhaus
Computational Systems Biology, Rhineland-Palatinate Technical University, Kaiserslautern, Germany
BIR
Birgitta König-Ries
Friedrich-Schiller-Universität Jena