Skip to content

Enhancing e!DAL-PGP

Towards a FAIR, Interoperable and AI-Ready Submission Platform for Plant Science Research Data

The rapid increase in the volume, heterogeneity, and complexity of plant science research data requires robust research data management (RDM) and data publication infrastructures that support FAIR data principles [1], metadata standards, and scalable submission workflows. While numerous domain-specific repositories exist for selected data types, many datasets generated in the plant sciences remain difficult to publish, annotate, and integrate into reusable data ecosystems due to volume, delayed and fragmented submission workflows and inconsistent metadata practices. The de.NBI service and GFBio data center e!DAL-PGP [2] addresses this challenge by providing a repository for long-term resolvable, DOI features publication and exposing of heterogeneous plant research datasets for machine actionable access or harvesting. However, the current submission infrastructure is based on legacy desktop clients requiring platform-specific deployment and maintenance, thereby limiting usability, and long-term sustainability.

To address these limitations, the BioHackathon Germany 2025 project “Enhancing e!DAL-PGP: A Modern Data Submission Platform for Plant Science Research Data” focused on the development of a modern, interoperable, and FAIR-oriented submission ecosystem for plant science research data. The project brought together contributors from de.NBI [3], DataPLANT [4], FAIRagro [5] and NFDI4Biodiversity [6] to collaboratively design and prototype solutions that improve usability, metadata quality, interoperability, and automated data publication workflows.

The project was structured into three complementary workstreams: (i) development of a web-based submission platform for e!DAL-PGP, (ii) integration of automated validation and submission workflows with the PLANTdataHUB ecosystem of DataPLANT [7], and (iii) development of documentation, training resources, and community support structures to facilitate sustainable adoption.

A major outcome of the project was the implementation of a prototype web-based submission client designed to replace the desktop-based submission client. The new interface guides users through the complete submission process using a structured multi-step wizard that emphasizes usability, metadata completeness, and FAIR-compliant data publication practices.

The redesigned e!DAL-PGP submission workflow starts with the mandatory acceptance of the Data Deposit and License Agreement (DLA), ensuring legal clarity and repository compliance before metadata entry or data upload are initiated. Subsequent steps transform the dataset documentation into highly interoperable, machine-readable assets. Users are guided through structured forms capturing core descriptors like titles, authors, abstracts, and licensing terms, with safeguards enforcing mandatory fields. To boost semantic interoperability for downstream AI discovery, the system embeds an autocomplete feature powered by the TS4NFDI terminology service [8] and TIB’s DataPLANT ontology collection, while still permitting custom annotations. Precise author management is achieved by integrating ORCID and ROR persistent identifiers. Scalable data upload accommodates massive, heterogeneous datasets through browser-based uploads or direct, S3-compatible storage such as NFDI4Biodiversity’s Aruna Engine. Finally, automated validation routines block hidden files and enforce transparent directory structures instead of opaque archives, concluding with a dynamic citation preview before final submission.

In parallel, the project established an interoperable pipeline between e!DAL-PGP and the PLANTdataHUB ecosystem, focusing on automated submission and continuous integration (CI) workflows for Annotated Research Contexts (ARCs). To bridge data validation and repository submission, the project expanded the ARC ecosystem to support both F#- and Python-based validation frameworks. This unified approach automates quality checks against mandatory DataCite metadata standards directly within automated workflows, generating standardized validation reports and machine-readable summaries before datasets ever reach the repository. Once a dataset successfully passes this automated validation safeguard, an integrated submission pathway handles the data transfer. Relevant metadata such as titles, licenses, descriptions, and author attributions is automatically extracted from the validated ARC RO Crate and used to prepopulate the e!DAL-PGP submission interface. By linking automated CI validation directly to the submission process, this workflow significantly reduces manual data entry, prevents errors, and preserves user control over final publication, serving as a model for FAIR-compliant pipelines across distributed life science infrastructures.

To support sustainable community adoption, the project additionally initiated the development of a dedicated e!DAL-PGP Knowledge Base based on the Astro and Starlight frameworks. The knowledge base will replace the existing project webpage, is intended as a centralized entry point for users, developers, and data stewards and includes technical documentation, contribution guidelines, and best-practice recommendations

Overall, the BioHackathon project demonstrates how collaborative and community-driven development formats can accelerate the modernization of research data infrastructures and foster interoperability across networks such as the NFDI and ELIXIR. The presented work contributes to the development of FAIR, interoperable, and AI-ready data publication workflows by combining modern web technologies, ontology-supported metadata annotation, automated validation infrastructures, and machine-actionable repository integration. The poster presents the developed architecture, implemented prototypes, integration concepts, and lessons learned during the collaborative development process, highlighting future directions towards scalable and sustainable research data infrastructures for plant science and beyond.

[1]Wilkinson MD, Dumontier M, Aalbersberg IJ, et al. The FAIR Guiding Principles for scientific data management and stewardship. Scientific Data 2016;3:160018.
[2]Arend D, König P, Junker A, Scholz U, Lange M. The on-premise data sharing infrastructure e!DAL: Foster FAIR data for faster data acquisition. GigaScience 2020;9:giaa107.
[3]Barysch S, Buchhalter I, Dammann-Kalinowski T, et al. Position paper on cooperation between NFDI and de.NBI \& ELIXIR-DE. 2025.
[4]Suchodoletz DV, Mühlhaus T, Krüger J, Usadel B, Rodrigues CM. DataPLANT – Ein NFDI-Konsortium der Pflanzen-Grundlagenforschung. Bausteine Forschungsdatenmanagement 2021:46--56.
[5]Ewert F, Specka X, Anderson JM, et al. FAIRagro - A FAIR Data Infrastructure for Agrosystems (proposal). 2023.
[6]Glöckner FO, Diepenbroek M, Felden J, et al. NFDI4BioDiversity - A Consortium for the National Research Data Infrastructure (NFDI). 2020.
[7]Weil HL, Schneider K, Tschöpe M, et al. PLANTdataHUB: a collaborative platform for continuous FAIR data sharing in plant research. The Plant Journal 2023;116:974--988.
[8]Baum R, Bouazouni S, Fillies J, et al. Terminology Services in the DACH Region Landscape – What Are the Essential Requirements?. heiBOOKS; 2025.
MAN
Manuel Feser
Leibniz Institute of Plant Genetics and Crop Plant Research (IPK) Gatersleben, Germany
JON
Jonathan Bauer
eScience Departement, Computer Center, Albert Ludwig University of Freiburg, Germany
Avatar Tailwind CSS Component
Sebastian Beier
Institute of Bio- and Geosciences (IBG-4 Bioinformatics), Bioeconomy Science Center (BioSC), CEPLAS, Forschungszentrum Jülich GmbH, 52425 Jülich, Germany
Avatar Tailwind CSS Component
Dominik Brilhaus
Cluster of Excellence on Plant Sciences (CEPLAS), Faculty of Mathematics and Natural Science, Heinrich Heine University Düsseldorf, Düsseldorf, Germany
JAG
Jagadeeshwar Reddy Etukala
Leibniz Institute of Plant Genetics and Crop Plant Research (IPK) Gatersleben, Germany
ELE
Elena Rey Mazón
Leibniz Institute of Plant Genetics and Crop Plant Research (IPK) Gatersleben, Germany
MAT
Matthias Lange
Leibniz Institute of Plant Genetics and Crop Plant Research (IPK) Gatersleben, Germany
DEN
Dennis Psaroudakis
Institute of Agricultural and Nutritional Sciences, Martin Luther University Halle-Wittenberg, Germany
KEV
Kevin Schneider
Computational Systems Biology, Rhineland-Palatinate Technical University, Kaiserslautern, Germany
HEL
Helena Schnitzer
Forschungszentrum Jülich GmbH - IBG-5; de.NBI & ELIXIR-DE
UWE
Uwe Scholz
Leibniz Institute of Plant Genetics and Crop Plant Research (IPK) Gatersleben, Germany
HEI
Heinrich Lukas Weil
Computational Systems Biology, Rhineland-Palatinate Technical University, Kaiserslautern, Germany
Avatar Tailwind CSS Component
Daniel Arend
Leibniz Institute of Plant Genetics and Crop Plant Research (IPK) Gatersleben, Germany