Enhancing e!DAL-PGP
The rapid increase in the volume, heterogeneity, and complexity of plant science research data requires robust research data management (RDM) and data publication infrastructures that support FAIR data principles [1], metadata standards, and scalable submission workflows. While numerous domain-specific repositories exist for selected data types, many datasets generated in the plant sciences remain difficult to publish, annotate, and integrate into reusable data ecosystems due to volume, delayed and fragmented submission workflows and inconsistent metadata practices. The de.NBI service and GFBio data center e!DAL-PGP [2] addresses this challenge by providing a repository for long-term resolvable, DOI features publication and exposing of heterogeneous plant research datasets for machine actionable access or harvesting. However, the current submission infrastructure is based on legacy desktop clients requiring platform-specific deployment and maintenance, thereby limiting usability, and long-term sustainability.
To address these limitations, the BioHackathon Germany 2025 project “Enhancing e!DAL-PGP: A Modern Data Submission Platform for Plant Science Research Data” focused on the development of a modern, interoperable, and FAIR-oriented submission ecosystem for plant science research data. The project brought together contributors from de.NBI [3], DataPLANT [4], FAIRagro [5] and NFDI4Biodiversity [6] to collaboratively design and prototype solutions that improve usability, metadata quality, interoperability, and automated data publication workflows.
The project was structured into three complementary workstreams: (i) development of a web-based submission platform for e!DAL-PGP, (ii) integration of automated validation and submission workflows with the PLANTdataHUB ecosystem of DataPLANT [7], and (iii) development of documentation, training resources, and community support structures to facilitate sustainable adoption.
A major outcome of the project was the implementation of a prototype web-based submission client designed to replace the desktop-based submission client. The new interface guides users through the complete submission process using a structured multi-step wizard that emphasizes usability, metadata completeness, and FAIR-compliant data publication practices.
The redesigned e!DAL-PGP submission workflow starts with the mandatory acceptance of the Data Deposit and License Agreement (DLA), ensuring legal clarity and repository compliance before metadata entry or data upload are initiated. Subsequent steps transform the dataset documentation into highly interoperable, machine-readable assets. Users are guided through structured forms capturing core descriptors like titles, authors, abstracts, and licensing terms, with safeguards enforcing mandatory fields. To boost semantic interoperability for downstream AI discovery, the system embeds an autocomplete feature powered by the TS4NFDI terminology service [8] and TIB’s DataPLANT ontology collection, while still permitting custom annotations. Precise author management is achieved by integrating ORCID and ROR persistent identifiers. Scalable data upload accommodates massive, heterogeneous datasets through browser-based uploads or direct, S3-compatible storage such as NFDI4Biodiversity’s Aruna Engine. Finally, automated validation routines block hidden files and enforce transparent directory structures instead of opaque archives, concluding with a dynamic citation preview before final submission.
In parallel, the project established an interoperable pipeline between e!DAL-PGP and the PLANTdataHUB ecosystem, focusing on automated submission and continuous integration (CI) workflows for Annotated Research Contexts (ARCs). To bridge data validation and repository submission, the project expanded the ARC ecosystem to support both F#- and Python-based validation frameworks. This unified approach automates quality checks against mandatory DataCite metadata standards directly within automated workflows, generating standardized validation reports and machine-readable summaries before datasets ever reach the repository. Once a dataset successfully passes this automated validation safeguard, an integrated submission pathway handles the data transfer. Relevant metadata such as titles, licenses, descriptions, and author attributions is automatically extracted from the validated ARC RO Crate and used to prepopulate the e!DAL-PGP submission interface. By linking automated CI validation directly to the submission process, this workflow significantly reduces manual data entry, prevents errors, and preserves user control over final publication, serving as a model for FAIR-compliant pipelines across distributed life science infrastructures.
To support sustainable community adoption, the project additionally initiated the development of a dedicated e!DAL-PGP Knowledge Base based on the Astro and Starlight frameworks. The knowledge base will replace the existing project webpage, is intended as a centralized entry point for users, developers, and data stewards and includes technical documentation, contribution guidelines, and best-practice recommendations
Overall, the BioHackathon project demonstrates how collaborative and community-driven development formats can accelerate the modernization of research data infrastructures and foster interoperability across networks such as the NFDI and ELIXIR. The presented work contributes to the development of FAIR, interoperable, and AI-ready data publication workflows by combining modern web technologies, ontology-supported metadata annotation, automated validation infrastructures, and machine-actionable repository integration. The poster presents the developed architecture, implemented prototypes, integration concepts, and lessons learned during the collaborative development process, highlighting future directions towards scalable and sustainable research data infrastructures for plant science and beyond.