Skip to content

Legal Metadata in FAIRagro

FAIRagro is one of the consortia within the German National Research Data Infrastructure (NFDI). The consortium brings together a broad range of Research Data Infrastructures (RDIs), scientific institutions, and domain experts covering diverse domains such as plant science, soil science, biodiversity research, agricultural systems, and environmental monitoring. [1] A central objective of FAIRagro is to enhance the findability, accessibility, interoperability, and reusability (FAIRness) of research data across heterogeneous infrastructures and disciplinary boundaries. In order to achieve this, FAIRagro develops harmonised metadata standards, technical integration workflows, and domain-specific FAIR Digital Object approaches. Within this context, legal metadata play an increasingly important role, as transparent information about licensing conditions, access restrictions, and regulatory obligations is essential for enabling legally compliant and FAIR reuse of agrosystems data.

Currently, RDIs in agrosystems research differ considerably in their handling of licensing, access restrictions, data protection, and regulatory compliance. These differences create challenges for the FAIR-compliant reuse of research data across infrastructures. FAIRagro therefore develops a framework for representing legal metadata in a standardised, machine-actionable, and interoperable manner that will be integrated in the FAIRagro middleware. Three central application areas for legal metadata were identified: licences, data access conditions, and regulatory requirements. Each of these areas requires different forms of metadata representation. The framework aims not only to improve transparency for users but also to support interoperability and enable future integration into automated workflows and machine-actionable services.

The FAIRagro middleware is intended to integrate legal metadata from participating RDIs whenever they are available in existing metadata records. These metadata will be harmonised and translated into JSON-LD representations that reference reusable ODRL policies through the Schema.org property usageInfo. Where structured legal metadata are not yet available, FAIRagro will work with the respective RDIs to identify the relevant legal information and map it to standardised policy templates. This approach is intended to enable a consistent representation of legal information across heterogeneous infrastructures while minimising implementation effort for individual repositories.

The first application area concerns licensing information. An analysis of clusters of agrosystems data in accordance with relevant legal categories across participating RDIs demonstrated that licences from the Creative Commons (CC) family are predominantly used. [2] Existing practices include the use of CC0, CC-BY, and CC-BY-NC licences or licensing models according to the Creative Commons principles. FAIRagro therefore aims to harmonise licensing information using standardised metadata representations while still allowing flexibility for infrastructure-specific requirements and domain-specific needs. In the proposed framework, licensing information is represented at the dataset level and embedded directly into Schema.org metadata. The Schema.org property usageInfo is employed to link datasets to detailed machine-readable legal policies. Schema.org is used, as it is the standard for metadata in FAIRagro and beyond [3, 4]. The property usageInfo is used, as it is designed to be used alongside the license property and still is part of the schema.org definition, which is not the case for an also possible additionalProperty. The policies are represented using the Open Digital Rights Language (ODRL) and are referenced through resolvable URLs or persistent identifiers. By using ODRL the approach allows for efficient machine actionability compared to approaches which use plain text to document regulatory information such as Croissant-ML RAI [5].

ODRL was selected as the central standard for legal metadata representation because it provides a flexible and interoperable model for expressing permissions, prohibitions, duties, and constraints. ODRL policies can therefore describe both simple and complex legal conditions in a structured and machine-actionable manner [6]. FAIRagro uses ODRL not only for licensing statements but also for the representation of data access conditions and regulatory obligations. Reusable ODRL policy templates for the most frequently used access and licensing scenarios are versioned and maintained in the FAIRagro Git repository in order to ensure transparency, reuse, and long-term maintainability [7].

Currently, FAIRagro provides reusable policy templates for three representative access scenarios: unrestricted open access, authentication-required access, and access upon request. These templates cover the most common legal situations encountered across participating RDIs and can be reused or extended for repository-specific requirements. [8]

The second application area focuses on data access conditions. Access restrictions are highly heterogeneous across RDIs, ranging from open access to authentication-based access and manual request procedures. FAIRagro therefore develops standardised ODRL policies for these common access scenarios, which are also made available through the FAIRagro Git repository.

With the policy templates integrated into dataset metadata, the FAIRagro SearchHub can display access conditions, licence information, and links to policies in a human-readable way, allowing users to better understand how datasets may be accessed or reused. At the same time, the machine-readable representation enables future integration into automated workflows.

The current implementation focuses on rendering legal information rather than enforcing legal policies automatically. This allows users to discover restricted datasets, understand applicable licences and access conditions, and navigate to the appropriate repository workflows while establishing the technical foundation for future machine-actionable policy enforcement.

The third application area concerns regulatory requirements arising, for example, from international agreements and domain-specific legal frameworks. In agrosystems research, datasets may be subject to obligations derived from the Nagoya Protocol on Access and Benefit-Sharing or from the International Treaty on Plant Genetic Resources for Food and Agriculture. Such frameworks often require documentation regarding the origin of biological material, material transfer agreements, or usage restrictions. FAIRagro therefore explores how these obligations can be partially represented using structured metadata and ODRL policies. Although not all compliance information can currently be expressed entirely within metadata records, the proposed framework establishes a foundation for improving transparency and traceability of legally relevant information.

The proposed approach contributes to the development of interoperable legal metadata standards within the NFDI and beyond. By combining FAIR Digital Object concepts and reusable ODRL policy templates, FAIRagro establishes a practical and extensible framework for legal metadata in agrosystems research. The framework enables transparent communication of legal conditions, improves the discoverability of restricted datasets, and supports the FAIR principles by making legal information more accessible and machine-readable.

The presentation will demonstrate the workflow from legal metadata provided by RDIs to harmonised JSON-LD representations, reusable ODRL policies, and their integration into the FAIRagro SearchHub. Concrete examples from the implemented policy templates will illustrate how legal information can be represented consistently across heterogeneous infrastructures. In addition, the talk will discuss current limitations, interoperability challenges, and future developments, including machine-actionable policies and community feedback on extending the approach to other life science research infrastructures.

Future work will focus on extending policy templates, harmonising legal metadata across additional infrastructures, and evaluating possibilities for more advanced policy enforcement and workflow integration. The presentation aims on feedback and possible hints to solutions which we might have overseen.

[1]Specka X, Martini D, Weiland C, et al. FAIRagro: Ein Konsortium in der Nationalen Forschungsdateninfrastruktur (NFDI) für Forschungsdaten in der Agrosystemforschung. Informatik Spektrum 2023;46:24--35.
[2]Lesch S, Scheuner C, Singson LS. Clusters of agrosystems data in accordance with relevant legal categories. 2024.
[3]Möller L, Ernst M, Fichtmüller D, et al. Advancing FAIR Biodiversity Data: Bioschemas Implementation in NFDI4Biodiversity. F1000Research; 2025.
[4]Castro LJ, Fluck J, Arend D, et al. Schema.org as a Lightweight Harmonization Approach for NFDI. Proceedings of the Conference on Research Data Infrastructure 2023;1.
[5]Jain N, Akhtar M, Giner-Miguelez J, et al. A Standardized Machine-readable Dataset Documentation Format for Responsible AI. arXiv; 2024.
[6]Feser M, Beyvers S, Arend D, et al. Simplifying and Standardizing the Creation of Data Use Agreements for Life Sciences and Beyond - BH Germany2024. BioHackrXiv; 2025.
[7]fairagro/odrl\_policies. FAIRagro; 2026.
CON
Constantin Breß
Leibniz-institute for Information Infrastructure (FIZ Karlsruhe), Karlsruhe, Germany
CAR
Carmen Scheuner
Senckenberg Museum of Natural History, Görlitz, Germany
STE
Stephan Lesch
Senckenberg Museum of Natural History, Görlitz, Germany