Skip to content

METRIN-KG

A FAIRly Rooted Knowledge Graph of Plant Metabolomes, Traits, and Interactions

Vast quantities of mutli-quality, specialised datasets exist across various sub-fields in life sciences, yet the nature of these resources prevents researchers from utilizing the full explanatory power when combined. For instance, biodiversity knowledge suffers from several shortfalls including the Raunkiæran and Eltonian [1], that are primarily focussed on the lack of connected information on species traits, interactions, and their ecological functions. Knowledge graphs represent heterogeneous entities and relationships in a semantically rich, machine-readable and, if planned well, in a machine-understandable AI-ready form, offering a feasible solution.

We present METRIN-KG (MEtabolomes, TRaits, and INteractions Knowledge Graph), a FAIR-compliant knowledge graph developed under the Earth Metabolome Initiative (EMI) that integrates plant metabolomics data with high-dimensional plant trait information and their pairwise interaction records [2].

All organisms on Earth are composed of dense networks of chemical, and ecological relationships. The metabolome, that is the vast array of compounds produced by an organism occupies a central position in this network: it provides links between biotic interactions, genotypes and phenotypes. In the Plantae kingdom alone, the collective metabolite diversity is estimated to span between 1.5 and 25.7 million unique compounds across approximately 400,000 plant species [3], representing a rich and an underexplored source of information available to researchers in drug discovery, ecology, agriculture, and conservation biology.

Despite this richness, metabolomics datasets have previously been generated and utilized in isolation. Experimental results are rarely connected to the broader biological context of the organisms that generated them, for instance their functional traits, phylogenetic relationships, ecological roles, or known interactions with other species. METRIN-KG closes this gap, providing a semantic integration layer that allows for a distributed, heterogeneous datasets into a coherent knowledge resource.

METRIN-KG integrates data from three major, independently maintained resources. The metabolomics data is drawn from the Experimental Natural Products Knowledge Graph (ENPKG) [4], a published resource containing enriched metabolome datasets from over 1,600 tropical plant extracts produced using mass spectrometry. The plant trait data is sourced from the TRY Plant Trait Database [5], a repository of curated functional plant traits. Biotic interaction data are taken from the Global Biotic Interactions (GloBI) database [6], which aggregates pairwise interaction records including herbivory, pollination, parasitism, and predation across a broad taxonomic range.

The integration is carried out using the Resource Description Framework (RDF) in a knowledge graph, structured using a combination of established ontologies and the Earth Metabolome Ontology, a purpose-built ontology for metabolomics data, developed to capture domain-specific concepts not adequately represented in existing controlled vocabularies. The graph is hosted with SPARQL query support, enabling federated queries across metabolome, trait, and interaction data from the endpoint.

All input and output files, RDF datasets, and mapping pipelines related to METRIN-KG are accessible throughZenodo, and the GitHub codebase is maintained under the GNU General Public License v3.0, ensuring full transparency and reproducibility.

Entities within METRIN-KG are assigned persistent, resolvable identifiers by mapping to established concepts like different organismal, chemical and material sample taxonomies. Interoperability is achieved through the reuse of over twenty established domain specific ontologies spanning biology, chemistry, ecology, and experimental metadata. Metadata describing datasets, studies, and provenance are fully represented within the graph through subject-predicate-object triples, enabling users to trace the origin and context of every triple.

The Earth Metabolome Ontology, developed under a CC0 license, captures concepts specific to metabolomics sampling workflows, extraction protocols, and annotation confidence levels that are not served by pre-existing ontologies.

The SPARQL endpoint for METRIN-KG is accessible via an editor interface. Moreover, to lower the barrier to entry for researchers without semantic web experience, we applied ExpasyGPT, a large language model driven tool based on lightweight metadata that allows for querying knowledge graphs in natural language. ExpasyGPT aids in constructing SPARQL queries in response to users’ text-format questions. This helps in mitigation of well-known LLM issues such as hallucinations, lack of domain-specific knowledge, and black-box behavior. Consequently, this makes answers verifiable and reproducible. METRIN-KG’s SPARQL endpoint allows for downloading results in tabular format, while still retaining provenance information of each entity.

To demonstrate the analytical potential of METRIN-KG, we present a series of representative case studies spanning drug discovery, ecology, and biodiversity.This includes simple queries like retrieving information on near-threatened IUCN status, and listing plants producing chemicals useful in agriculture and human health. In addition, the case studies also include complex queries like listing 3-way relationships between plant hosts, allelopaths and parasites. Users can also query for plant species known to produce a specific class of metabolites, retrieve their functional traits related to leaf chemistry and growth strategy, and simultaneously obtain records of their biotic interactions while also pulling wikidata instances for those species through a federated query. This kind of multi-dimensional retrieval, which would require coordinating multiple separate API calls and partially automated data harmonisation in a traditional workflow, is rendered straightforward through the knowledge graph frame.

METRIN-KG’s architecture is designed to be extensible: additional data sources, ontologies, and taxonomic scopes can be incorporated incrementally, allowing for flexibility.

Our work demonstrates that the concept of knowledge graphs, when applied with detailed attention to FAIR principles, ontology reuse, and open infrastructure, can advance the integration of heterogeneous life science datasets. The current release of METRIN-KG presents a path for the future development, where we will extend metabolome coverage to the full breadth of plant species represented in the ongoing EMI work, and will subsequently incorporate metabolite data from non-plant kingdoms, bringing the vision of a global, cross-kingdom metabolome knowledge graph closer to realisation.

knowledge graph, metabolomics, plant traits, biotic interactions, FAIR data, Earth Metabolome Initiative, RDF, SPARQL, ontology, biodiversity informatics

[1]Hortal J, Bello FD, Diniz-Filho JAF, Lewinsohn TM, Lobo JM, Ladle RJ. Seven Shortfalls that Beset Large-Scale Knowledge of Biodiversity. Annual Review of Ecology, Evolution, and Systematics 2015;46:523--549.
[2]Tandon D, De Farias TM, Allard P, Defossez E. METRIN-KG: A knowledge graph integrating plant metabolites, traits, and biotic interactions. GigaScience 2026;15:giag051.
[3]Hart CE, Gadiya Y, Kind T, et al. Defining the limits of plant chemical space: challenges and estimations. GigaScience 2025;14:giaf033.
[4]Gaudry A, Pagni M, Mehl F, et al. A Sample-Centric and Knowledge-Driven Computational Framework for Natural Products Drug Discovery. ACS Central Science 2024;10:494--510.
[5]Kattge J, Bönisch G, Díaz S, et al. TRY plant trait database – enhanced coverage and open access. Global Change Biology 2020;26:119--188.
[6]Poelen JH, Simons JD, Mungall CJ. Global biotic interactions: An open infrastructure to share and analyze species-interaction datasets. Ecological Informatics 2014;24:148--159.
DIS
Disha Tandom
Institute of Biology, University of Neuchâtel, CH-2000 Neuchâtel, Switzerland
TAR
Tarcisio Mendes De Farias
SIB Swiss Institute of Bioinformatics, CH-1015 Lausanne, Switzerland
PIE
Pierre-Marie Allard
Department of Biology, University of Fribourg, CH-1700 Fribourg, Switzerland
EMM
Emmanuel Defossez
Institute of Biology, University of Neuchâtel, CH-2000 Neuchâtel, Switzerland