Towards FAIR NGS Data Management Across the Data Life Cycle
Sequencing data underpins the generation of a wide range of so-called OMICS data types. Thereby, OMICS refers to the comprehensive characterization and quantification of biological entities using high-throughput technologies. Next-generation sequencing (NGS) represents a key high-throughput technology and produces massive volumes of data. Core molecular layers, such as the genome and transcriptome, are routinely analysed using sequencing approaches. Beyond these layers, numerous sequencing-based protocols exist to characterize regulatory mechanisms and functional interactions, for example, by quantifying chemical modifications of DNA and RNA. Almost all NGS applications generate large data volumes per sample, ranging from gigabytes to terabytes. Moreover, different sequencing technologies yield data with distinct characteristics, often requiring specialized, technology-specific processing pipelines to be executed within high-performance computing environments. Repositories for the permanent deposition of public NGS data are well established and have implemented many mechanisms supporting the FAIR principles (Findable, Accessible, Interoperable, Reusable). Major resources include the European Nucleotide Archive (ENA), which is synchronized with its US and Japanese counterparts. For controlled-access human data, these are complemented by the European Genome-phenome Archive (EGA) as well as national infrastructures such as those provided by the NFDI consortium GHGA. While these repositories provide a strong foundation for FAIR data stewardship and long-term accessibility, achieving full FAIRness remains an ongoing process that depends on data quality, metadata completeness, community standards, and repository-specific implementations. Despite this, ensuring FAIR management of NGS data throughout the entire data life cycle, i.e., the sequence of stages that data goes through from its creation and storage to processing, analysis, sharing, and eventual archiving or deletion, remains challenging for researchers, sequencing facilities, IT services, and data stewards alike. We see two sets of difficulties: (a) the lack of overarching approaches that address the heterogeneous needs of these stakeholders. Contributing factors include the sheer raw data volume and the diversity of biological domains and experimental contexts in which NGS data are generated. This is reflected in multiple NFDI consortia covering NGS-based OMICS data. Examples include NFDI4Microbiota, GHGA, DataPLANT, NFDI4BIOIMAGE, and NFDI4Biodiversity. While several of these initiatives develop research data management (RDM) solutions, they typically focus on the needs of a single community. Biology-centered research institutions, however, routinely generate data spanning multiple such domains. An example is human–microbe interactions, which are extensively studied in the context of infection or microbial colonization. The second (b) key difficulty is that there is a need for a data management closer to the institutions, e.g., for data that are not allowed to leave the premises of the institution, or data that are simply too immature to handle and share using a national data venue. Summarizing, there is a need for structures that offer consistent linkage between metadata over time, close to the institution, making data ready for use within the NFDIs. We are aware that the NFDIs aim at further harmonisation and integration. However, we see the need for shorter-term efforts that increase the institutions’ readiness for collaboration with today’s and future setups of the NFDI(s).
We conducted a needs analysis within our institutes and their scientific networks. Requirements from different stakeholder groups were collected and investigated. Established NGS repositories for long-term archiving were reviewed, and existing RDM solutions developed within the NFDI were assessed with respect to interoperability, coverage of identified requirements, and integration with established NGS repositories.
We observe that while several NFDI consortia have developed RDM solutions, none currently support overarching management of multi-domain OMICS data. National efforts are paralleled by institution-specific solutions that address subsets of the overall requirements but do not scale across domains or institutions. ENA and EGA are identified as the central public repositories for the archiving of non-human and access-restricted human NGS data, respectively. A key requirement emerging from our analysis is the ability to broker raw data from institutional RDM systems to these repositories. In parallel, NGS data must often be retained on-premises and distributed across heterogeneous storage systems, while still enabling central access and controlled sharing between institutions. Access control must support fine-grained authorization at the level of individual data items. Importantly, this entails sharing selected metadata openly while restricting access to the underlying data, which must be explicitly requested and granted, possibly involving a data access agreement. Comprehensive metadata capture is essential and must encompass NGS-specific technical metadata, center- and project-specific information, domain-specific sample metadata, and links to electronic lab books. Additionally, selected key result data should be recorded, shared, and assigned persistent identifiers such as DOIs. Access restrictions and detailed permission models are required at all stages of analysis and across institutional boundaries. Based on the results of our needs analysis, we present a concept for a decentralized infrastructure that supports the sharing of NGS data and associated metadata across institutions, independent of physical storage systems and locations. For metadata management, we propose the use of FAIRDOM SEEK as a central cataloguing platform. SEEK enables structured representation of heterogeneous research assets, including samples, tabular results data, associated metadata, and protocols, while preserving their relationships within studies and investigations based on the ISA (Investigations, Studies, Assets) framework. It further supports fine-grained permission and access-control mechanisms for private and pre-publication data, semantic annotation, and DOI assignment, thereby enabling FAIR-compliant metadata management across the NGS data life cycle. SEEK is proposed as a suitable candidate because several of the authors actively contribute to the development and operation of SEEK-based infrastructures and therefore have direct experience with its deployment in research data management environments. While other metadata catalogues may also be applicable, SEEK was selected as a concrete and mature example that addresses the requirements identified in our analysis. Building on SEEK-based implementations by ELIXIR Belgium for their DataHub, we propose incorporating SEEK-based brokering functionalities that integrate EBI repository metadata standards, facilitating seamless submission of NGS data to the ENA and EGA archives. For decentralized NGS raw data storage and access, we propose the integration of Aruna as a federation-first data management system. Aruna enables federated data management across distributed infrastructures, supporting metadata replication, data locality-aware compute integration, and secure governance across institutional boundaries. This allows data to remain primarily within institutional domains while still being shareable in a coordinated, FAIR-compliant manner. Aruna acts as a data abstraction layer, harmonizing data access and integration across diverse storage mechanisms, including POSIX filesystems and object storage. Aruna was selected as a concrete implementation because authors actively contribute to its development and therefore have direct experience with its capabilities and deployment. In addition, Aruna serves as the data management platform within NFDI4Microbiota, providing a relevant reference implementation in the German NFDI landscape. In the proposed architecture, SEEK and Aruna fulfill complementary roles: SEEK serves as the metadata catalogue and coordination layer, whereas Aruna provides the abstraction and federation layer for distributed raw data storage and access. Together they enable consistent linkage between metadata and physically distributed datasets across institutional boundaries. This approach enables seamless and harmonized access to NGS data stored in heterogeneous environments, ranging from local file systems to cloud and object storage, while ensuring consistent linkage between data and metadata across the entire data life cycle.
With this contribution, we aim to stimulate discussion and community alignment on requirements and design choices of cross-domain, institution-centered NGS data management across the data life cycle. This will serve as a basis for positioning the proposed concept within the broader national and international RDM landscape to build on existing efforts, complement ongoing initiatives, and identify synergies. Implementing FAIR principles consistently across the entire NGS data life cycle is crucial for enabling integrative bioinformatics and data-driven discovery. NGS data can only be combined across OMICS layers, studies, institutions, and scientific domains if the raw data and key results are reliably findable, interoperable, and reusable. A data life-cycle-oriented, FAIR-by-design infrastructure not only improves reproducibility and long-term usability of NGS data, but also lowers technical barriers for cross-domain integration and secondary analyses. This is particularly important as modern bioinformatics increasingly relies on the joint analysis of heterogeneous OMICS datasets to address complex biological and biomedical questions.