Skip to content

An interactive platform for exploring structural variation across the barley panproteome

Proteins play a central role in translating genetic information encoded in genomes into molecular functions and phenotypic traits. Genetic variation can influence protein sequence, structure, stability, localization, abundance and interaction specificity, thereby shaping biological diversity within a species, yet the systematic characterization of how naturally occurring allelic diversity propagates through protein structural space remains a critical and largely unaddressed challenge in plant genomics. Population genomic approaches can identify protein variants across accessions, but they provide limited mechanistic insight into downstream functional consequences. Bridging this gap requires frameworks that integrate genomic, proteomic, and structural information at scale.

The increasing availability of pangenomes offers an opportunity to move beyond single-reference representations and capture species-wide genetic diversity across large, diverse panels of accessions. Extending this concept to the proteome, a panproteome analysis provides a framework for systematically studying how genetic variation may reshape protein function across a species. Recent advances in protein structure prediction, most notably in models such as AlphaFold [1, 2], have enabled proteome-scale structural modelling [3], opening new opportunities to investigate how genomic variation is reflected in the three-dimensional protein universe of a species. Despite this potential, the structural consequences of naturally occurring allelic variation at panproteome scale remain poorly characterized, partially because existing methods for comparing large numbers of protein structures typically reduce complex multidimensional relationships to global similarity scores or rely on representations that do not show biologically meaningful variation in local conformation or structural divergence [4, 5].

Here, we present a framework for analyzing a predicted species-wide panproteome, applied to the barley (Hordeum vulgare) pangenome version 2 [6], which comprises 76 fully sequenced genotypes divided into 23 wild accessions, 36 landraces, and 17 cultivar lines. From over 3.79 million encoded proteins across this panel, we derived a non-redundant panproteome of 856,971 unique protein sequences, or proteoforms, a term that encompasses allelic variants, paralog-derived sequences, copy-number-associated products, and putative pseudogene-derived proteins. Barley was selected as a model system based on its global agricultural importance, broad adaptive range, and substantial genetic, phenotypic diversity, and data availability, making it well suited to a panproteome-scale investigation.

Our framework comprises two complementary analytical components. The first addresses proteoform classification, mapping proteoforms across genotypes and distinguishing allelic variants, paralogs, and putative pseudogene-derived sequences. The second addresses structural comparison at scale. Predicted structures were represented by Cα atoms from residues with pLDDT > 50, and residue-centered environments within 15 Å were compared using an lDDT-inspired score evaluated at distance-deviation thresholds of 0.5, 1, 2, and 4 Å, with greater weight assigned to smaller structure deviations. The resulting residue-level similarity matrices were decomposed into non-overlapping Smith–Waterman alignments with affine gap penalties, and high-quality local alignments were combined into a pairwise structural-similarity score normalized by structure length and aligned coverage. Structures were subsequently grouped by graph-based threshold clustering, and residue mappings within each cluster were tested for transitive consistency to identify shared structural blocks and distinguish terminal extensions, inter-block insertions, and other patterns of localized structural divergence. Representative structures were selected according to mean within-group similarity. By resolving structural similarity into conserved subregions rather than collapsing each comparison into a single global measure, this approach captures relationships such as partial structural conservation, domain rearrangements, and localized conformational differences that may be obscured by conventional global similarity scores.

More than 450,000 of the 856,971 proteoforms of the panproteome have currently been predicted. Structural analyses completed so far encompass proteoforms associated with approximately 4,000 of the 10,405 functional annotations represented in the dataset and have yielded approximately 18,000 structural clusters. Structural prediction, clustering, and transitivity analyses are continuing across the panproteome. Together with the proteoform classification, the analyses help to reveal that natural allelic diversity can meaningfully reorganize protein structural landscapes in ways that are biologically interpretable, with implications for understanding molecular adaptation across accessions.

A persistent challenge in panproteome research is not only generating data but enabling its exploration. Panproteome datasets are characterized by large numbers of related proteins whose relationships can cut across multiple levels simultaneously: allelic variation, gene duplication, partial annotation, structural similarity, and shared biological function. No single label or similarity measure captures this multidimensional structure. Meaningful interpretation requires integrating proteoform presence and absence of patterns across genotypes, predicted structures, functional annotations, structural cluster assignments, and biological data in a unified and interactive environment. To address this, we developed a web-based exploration platform built on the barley panproteome dataset, designed to enable researchers to navigate allelic and structural diversity without requiring specialist computational infrastructure. The platform links annotation, proteoform distribution within and across genotypes, structural cluster groups, and phenotypic data in a single analytical environment, with particular attention to the approximately 25% of proteoforms that lack functional annotation and for which exploratory, to which this kind of structure-guided approaches are essential.

Taken together, this work establishes a reusable and extensible framework for panproteome-scale structural analysis and lays the groundwork for connecting allelic diversity to molecular function in crop species and beyond. The approaches developed here, for proteoform classification, structure-based comparison, and interactive exploration, are applicable to any species for which a pangenome and proteome-scale structure predictions are available. They represent a concrete step toward systematically interpreting the functional consequences of natural genetic variation in the language of protein structure, and toward making panproteome datasets accessible and analytically tractable for the broader research community.

[1]Jumper J, Evans R, Pritzel A, et al. Highly accurate protein structure prediction with AlphaFold. Nature 2021;596:583--589.
[2]Abramson J, Adler J, Dunger J, et al. Accurate structure prediction of biomolecular interactions with AlphaFold 3. Nature 2024;630:493--500.
[3]Varadi M, Anyango S, Deshpande M, et al. AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models. Nucleic Acids Research 2022;50:D439--D444.
[4]van Kempen M, Kim SS, Tumescheit C, et al. Fast and accurate protein structure search with Foldseek. Nature Biotechnology 2024;42:243--246.
[5]Barrio-Hernandez I, Yeo J, Jänes J, et al. Clustering predicted structures at the scale of the known protein universe. Nature 2023;622:637--645.
[6]Jayakodi M, Lu Q, Pidon H, et al. Structural variation in the pangenome of wild and domesticated barley. Nature 2024;636:654--662.
VIC
Victor Henrique Rabesquine Nogueira
Leibniz Institute of Plant Genetics and Crop Plant Research (IPK) Gatersleben, Germany
AMA
Amanda Souza Camara
Leibniz Institute of Plant Genetics and Crop Plant Research (IPK) Gatersleben, Germany