
Doreen Ware
Doreen Ware is a molecular and computational biologist with the USDA Agricultural Research Service (ARS) and an Adjunct Professor at Cold Spring Harbor Laboratory. Her research focuses on plant genomics, comparative genomics, genetic variation, and the development of computational resources and data infrastructure that connect genomic information with biological function and agricultural traits.
Her group develops and supports community resources including Gramene and SorghumBase and works on the integration of reference genomes, pan-genomes, genetic variation, gene expression, phenotypes, and functional genomic data across agriculturally important species. An important component of this work is the development and adoption of standards that make biological data more interoperable, reusable, and accessible across research communities.
Ware contributed to the sequencing and analysis of the plant genomes including sorghum, maize and rice, and has participated in numerous large-scale plant genomics and comparative genomics projects. Her current research is increasingly focused on developing integrated and AI-ready biological datasets and infrastructure that can support predictive approaches in plant biology and agriculture.
She is a Fellow of the American Association for the Advancement of Science (AAAS) and previously served as Acting Chief Scientific Information Officer for USDA ARS from 2014–2017.
From Genomes to Prediction: Building the Data Infrastructure for Future Biological Discovery
Section titled “From Genomes to Prediction: Building the Data Infrastructure for Future Biological Discovery”Life science is generating data at an increasing scale and complexity. In plant science, we have moved from individual reference genomes to pan-genomes and population-scale variation, while also generating new functional, phenotypic, environmental, and single-cell datasets. These technologies provide new ways to address long-standing scientific questions, but also create challenges for how we describe, connect, maintain, and reuse the resulting data.
Our research is addressing some of these challenges through community resources such as Gramene and SorghumBase, connecting genomes and genetic variation with expression, function, phenotypes, and comparative genomics. Standards, persistent identifiers, common vocabularies, interoperable metadata, and reproducible workflows are increasingly important for connecting data across species, technologies, and research communities.
Looking forward, our infrastructure will need to support both new types of biological data and new ways of using these data. This includes making resources interoperable, computationally accessible, and AI-ready, while maintaining provenance and connections to the underlying biology. This talk will discuss how advances in genomics and functional biology are shaping these needs and how we can build sustainable community resources that support both current research and future biological discovery.