Native Crop Pangenomes for Resilient Varieties
Graph Genomes and Pangenomes of Native Crops for Designing Climate-Resilient Varieties
When a native variety survives drought, heat, salinity, or disease pressure, part of the answer lies in genes that are not always visible in a commercial reference genome. A single reference genome functions like an essential but incomplete map; it places reads, variants, and genes along one linear path, but if an allele or DNA segment is absent from that path, it may be misaligned or excluded from the analysis. For climate-resilient agriculture, this limitation is not merely a technical issue in the laboratory. It is tied to variety selection, food security, the economics of water, and investment decisions across the seed value chain. As climate stresses put more pressure on production, the value of seeing the hidden layers of genetic diversity left outside a single reference becomes increasingly clear.
– Yao Zhou and colleagues, Institute of Botany, Chinese Academy of Sciences, and international collaborators: “A variant map based on a single linear reference creates reference bias.”
This is where the pangenome becomes important, because instead of relying on one representative genome, it keeps the genes, alleles, non-reference sequences, structural variants, and haplotype paths of multiple genotypes within a shared framework. A graph genome is the advanced computational representation of that same diversity; insertions, deletions, inversions, and different allelic paths are stored as nodes and paths, and reads are no longer restricted to a single fixed route. This distinction matters because the pangenome is the biological and data-level concept of diversity, while the graph genome is the computational tool for mapping, calling, and analyzing that diversity. For native crops, this shift means moving beyond the simple documentation of germplasm toward data-driven design of varieties that are more resistant to real field stresses.
Why Is a Single Reference Genome Not Enough for Climate-Resilient Native Crops?
Native crops and their wild relatives often carry diversity that has remained hidden from breeders through domestication, commercial selection, or the concentration on high-yielding varieties. If a breeding program relies only on a linear reference, it may undervalue or completely miss genes related to defense, flavor, stress tolerance, or regional adaptation that are absent from that reference. The 2019 tomato example showed that analyzing 725 accessions of cultivated tomato and its close relatives led to the identification of 4,873 protein-coding genes in non-reference sequences. This number shows that part of the diversity important for plant breeding lies outside the frame of a single reference genome and, without a pangenome, does not have an effective presence in the final analysis.
The tomato graph pangenome made the second layer of this issue even clearer. In that project, 838 genomes and 32 new reference-level assemblies were used, and more than 19 million variants were recorded. The estimated heritability of molecular traits increased from 0.33 in the linear genome to 0.41 in the graph pangenome, equivalent to a 24 percent increase. In the same study, the linear reference recovered only about 20 percent of the structural variants identified by the graph pangenome, and this gap showed that structural variation is severely underrepresented in a single reference.
– Yao Zhou and colleagues, Institute of Botany, Chinese Academy of Sciences, and collaborators: “This study demonstrates the power of a graph pangenome in understanding the heritability of complex traits and crop improvement.”
For native crops, the value of this technology is not limited to discovering new genes. Its deeper value lies in connecting gene bank resources, field phenotyping data, climate data, and multi-omics data to a breeding pathway that can support real decisions. If a variety shows more stable performance in a dry or hot climate, the pangenome should reveal which variant, which haplotype, and which regulatory network can be tracked in the breeding program. This logic moves plant breeding closer to data-driven design rather than selection based only on general field observation. The practical outcome is a list of breeding targets associated with specific stresses such as drought, heat, salinity, disease, and nutritional quality.
How Does a Graph Pangenome Turn Hidden Variants into Breeding Data?
– From Long Reads to Graph-Aware Mapping
New plant pangenomes increasingly rely on long-read sequencing, chromosome-level assemblies, Hi-C, and graph-aware mapping. Long reads are essential for detecting large insertions, repetitive regions, and structural variants, because many of these segments cannot be accurately reconstructed with short-read data or a linear reference. Hi-C also uses chromosomal contact data to help order and orient contigs, and in the potato example, haplotype assemblies were converted into pseudochromosomes with its support. Graph-aware mapping also connects reads to a multi-path structure and reduces errors around structural variants.
The quality of such a system can be assessed through assembly and variant-calling metrics. In the tomato graph pangenome, the SL5.0 version had a contig N50 of 41.7 megabases, showing an approximately sevenfold increase compared with SL4.0. The average BUSCO score for Solanales genes in the assemblies of the same study was reported as 96.2 percent, a metric that matters for assessing the completeness of single-copy orthologous genes. In simulated 10× data, the graph pangenome achieved an F1 score of 0.966 for SNPs, 0.941 for indels, and 0.840 for SVs, while the linear reference reached 0.931, 0.897, and 0.474, respectively.
– Haplotype as the Unit of Variety Design
Climate-resilient design cannot be completed by identifying a single gene, because many complex traits are controlled by a set of alleles and haplotype paths. In diploid potato, a phased graph pangenome was built from 60 haplotypes, and the length represented in the graph was 3,076 megabases, while the linear reference covered 742 megabases. In the same case, 133,264 structural variants with a threshold of 50 base pairs or more were reported, and 87.0 percent of the structural variants were multiallelic. The identification of 19,625 potentially deleterious structural variants also showed that haplotype design can help reduce undesirable genetic load in predictive breeding.
– Lin Cheng and colleagues, authors of the Nature paper on the potato graph pangenome: “There is a need to move from single reference genomes to phased pangenome references.”
For native crops, this transition means more than improving computational accuracy. If a gene bank only stores a sample but its haplotype path, stress phenotype, and climate data are unclear, breeders cannot quickly incorporate it into a selection program. A phased graph genome makes it possible to describe a germplasm accession not only by name and origin, but also by its allelic paths and the risk of deleterious variants. However, high heterozygosity, multiallelic variants, and the need for accurate phasing make direct application of this model difficult without bioinformatics infrastructure and computational breeding capacity.
Global Pangenome Case Studies for Heat and Drought Tolerance and Habitat Adaptation
Pearl millet is one clear example of the connection between a graph pangenome and heat stress. In this study, 10 new chromosome-level assemblies, along with one existing assembly, were used to build a graph pangenome, and 424,085 structural variants were recorded. The technical outcome of the study pointed to the expansion of the RWP-RK transcription factor family and endoplasmic-reticulum-related genes in heat tolerance. The significance of this example for native crops is that climate resilience is often distributed across both small and large variants, and without a graph pangenome, a major portion of it remains on the margins of analysis.
– Haidong Yan and colleagues, authors of the Nature Genetics paper on pearl millet: “We constructed a graph-based pangenome with ten chromosome-scale genomes and one existing assembly.”
Wheat also shows that a pangenome can read breeding history and habitat adaptation together. A Nature study on 17 Chinese cultivars at chromosome-level resolution identified 249,976 structural variants, of which 122,567, or 49.03 percent, were larger than 5 kilobases. The paper also discussed the association of variants near centromeres with reduced crossover and the role of changes in VRN-A1 in the transition of wheat from spring to winter types. For Iran, this data is important at the level of methodological logic, because directly generalizing it to Iranian wheat without Iranian cultivars, Iranian climate conditions, and local phenotyping does not create a scientifically reliable path.
– Chengzhi Jiao and colleagues, Chinese Academy of Agricultural Sciences, China Agricultural University, and collaborators: “We report chromosome-scale assemblies of 17 wheat cultivars that record the breeding history of China.”
In barley, the scale of the project has moved closer to a population-level stage. The expanded barley pangenome included 76 chromosome-scale sequences generated from long reads and short-read data from 1,315 barley genomes. This structure shows that a pangenome is not merely a collection of a few high-quality assemblies; it can integrate wild and domesticated genomes within one framework to discover variants related to adaptation and breeding. For countries where native crops and their wild relatives are important genetic assets, this model shows a path for turning a gene bank into an active discovery and selection system.
In maize, a Nature Genetics study using 25 new assemblies from germplasms with major differences in drought resistance, along with 31 additional genomes, linked three genes—ZmUGE2, ZmSIL2, and ZmASI3—to drought resistance at different stages of growth. The importance of this case lies in connecting the pangenome with multi-omics data, because the paper referred to RNA-seq, ChIP-seq, and transgenic data in NCBI and Zenodo. This example shows that an applied pangenome is not just a genome file; it is a network of genome, gene expression, regulation, and phenotype. Therefore, gene discovery without multi-environment phenotyping and regulatory data is not enough to design a climate-resilient variety.
Gene Bank Standards and Data Reproducibility in a National Pangenome Infrastructure
Pangenome infrastructure begins with the genetic sample, which is why the gene bank sits at its center. FAO standards for plant gene banks cover areas such as orthodox seed conservation, field collections, in vitro culture, and cryopreservation, and they are voluntary in nature. The 2022 practical version of these standards also provides step-by-step guidance for managing orthodox seeds, field collections, and in vitro culture. Connecting such standards to a pangenome project means that before each sample enters genomic analysis, its origin, conservation status, metadata, and traceability must be reliable.
– Food and Agriculture Organization of the United Nations: “Well-managed gene banks conserve genetic diversity and make it available to breeders.”
Scientific publication of a pangenome is also incomplete without data accessibility and reproducibility. In the 2025 maize study, raw data were released in NCBI SRA under BioProject PRJNA1103102, and the assembly, annotation, and structural variants were placed in Zenodo. The custom codes from the same study were also released on GitHub and Zenodo, showing that data standards in pangenomics are not limited to sequence files. For national programs, this model means that every project must be designed from the beginning around a data repository, code versioning, phenotypic metadata, and the possibility of reanalysis by independent groups.
The appropriate funding model for building an open version of the pangenome of strategic crops in the global record is closer to a public academic consortium. Tomato, wheat, barley, and maize projects have moved forward through the participation of multiple universities or institutes and the release of data in open repositories, although their budget figures have not been quantified in usable data. This characteristic matters for technology investment, because the initial output of a pangenome is not necessarily a ready-to-sell variety. It is validated data, markers, haplotype paths, and prioritized germplasm. Therefore, the return on value in such a project does not come through the short-term sale of a product; it emerges through the creation of decision infrastructure for plant breeding.
Iran’s Path for Turning Native Germplasm into an Applied Pangenome Program
For Iran, the clear international starting point is the institutional identification of the National Plant Gene Bank of Iran in WIEWS and Genesys. The National Plant Gene Bank of Iran is registered under the code IRN029, and its type is reported as governmental. The Seed and Plant Improvement Research Institute is also registered in Genesys under the code IRN001, but the number of accessions recorded for it is shown as zero. This zero should not be interpreted as the absence of domestic resources; the more precise meaning is that, at the Genesys level, publicly searchable records for those institutions are not available.
This situation clarifies Iran’s implementation priority. Before making any major claim about the pangenome of native crops, the chain of sample, metadata, phenotype, climate, and sequencing must be connected. If a germplasm record cannot be retrieved on a public platform, if stress phenotypes have not been recorded across multiple environments, or if the data path from the gene bank to the breeder is not traceable, the pangenome project becomes a catalog of variants. Applied design should begin with crops that are important for food security and for which multi-environment phenotyping, standardized sampling, and bioinformatics analysis are feasible.
The main risk for Iran is the rushed generalization of findings from tomato, wheat, millet, barley, maize, or potato to domestic varieties. Each global case study provides a methodological model, but a beneficial allele in one population and climate does not necessarily have the same effect in another native population without genotype-by-environment testing. For this reason, a national pangenome must move forward with climate phenotyping, multi-omics data, and local breeding trials. The right path is to build the data infrastructure and the breeding decision pathway at the same time, not to produce massive amounts of sequence data without a clear agronomic question.
In implementation design, the first output should not be assumed to be an immediate commercial variety. Instead, it should be a usable data package for breeders. This package can include a stable sample identifier, gene bank conservation status, assembly quality, haplotype paths, structural variants, multi-environment phenotypic data, and the connection of those data to dominant stresses. Such a package is consistent with the logic of global projects, because the value of a graph pangenome becomes visible when it moves beyond variant discovery and reaches breeding markers, genomic selection, and germplasm prioritization. For Iran, this design brings the capacity of the gene bank closer to an active plant-breeding infrastructure rather than raw genetic storage.
The practical conclusion for Iran is that graph genomes and pangenomes should be viewed as decision infrastructure in plant breeding, not merely as expensive genomics projects. A single reference is enough to enter the analysis, but it is not enough to discover hidden diversity, structural variants, non-reference genes, and climate-resilient haplotype paths. A successful program must consist of a standardized gene bank, open or semi-open data, high-quality assembly, graph-aware mapping, multi-environment phenotyping, and computational reproducibility. Only within such a structure can native resources move beyond raw genetic assets and become tools for designing varieties for food security and climate adaptation.