Skip to content

Menu

Search

🔍

Indo-Aryan qpAdm Admixture

Updated August 2026 · AADR v66

Introduction

This study presents a detailed analysis of Steppe admixture in Indo-Aryan populations, with a particular focus on the Jatt Sikh population in Punjab, across three Bronze Age windows: the Early Bronze Age (3300-2600 BCE), Middle Bronze Age (2900-2350 BCE), and Late Bronze Age (1900-1200 BCE). Utilizing both personal genetic data and the Allen Ancient DNA Resource, we examine the Steppe, Iranian farmer-related, and Ancient Ancestral South Indian (AASI) components as each window supplies the Steppe source in turn. Our findings show a Steppe share of 33-38%, an Iranian farmer-related share of 36-39%, and an AASI share of 25-29%. The qpAdm fits alone cannot distinguish among Steppe cultures, though combined with Y-chromosome and chronological evidence they make the Sintashta–Andronovo horizon the best-supported source; a Bactria-Margiana Archaeological Complex (BMAC) contribution cannot be resolved as a separate stream.

Tools Used

The analysis rests on the following tools and datasets:

  • ADMIXTOOLS 2 (R package; Maier et al., 2023) — qpAdm model fitting, f-statistics, and cached f2 blocks. docs
  • Allen Ancient DNA Resource — ancient and present-day genotypes. Fitted on v54.1.p1 1240k, re-fitted on v66. AADR
  • 23andMe v5 chip — target genotypes, over 500,000 SNPs.
  • AncestryDNA — second target genotype set, over 500,000 SNPs, used to cross-check the first.
  • Big Y-700 — Y-DNA haplogroup: R1a-Z93 > R-L657 > R-FTF40903.
  • ADAMIXTURE 1.7.5 — unsupervised ancestry clustering, K = 5 to 12, on AADR v66. repo
  • f4-statistics — tests of what each source population itself carries.
  • IllustrativeDNA G25 — independent admixture check. results

Key Findings

Across all three Bronze Age windows, the genome resolves into the same three ancestry streams:

Ancestry components of a Jatt Sikh genome across the Bronze Age windowsGrouped bar chart of three qpAdm models on AADR v66, sharing one Iranian farmer-related source, Iran_ShahTepe_BA (3235-3150 BCE), old enough to stand in all three windows, and one deep South Asian source, the Onge (ONG), the Andamanese hunter-gatherer-related (AHG) population standing for the unsampled AASI lineage. Error bars are one standard error. Early Bronze Age with Russia_Samara_EBA_Yamnaya (3215-2568 BCE): Steppe 35.7% plus or minus 3.8, Iranian farmer-related 39.2% plus or minus 3.8, AHG-related 25.2% plus or minus 1.5, p = 0.414. Middle Bronze Age with Poland_CordedWare (2752-2287 BCE): Steppe 36.4% plus or minus 3.8, Iranian farmer-related 35.9% plus or minus 4.0, AHG-related 27.7% plus or minus 1.5, p = 0.628. Late Bronze Age with Russia_Samara_LBA_Srubnaya (1850-1525 BCE): Steppe 33.7% plus or minus 3.6, Iranian farmer-related 37.4% plus or minus 4.0, AHG-related 28.9% plus or minus 1.5, p = 0.273. The Steppe share does not change significantly across the three windows: the Early to Late Bronze Age difference is 2.0 plus or minus 5.3 percentage points, 0.37 standard errors. Adding an explicit European farmer term as a fourth source fails: the term is unresolvable in the Early Bronze Age (Z = 0.61) and negative in the Middle and Late Bronze Age (-4.2% and -10.1%).0102030405060Ancestry proportion (%)Steppe ancestryIranian farmer-related ancestryAASI (via Onge)35.7%Russia_Samara_EBA_Yamnaya39.2%Iran_ShahTepe_BA25.2%ONG (Onge)36.4%Poland_CordedWare35.9%Iran_ShahTepe_BA27.7%ONG (Onge)33.7%Russia_Samara_LBA_Srubnaya37.4%Iran_ShahTepe_BA28.9%ONG (Onge)Early Bronze Age(3300-2600 BCE)p = 0.414Middle Bronze Age(2900-2350 BCE)p = 0.628Late Bronze Age(1900-1200 BCE)p = 0.273Time period
Ancestry components of a Jatt Sikh genome across the Bronze Age windows
WindowSteppe sourcepSteppe (%)Iranian farmer-related (%)AASI (%)
Early Bronze AgeRussia_Samara_EBA_Yamnaya0.41435.7 ± 3.839.2 ± 3.825.2 ± 1.5
Middle Bronze AgePoland_CordedWare0.62836.4 ± 3.835.9 ± 4.027.7 ± 1.5
Late Bronze AgeRussia_Samara_LBA_Srubnaya0.27333.7 ± 3.637.4 ± 4.028.9 ± 1.5
  • The Steppe ancestry proportion is stable across all three time windows. Whether the Steppe source is drawn from the Early Bronze Age (Yamnaya), the Middle Bronze Age (Corded Ware), or the Late Bronze Age (Srubnaya), the representative estimate stays between 33.7% and 36.4% (33.3–38.1% across all ten sources tested). The Early-to-Late difference is 2.0 ± 5.3 percentage points (0.37 standard errors), which is not statistically significant. This constancy is partly a property of the sources themselves — all ten descend from one Steppe gene pool, so any member fits interchangeably — which means it bounds the combined Steppe contribution tightly but cannot, on its own, count arrival pulses. What it establishes is a stable Steppe fraction drawn from an already-formed Steppe gene pool, however many crossings delivered it.
  • Paternal lineage and chronology, with consistent autosomal fits, make a Sintashta–Andronovo (Steppe_MLBA) source the best-supported hypothesis. First, the autosomal fits are fully consistent with Steppe_MLBA sources, the class published work identifies for South Asia (Narasimhan et al., 2019), though on autosomal similarity alone they also admit earlier Steppe populations; the following two constraints do the narrowing. Second, the paternal lineage documented here, R1a-Z93 > L657, belongs to the Asian branch of the R1a-M417 expansion. Its European sister branch (Z282) accompanied Corded Ware into Europe, a division established by genotyping over 16,000 Eurasian samples (Underhill et al., 2015), and the South Asian subclades, including L657, date to the Bronze Age (Silva et al., 2017). Third, the horizon's southward expansion through Central Asia (c. 2100–1500 BCE) is the only Steppe movement that reaches the Punjab in the required window. Earlier Yamnaya-like populations such as Afanasievo also pass these fits, because a qpAdm p-value tests whether a model is plausible rather than which source is the true one (Harney et al., 2021); they fail all three constraints above. The data cannot yet resolve which branch within the horizon was involved.
  • Three ancestry sources are sufficient to model the genome. Steppe, Iranian farmer-related, and AASI ancestry together account for the genome in all three windows (p = 0.27–0.63). Adding a fourth source does not help. An explicit European farmer term is unresolvable in the Early Bronze Age (Z = 0.61) and negative in the Middle and Late Bronze Age (−4.2% and −10.1%). That ancestry is already inside the Steppe source: the Steppe_MLBA populations formed as mixtures of Yamnaya-related and European farmer ancestry (Haak et al., 2015; Allentoft et al., 2015).
  • No distinct BMAC (Oxus) contribution is detected. Bronze Age Oxus samples (Gonur, Sapalli-tepe, Ulug-depe) carry the same Central Asian farmer-related ancestry as the other farmer sources tested. Swapping them into the farmer slot changes nothing of substance (detailed under Investigating BMAC Influence). The genome is fully accounted for without any BMAC-specific term. Narasimhan et al. (2019) reached the same conclusion for South Asia generally: the main BMAC population contributed little ancestry, and the Steppe stream crossed the region without acquiring it.
  • Unsupervised clustering does not recover a discrete AASI component. In ADAMIXTURE runs at K = 5 through 12, the deep South Asian ancestry never forms its own cluster. It surfaces instead as composite components anchored by the nearest deep-lineage references in the panel, mainly Papuan, Andamanese, and Upper Palaeolithic East Asian genomes, together about 15% of the genome. This is expected rather than anomalous: no ancient AASI genome has been sequenced, so the lineage is reconstructed, not sampled (Narasimhan et al., 2019). The qpAdm models above bridge the gap by using the Onge of the Andaman Islands as the AASI source. The Onge are the closest available population but a long-diverged relative, not AASI itself, so the 25–29% AASI estimate is measured through that proxy rather than against the lineage directly. The Onge are nonetheless the right instrument for the job: they carry no detectable West Eurasian ancestry and were identified as the closest sampled relatives of the deep South Asian component (Reich et al., 2009), which is what makes the farmer and AASI streams separable at all. The deep-branch symmetry tests below refine that picture — the Onge behave as a deeply diverged sibling of AASI rather than a strict clade member — one more reason the AASI figure is reported as proxy-measured.
  • Large calculator "EEF" components and a near-zero qpAdm EEF stream are both correct. The clustering's Anatolian-farmer-related component is large (22–27% across K), yet the qpAdm models require no separate European farmer stream (0–4%, unresolved). Both are correct. Calculator-style components measure total Anatolian-related affinity, which arrives folded inside the Iranian farmer and Steppe streams; qpAdm measures distinct ancestry streams, and no independent European farmer migration is needed to explain the genome.

Model-Free Clustering (ADAMIXTURE)

The qpAdm models in this analysis all specify their sources. As a model-free cross-check, ADAMIXTURE clustering was run on AADR v66 at K = 5 to 12, letting the components emerge from the reference panel with no source assumptions. ADAMIXTURE represents each genome as a mixture of K components learned from the data; the components are statistical constructs of the panel, not ancient populations, and reading them well means watching how they behave as K changes rather than fixating on any single K.

The sweep tells a story in itself. At K = 5 the Iranian farmer and Steppe streams fuse into a single component carrying 48% of the genome. K = 8 splits them apart, and briefly resolves a fused EEF-plus-Steppe component (22.8%) whose shape is exactly the Corded Ware fusion documented in the Results: the clustering rediscovers, on its own, the population structure the supervised models specify. By K = 10 and 12 the streams under discussion are cleanly separated, at a cost: a formal prediction criterion prefers K = 5, and the growing unassigned slivers at high K are the over-splitting it warns about. Four panels spanning the sweep (K = 5, 8, 10, 12) are shown for that reason; no single K is privileged.

ADAMIXTURE component shares across KStacked horizontal bars of the target's unsupervised ADAMIXTURE component shares at K equals 5, 8, 10, and 12, on AADR v66. At K=5 a combined Iranian-farmer-plus-Steppe composite carries 48.0 percent, EEF-related 27.3, a deep South Eurasian composite anchored by Papuan and Tianyuan 11.1, Ancient North Eurasian related 8.3, and a deep African plus Andamanese composite 5.3. At higher K the composite splits: Western Steppe Herder stabilises at 25.7 to 29.3 percent, Iranian farmer-related at 17.2 to 20.2, EEF-related affinity at 22.5 to 27.3, while the deep components stay near 15 percent combined and no discrete AASI cluster ever forms. The steppe share in unsupervised clustering, about 26 to 29 percent, brackets the qpAdm estimate of about 34 percent from below, and the large EEF-related component measures total Anatolian-farmer affinity carried inside the farmer and Steppe streams, not a separate ancestry stream.K=5K=8K=10K=1248.0%27.3%11.1%8.3%28.8%22.8%20.2%10.6%7.7%25.7%17.2%22.5%10.7%7.8%29.3%19.1%23.5%10.7%7.8%Steppe (WSH)Iranian/BMAC farmerEuropean farmer (EEF) affinityIranian + Steppe compositeEEF + Steppe compositeDeep South Eurasian compositeDeep African + AndamaneseANE / Native American compositeunassigned
Model-free component shares at K = 5, 8, 10, and 12 (AADR v66)
Ancestry stream (rolled up)K=5K=8K=10K=12
Western Steppe Herder28.825.729.3
Iranian-Neolithic / BMAC farmer20.217.219.1
European farmer (EEF) affinity27.32.822.523.5
Iranian + Steppe composite48.06.0
EEF + Steppe composite22.8
Deep South Eurasian composite (Papuan, Tianyuan, Ami anchors)11.110.610.710.7
Deep African + Andamanese composite5.34.24.24.2
ANE / Native American composite8.37.77.87.8
unassigned2.76.05.4

First, the Steppe component stabilises at 26–29% once K separates it from the farmer streams, bracketing the qpAdm estimate from below; the two methods agree that roughly a third of this genome traces to the Steppe. Second, the EEF-related component is large (22–27%) at every K, once the K = 8 EEF-plus-Steppe composite is counted with it: this is the total Anatolian-farmer affinity discussed under Key Findings, carried inside the farmer and Steppe streams rather than arriving as its own migration. Third, and most diagnostic, no discrete AASI cluster ever forms at any K. The deep ancestry surfaces only as composites anchored by Papuan, Tianyuan, and Andamanese references, which is exactly what an unsampled lineage from a near-simultaneous three-way split of the deep eastern lineages into AASI, Andamanese, and East Asian branches (the eastern trifurcation) should look like when forced through a reference panel that does not contain it.

The trifurcation reading was tested directly rather than left as interpretation. In f4 tests of the form f4(Sohi, Mbuti; Andamanese, X), where Sohi is the label for the target genome, the genome is statistically symmetric between the Andamanese branch and the East Asian branches (Ami, Dai, and the 40,000-year-old Tianyuan genome; |Z| ≤ 0.5 in every comparison), with only the pre-split Ust-Ishim genome showing the weak asymmetry its earlier divergence predicts (Z = 2.1). The deep component is at the resolution floor of current references: no available population can subdivide it further.

Population context comes from the same toolkit. Compared against 87 unrelated Punjabi genomes from the 1000 Genomes Project with f4 tests of the form f4(Mbuti, Sintashta; Sohi, Punjabi), the target shares more drift with Sintashta than every one of the 87, significantly so for 77 of them (median Z = −5.2), while no Punjabi sample is significantly more Steppe-shifted than the target. A Jatt genome sitting at the high-Steppe end of the Punjabi distribution is what the community's agro-pastoralist history predicts, and it is the same placement the qpAdm proportions give in absolute terms.

Methodology

The sections above present the findings; this one documents how they were produced, in enough detail to re-run them.

Data

The target is a Jatt Sikh genome genotyped independently on two consumer panels, 23andMe v5 and AncestryDNA, each with over 500,000 SNPs. The two panels serve as mutual checks. qpAdm compares the target against ancient sources over the intersection of SNPs present in both the query genotype and the Allen Ancient DNA Resource 1240k panel. Because the Bronze Age ancient samples are the limiting factor, the informative overlap is nearly identical for the two chips, and their ancestry estimates converge to the same values with slightly different standard errors. The current models were fitted on the merged target against AADR version 66; raw outputs are in the Supplementary Data.

Model fitting

Admixture models were fitted with qpAdm in ADMIXTOOLS 2 (Maier et al., 2023), run in genotype mode with allsnps = TRUE: each f-statistic uses the SNPs available for its own populations, the appropriate regime when a single array-genotyped target meets ancient capture data of heterogeneous coverage. The regime is stamped in every raw model receipt in the Supplementary Data. qpAdm models the target as a mixture of a small number of source populations (the "left" set) and evaluates the model against a set of reference populations (the "right" set) chosen to be differentially related to the sources. The right set does the discriminating work, so its composition matters more than its size.

The reference set used here contains twelve populations, one anchor per ancestral stream the sources draw from: deep African outgroups (Mbuti, Mota), Upper Palaeolithic Eurasians (Ust-Ishim, Tianyuan, Kostenki), Papuan for the deep eastern lineage, Karitiana for Ancient North Eurasian ancestry, an East Asian Neolithic anchor (Mongolia_N), Caucasus and Western hunter-gatherers (Kotias, Iron Gates), and two farmer anchors (Levant PPNB, Anatolian Neolithic). This follows the design guidance of Harney et al. (2021): no population directly ancestral to a source, and the AASI proxy (Onge) kept out of the reference set it is judged against. The set was validated before use: it is stable on the accepted model and correctly rejects deliberately wrong sources, including an East-Asian-shifted steppe population and an EEF-bearing one.

Source selection

Each Bronze Age window supplies its own Steppe source: five Yamnaya and Afanasievo populations for the Early Bronze Age, three Corded Ware populations for the Middle, and Sintashta and Srubnaya for the Late. The figure shows one representative per window; the full grid of 70 models (7 farmer candidates × 10 Steppe sources, distributed 5/3/2 across the windows) is in the Supplementary Data.

The Iranian farmer-related source is held fixed across all three windows at Iran_ShahTepe_BA (Gorgan plain, 3235-3150 BCE), so that the only variable in the comparison is the Steppe era. Three criteria led to this choice. First, chronology: the farmer source must predate the earliest window, and every Bactria-Margiana era candidate (2300-1000 BCE) is younger than the Early Bronze Age window itself. Second, feasibility: the one other old-enough candidate, Chalcolithic Anau (Turkmenistan_C), fails every Early Bronze Age combination (p = 0.011-0.037), while ShahTepe is feasible in all ten window-by-Steppe combinations. Third, data quality: ShahTepe is the best-covered candidate in the set (7 individuals at 2.14x median coverage; the Gonur group has 34 individuals at 0.01x).

The deep South Asian source is the Onge (ONG). As discussed under Key Findings, the Onge stand in for the unsampled AASI lineage; they carry no detectable West Eurasian ancestry, which is what allows the farmer and AASI streams to be separated.

Model evaluation

A qpAdm p-value tests whether the proposed model is consistent with the data given the reference set. A model is retained when p exceeds 0.05 and every ancestry coefficient is positive and resolved; a high p-value does not prove the model, and p-values do not rank passing models against one another (Harney et al., 2021). Models were additionally stress-tested by adding fourth sources (European farmer, Bactria-Margiana populations) to check whether three streams suffice, and cross-checked against the model-free ADAMIXTURE clustering shown above, which requires no source assumptions.

Finally, because the target is a consumer genotyping array co-analysed with capture and shotgun ancient data, the models were checked for sequencing-platform bias using the Compatibility SNP panel of Fournier, Fulton, and Reich (2026), which restricts analysis to positions with minimal technology-specific bias. Re-fitting the three figure models on the 246,529 SNPs shared with the panel (59% of the full set), under matched extraction settings for both arms, moves no ancestry estimate by more than 2.0 percentage points, within one standard error in every case, while standard errors widen by about 10% as expected from the reduced SNP count. The reported proportions are therefore not artifacts of mixing genotyping platforms.

Results

The three subsections below work through the Bronze Age windows in order. Each follows the same shape: the historical setting, the fitted models for that window, the four-source stress test, and what the window adds to the overall picture.

Early Bronze Age (3300-2600 BCE): The Yamnaya Influence

The Yamnaya culture emerged on the Pontic-Caspian steppe in the Early Bronze Age. Known for their wheeled vehicles, advanced metallurgy, and distinctive kurgan burial practices, the Yamnaya people reshaped Eurasian genetics and culture. Recent genomic research shows that Yamnaya populations drew approximately four fifths of their ancestry from the Caucasus-Lower Volga (CLV) cline and the remainder from Ukrainian Neolithic hunter-gatherers (UNHG) (Lazaridis et al., 2025). The CLV cline represents a gradient of admixture between Caucasus hunter-gatherer (CHG) ancestry and steppe populations, while the UNHG component reflects interactions with local European groups. From this steppe base the Yamnaya spread widely across Eurasia (Haak et al., 2015).

Five Early Bronze Age Steppe sources were tested against the fixed farmer source (Iran_ShahTepe_BA) and the Onge. All five models fit, with closely similar proportions (AADR v66):

Early Bronze Age three-source model with YamnayaBar chart of the representative Early Bronze Age qpAdm model on AADR v66, p = 0.414. Steppe ancestry from Russia_Samara_EBA_Yamnaya: 35.7 percent plus or minus 3.8. Iranian farmer-related ancestry from Iran_ShahTepe_BA: 39.2 percent plus or minus 3.8. AASI ancestry via the Onge: 25.2 percent plus or minus 1.5. Error bars are one standard error.01020304050Ancestry proportion (%)35.7%39.2%25.2%SteppeIranian farmer-relatedAASI (via Onge)Russia_Samara_EBA_YamnayaIran_ShahTepe_BAONG (Onge)
Early Bronze Age three-source model (p = 0.414)
Steppe sourcepSteppe (%)Iranian farmer-related (%)AASI (%)
Russia_Samara_EBA_Yamnaya (shown in figure)0.41435.7 ± 3.839.2 ± 3.825.2 ± 1.5
Russia_Khakassia_Afanasievo0.63337.836.925.3
Russia_Altai_Afanasievo0.51038.136.525.3
Russia_Kalmykia_EBA_Yamnaya0.37636.637.925.5
Russia_Orenburg_EBA_Yamnaya0.52836.238.525.3

A four-source model adding an explicit European farmer term (Austria_N_LBK, n = 103) was also tested for this window. With ShahTepe as the farmer source the added coefficient is 2.1 ± 3.4 percent (Z = 0.61), indistinguishable from zero. The behaviour of this term across model variants is informative in both directions. With European-farmer-free farmer sources (Ganj Dareh, Anau, Fergana) alongside Yamnaya, the term stays at 0–4 percent with |Z| ≤ 1.3: no separate European farmer stream is required. With classic Steppe_MLBA sources such as Sintashta or Alakul in the Steppe slot, the term goes negative (−4 to −12 percent depending on the farmer): those populations carry more European farmer ancestry than the target's Steppe stream did. Only one configuration produces a positive term (6–7 percent, with the AASI-rich Indus-Periphery-like farmer Iran_ShahriSokhta_BA2, whose own composition redistributes the coefficients). Read together, the tests bound the European farmer content of the incoming Steppe stream: present, but below the level of the sampled classic Sintashta–Andronovo populations, pointing to a vehicle at the horizon's earlier, less-admixed edge.

Early Bronze Age four-source model with a European farmer termBar chart of the Early Bronze Age four-source qpAdm model on AADR v66, p = 0.963. Steppe ancestry from Russia_Samara_EBA_Yamnaya: 37.5 percent plus or minus 3.9. Iranian farmer-related ancestry from Iran_ShahTepe_BA: 33.6 percent plus or minus 4.9. AASI ancestry via the Onge: 26.9 percent plus or minus 1.8. European farmer ancestry from Austria_N_LBK: 2.1 percent plus or minus 3.4, Z = 0.61, with an error bar crossing zero: the coefficient is unresolvable, so a distinct European farmer stream is not supported.01020304050Ancestry proportion (%)37.5%33.6%26.9%2.1% ± 3.4Z = 0.61, unresolvableSteppeIranian farmer-relatedAASI (via Onge)European farmer (EEF)Russia_Samara_EBA_YamnayaIran_ShahTepe_BAONG (Onge)Austria_N_LBK
Early Bronze Age four-source model with ShahTepe as the farmer source (p = 0.963): the European farmer coefficient is indistinguishable from zero
Ancestry componentSourceEstimate (%)Z
SteppeRussia_Samara_EBA_Yamnaya37.5 ± 3.99.6
Iranian farmer-relatedIran_ShahTepe_BA33.6 ± 4.96.8
AASIONG (Onge)26.9 ± 1.815.2
European farmer (EEF)Austria_N_LBK2.1 ± 3.40.61

Standard errors are shown for the representative models; the full grid output is in the Supplementary Data.

Observations for this window:

  • The Steppe share is 36-38% under every Early Bronze Age source, including Afanasievo. This uniformity reflects the coherence of the early Steppe gene pool rather than identifying any of these populations as the migrating group; the transmitting culture is constrained separately (see Key Findings). The scale of the component supports a substantial demographic contribution from Steppe-related populations of the kind proposed by Anthony (2007).
  • The Iranian farmer-related component is the largest in this window (36-39%). It reflects genetic connections between South Asia and the Iranian plateau region that predate the Indo-Aryan migrations (Broushaki et al., 2016).
  • AASI ancestry is stable near 25% in every model (Basu et al., 2016).

Middle Bronze Age (2900-2350 BCE): Corded Ware

The Early Bronze Age sources above are chronological stand-ins; the Middle Bronze Age supplies the first admixed European Steppe populations. The Corded Ware culture represents a fusion of Steppe ancestry with European Neolithic farmer populations (Haak et al., 2015). Our qpAdm analysis for this window tests three Corded Ware populations (Poland, Czechia, and Esperstedt in Germany) as the Steppe source.

The Corded Ware populations themselves are admixed. Before modelling the target, we resolved Czech Corded Ware as a two-source mixture: Yamnaya-related Steppe ancestry plus European farmer-related ancestry represented by Globular Amphora (Allentoft et al., 2015). This matters for everything that follows: from the Middle Bronze Age onward, the Steppe source itself carries European farmer ancestry folded inside it.

Corded Ware as a two-source mixtureBar chart of a two-source qpAdm model of Czech Corded Ware. Yamnaya-related Steppe ancestry from Russia_Samara_EBA_Yamnaya: 71.1 percent plus or minus 1.5. European farmer-related ancestry from Ukraine_EBA_GlobularAmphora: 28.9 percent plus or minus 1.5. The Corded Ware culture is therefore itself a fusion of Steppe and European farmer ancestry, which is why Steppe_MLBA sources deliver European farmer ancestry without a separate EEF stream.020406080Ancestry proportion (%)71.1%28.9%Yamnaya-related SteppeEuropean farmer-related (EEF)Russia_Samara_EBA_YamnayaUkraine_EBA_GlobularAmphora
Czech Corded Ware resolved as Yamnaya plus European farmer ancestry
Ancestry componentSourceEstimate (%)
Yamnaya-related SteppeRussia_Samara_EBA_Yamnaya71.1 ± 1.5
European farmer-related (EEF)Ukraine_EBA_GlobularAmphora28.9 ± 1.5

This two-source structure is well established in the ancient-DNA literature. Corded Ware communities formed on the North European Plain around 2900 BCE as Yamnaya-related migrants mixed with local Neolithic farming populations, and genomic studies consistently recover roughly a quarter to a third European farmer ancestry in them (Haak et al., 2015; Papac et al., 2021). Our two-source fit above reproduces that published range from an independent direction, which is a useful check on the reference setup before the target is modelled.

With the decomposition in hand, the target was modelled with each of three Corded Ware populations in the Steppe slot. All three fit, with closely similar proportions; the figure shows the Poland_CordedWare model.

Middle Bronze Age three-source model with Corded WareBar chart of the representative Middle Bronze Age qpAdm model on AADR v66, p = 0.628. Steppe ancestry from Poland_CordedWare: 36.4 percent plus or minus 3.8. Iranian farmer-related ancestry from Iran_ShahTepe_BA: 35.9 percent plus or minus 4.0. AASI ancestry via the Onge: 27.7 percent plus or minus 1.5. Error bars are one standard error.01020304050Ancestry proportion (%)36.4%35.9%27.7%SteppeIranian farmer-relatedAASI (via Onge)Poland_CordedWareIran_ShahTepe_BAONG (Onge)
Middle Bronze Age three-source model (p = 0.628)
Steppe sourcepSteppe (%)Iranian farmer-related (%)AASI (%)
Poland_CordedWare (shown in figure)0.62836.4 ± 3.835.9 ± 4.027.7 ± 1.5
Germany_Esperstedt_CordedWare0.66536.035.528.5
Czechia_EBA_CordedWare0.32035.836.228.0

In the four-source test for this window, the explicit European farmer term goes negative (−4.2 ± 3.7 percent) and the model is infeasible. This is the expected consequence of the Corded Ware decomposition above: the Steppe source now carries European farmer ancestry of its own, so adding a separate term double-counts it.

Observations for this window:

  • The proportions are statistically unchanged from the Early Bronze Age window (Steppe 36%, farmer 36%, AASI 28%), even though the Steppe source is now an admixed European population. The models continue to see one Steppe signal.
  • The European farmer ancestry inside Corded Ware does not inflate any component. With the farmer slot held at ShahTepe across all windows, the accounting stays clean; earlier versions of this analysis, which used an AASI-rich farmer source, showed apparent shifts between windows that were artifacts of the source composition, not of the target's history.

Late Bronze Age (1900-1200 BCE): Andronovo Culture

This window corresponds to the Sintashta–Andronovo horizon itself: the period in which Steppe_MLBA populations moved south through Central Asia, appearing in the Turan corridor between roughly 1800 and 1500 BCE; in South Asia itself, directly dated Steppe ancestry first appears in the Swat valley by 1200–800 BCE (Narasimhan et al., 2019). Unlike the earlier windows, the Steppe sources tested here are contemporaries of the migration, not stand-ins from an earlier era.

Two well-sampled western members of this world serve as the Steppe sources: Sintashta (Chelyabinsk, the fortified-settlement culture at the horizon's root) and Srubnaya (Samara, its western sibling). Both are, like Corded Ware before them, fusions of Yamnaya-related and European farmer ancestry (Allentoft et al., 2015); the classic Andronovo branches of Kazakhstan (Alakul, Fedorovo) are their close eastern relatives. The figure shows the Srubnaya model; both sources fit with closely similar proportions.

Late Bronze Age three-source model with SrubnayaBar chart of the representative Late Bronze Age qpAdm model on AADR v66, p = 0.273. Steppe ancestry from Russia_Samara_LBA_Srubnaya: 33.7 percent plus or minus 3.6. Iranian farmer-related ancestry from Iran_ShahTepe_BA: 37.4 percent plus or minus 4.0. AASI ancestry via the Onge: 28.9 percent plus or minus 1.5. Error bars are one standard error.01020304050Ancestry proportion (%)33.7%37.4%28.9%SteppeIranian farmer-relatedAASI (via Onge)Russia_Samara_LBA_SrubnayaIran_ShahTepe_BAONG (Onge)
Late Bronze Age three-source model (p = 0.273)
Steppe sourcepSteppe (%)Iranian farmer-related (%)AASI (%)
Russia_Samara_LBA_Srubnaya (shown in figure)0.27333.7 ± 3.637.4 ± 4.028.9 ± 1.5
Russia_Chelyabinsk_MLBA_Sintashta0.24533.338.228.6

The four-source test is most decisive in this window: the explicit European farmer term reaches −10.1 ± 4.0 percent. Late Bronze Age Steppe populations such as Srubnaya and Sintashta carry substantial European farmer ancestry of their own (Wang et al., 2019), more than the target's Steppe stream requires; forcing an additional term overcorrects. This is the oversupply pattern discussed under Key Findings: the incoming Steppe stream carried European farmer ancestry, but less of it than the sampled classic Sintashta–Andronovo populations.

Observations for this window:

  • The proportions complete the stability pattern across all three windows (Steppe 33.7%, farmer 37.4%, AASI 28.9%). The chronological alignment is the meaningful part: Srubnaya and Sintashta (1900–1200 BCE) are the only tested Steppe sources that are contemporaries of the migration window itself (Narasimhan et al., 2019), coinciding with the period proposed for Indo-European language spread into South Asia (Anthony, 2007).
  • AASI ancestry remains near 29%, continuous with the indigenous substrate documented from the Indus period onward (Shinde et al., 2019).

Investigating BMAC Influence

The Bactria-Margiana Archaeological Complex (BMAC), or Oxus civilization, sat directly on the route the Steppe stream travelled into South Asia, and its urban centres interacted with both Steppe and South Asian populations. Whether it contributed ancestry to Indo-Aryan groups is therefore a natural question, and one this analysis tests directly rather than inherits.

The test mirrors the European farmer design used in the Results: each window's accepted three-source model is extended with a BMAC population as an explicit fourth source, against the same twelve-population reference set. Three BMAC representatives span the available samples: Gonur, the capital (Turkmenistan_BA1-1, 34 individuals at 0.01x coverage), Dzharkutan, a late steppe-admixed site (Uzbekistan_BA1-1, 9 individuals at 2.11x), and the Bustan/Sapalli group (Uzbekistan_BA, 31 individuals). If BMAC contributed a distinct stream, its coefficient should resolve positive; nine models (three windows, three sources) put that to the test on AADR v66.

Early Bronze Age four-source model with a BMAC termBar chart of the Early Bronze Age four-source qpAdm model adding Gonur (Turkmenistan_BA1-1) as a BMAC source, p = 0.321. Steppe from Russia_Samara_EBA_Yamnaya: 35.8 percent plus or minus 3.9, resolved. AASI via the Onge: 24.8 percent plus or minus 2.5, resolved. Iranian farmer-related from Iran_ShahTepe_BA: 33.0 percent plus or minus 42.0, unresolved. BMAC from Gonur: 6.4 percent plus or minus 43.0, Z = 0.15, unresolved. The farmer and BMAC error bars each span roughly 85 percentage points and both cross zero: the model cannot separate BMAC ancestry from Iranian farmer-related ancestry, while the Steppe and AASI estimates are unaffected.0-2020406080Ancestry proportion (%)35.8%33.0% ± 4224.8%6.4% ± 43SteppeIranian farmer-relatedAASI (via Onge)BMACRussia_Samara_EBA_YamnayaIran_ShahTepe_BAONG (Onge)Turkmenistan_BA1-1 (Gonur)
Early Bronze Age four-source model with a BMAC term (p = 0.321): the farmer and BMAC coefficients cannot be separated
WindowBMAC sourcepBMAC estimate (%)ZFeasible
Early Bronze AgeGonur0.3216.4 ± 43.00.15yes
Early Bronze AgeDzharkutan0.33414.6 ± 42.90.34yes
Early Bronze AgeSapalli0.40430.6 ± 39.50.77yes
Middle Bronze AgeGonur0.57318.2 ± 34.50.53yes
Middle Bronze AgeDzharkutan0.71137.6 ± 39.00.96yes
Middle Bronze AgeSapalli0.5263.9 ± 38.50.10yes
Late Bronze AgeGonur0.340−37.1 ± 39.1−0.95no
Late Bronze AgeDzharkutan0.1996.1 ± 72.80.08yes
Late Bronze AgeSapalli0.454−72.3 ± 67.1−1.08no

No model resolves a BMAC stream: every coefficient is statistically indistinguishable from zero (|Z| ≤ 1.1). The standard errors tell the more precise story. They run from ±34 to ±73 percentage points, ten to twenty times the ±3.4 of the European farmer term tested the same way, and when a BMAC source enters, the Iranian farmer coefficient's error inflates in step (±42 in the figure) while the Steppe and AASI estimates do not move at all. The model is telling us that BMAC ancestry and Iranian farmer-related ancestry are close to interchangeable from this genome's point of view; the two coefficients trade freely against each other, and the data cannot apportion ancestry between them.

The interchangeability is confirmed from the other direction: substituted directly into the farmer slot instead of added as a fourth source, BMAC populations fit well (twenty source combinations, p = 0.77–0.92, with proportions shifting only a few points). BMAC ancestry is, from South Asia's vantage, largely the same Central Asian farmer-related ancestry the model already carries. The contrast with the European farmer test is the internal control: the same design resolved that term crisply, so the failure to resolve a BMAC term reflects genuine genetic redundancy, not a weak instrument.

The conclusion matches the ancient-DNA record. Narasimhan et al. (2019) found that the main BMAC population contributed little ancestry to South Asian gene pools, and that gene flow ran detectably the other way: BMAC individuals carry 2–5% Andamanese hunter-gatherer (AHG)-related ancestry from South Asia, while a reciprocal BMAC signal in South Asians is undetectable. The Steppe stream crossed the Oxus world; genetically, it did not stay.

Additional Genetic Insights

qpAdm measures autosomal mixture only. Two independent data types check it: the paternal lineage from Big Y-700, and a G25 fit computed with entirely different machinery.

Y-DNA Haplogroup Analysis via Big Y-700

y-dna Indo-Aryan
Y-DNA Haplogroup R-FTF40903

The Big Y-700 test resolved the paternal lineage to R-FTF40903, a subclade of R1a-Z93 through L657. The lineage's placement can be inspected on both public phylogenies: the YFull tree and the FTDNA Discover haplotree. Its chronology, from published work:

  1. R1a-M417 expansion, c. 3500 BCE: the parent lineage expands with the early Steppe world and splits into a European branch (Z282, later carried by Corded Ware into Europe) and an Asian branch (Z93) (Underhill et al., 2015).
  2. Z93, c. 2900–2600 BCE: the Asian branch forms; ancient carriers appear in Sintashta and Andronovo contexts (Silva et al., 2017).
  3. L657, c. 2200 BCE: the largest South Asian subclade forms within Z93 > Z94, its date squarely inside the migration window (Silva et al., 2017).
  4. South Asian prevalence by the late 2nd millennium BCE: multiple closely related founder clades expand across the subcontinent, consistent with arrival through the 1800–1500 BCE corridor window the ancient-DNA record identifies.

One limit should be stated plainly: no sampled ancient individual has yet been confirmed to carry L657 itself; the ancient carriers sit on neighbouring Z93 branches, so the paternal evidence corroborates the horizon without pinning the transmitting population. The paternal line is in any case a co-witness rather than independent proof: a single lineage cannot measure ancestry proportions. Its value is that it names the same source world as the autosomal fits, through an entirely different inheritance system: Z93's sister branch travelled west with Corded Ware while Z93 itself went east and south with the Sintashta–Andronovo horizon, the two directions of one dispersal described under Key Findings.

IllustrativeDNA Results

illustrativeDNA Indo-Aryan
IllustrativeDNA G25 Models

The third cross-check uses IllustrativeDNA's G25 fit, a least-distance model in a 25-dimensional PCA space. Its two-way model gives:

  1. Andronovo Culture: 34.1%
  2. Indus Valley Civilization: 65.9%

This two-way model maps directly onto the Late Bronze Age qpAdm results, using an entirely different method and reference frame:

  • The Andronovo component (34.1%) matches the Steppe estimate in the Late Bronze Age model (33.7 ± 3.6%).
  • The Indus Valley Civilization component (65.9%) matches the combined Iranian farmer-related and AASI ancestry (37.4% + 28.9% = 66.3%), which is exactly what an IVC reference should absorb: the Indus population was itself a mixture of those two streams.

Two methods with different assumptions converging within a percentage point is meaningful corroboration. It also illustrates the reference-frame point made throughout this analysis: G25's "Indus Valley" component is not a fourth ancestry stream but a repackaging of two of the three streams the qpAdm models resolve separately.

A caution on where this cross-method agreement does and does not extend. G25 fits are constrained least-distance solutions in a 25-dimensional PCA space, and that machinery happily produces clean-looking percentages even from mutually collinear sources. Its deep "Neolithic ingredient" models (Zagros, EHG, CHG, EEF, and similar) have no valid qpAdm equivalent: run through f-statistics, those source sets are degenerate, returning negative or greater-than-100% weights with every coefficient statistically empty. We verified this directly on the present target. Proximal Bronze Age models built from real, genetically distinguishable populations, like the two-way fit above, are the level at which the two methods can check each other, and there they agree.

Conclusion

Modelled against AADR v66, with one fixed farmer source and the Onge standing for the unsampled AASI lineage, the Jatt Sikh genome resolves into three ancestry streams: roughly 35% Steppe, 36–39% Iranian farmer-related, and 25–29% AASI. The principal findings:

  • One Steppe gene pool. The Steppe proportion is statistically identical whichever era supplies the source (33.3–38.1% across all ten window-by-source combinations). Because every tested source descends from the same Steppe gene pool, this stability bounds the combined Steppe contribution tightly; how many crossings delivered it is a question these models cannot ask.
  • A Sintashta–Andronovo (Steppe_MLBA) source, best supported. The R1a-Z93 > L657 paternal lineage and corridor chronology constrain the source to the horizon, the autosomal fits are fully consistent with it, and the European farmer content of the stream points to an earlier, less farmer-admixed population within that horizon. The specific branch lies beyond the resolution of current samples.
  • Three sources are sufficient. Explicit fourth terms fail in informative ways: the European farmer term is unresolvable or negative depending on what the other sources already carry, and no BMAC term can be resolved at all, because BMAC ancestry is statistically interchangeable with the Iranian farmer stream from this genome's vantage.
  • The indigenous substrate persists. AASI ancestry stands near a quarter of the genome in every model. The migrations added to South Asia's genetic foundation; they did not replace it.

These results sit squarely within the current ancient-DNA picture. Steppe ancestry in South Asia runs highest in the northwest and declines southeastward, with genome-wide surveys of the Indian cline reporting Steppe proportions from near zero to ~45% (Kerdoncuff et al., 2025); a value near 35% is what a northwestern agro-pastoralist community is expected to carry. The earliest Steppe ancestry directly dated in South Asian remains appears in the Swat valley by 1200–800 BCE, consistent with arrival in the window modelled here (Narasimhan et al., 2019). On the Steppe side, the record is deep and consistent: Yamnaya formation is now traced to the Caucasus–lower Volga cline (Lazaridis et al., 2025), the steppe's later population sequence is documented across 137 ancient genomes (Damgaard et al., 2018), and the Early Bronze Age expansions eastward are known to have left little genetic trace south of the steppe (de Barros Damgaard et al., 2018), the same asymmetry these windows reproduce: the Early Bronze Age sources fit only as stand-ins, while the Middle-to-Late Bronze Age populations are the migration's contemporaries. The reference record is still growing, with over ten thousand newly reported West Eurasian ancient genomes in the latest release (Akbari et al., 2026).

What the genome records, it records plainly: an Iranian farmer-related and AASI foundation of the kind that built the Indus world, joined in the second millennium BCE by a Steppe stream whose paternal lineage still runs unbroken from the Bronze Age steppe to the Punjab. The clearest open question is branch-level attribution within the Sintashta–Andronovo horizon, and it will be settled by denser ancient sampling of the Central Asian corridor, not by further modelling of the samples that exist.

Supplementary Data

qpAdm Model Outputs

Complete outputs for the 144 qpAdm models behind the current figures, fitted on AADR v66 with the twelve-population reference set described under Methodology, plus the unsupervised clustering sweep:

Raw per-model outputs (AADR v66)

One receipt file per published model, in the figures' exact regime: full weights with standard errors, the rank-drop table behind each p-value, and the pop-drop table of all nested submodels. Each file states its own dataset, target, and complete left and right population lists.

Three-source window models:

  1. EBA — Samara Yamnaya (shown in figure)
  2. EBA — Khakassia Afanasievo
  3. EBA — Altai Afanasievo
  4. EBA — Kalmykia Yamnaya
  5. EBA — Orenburg Yamnaya
  6. MBA — Poland Corded Ware (shown in figure)
  7. MBA — Czechia Corded Ware
  8. MBA — Esperstedt Corded Ware
  9. LBA — Samara Srubnaya (shown in figure)
  10. LBA — Chelyabinsk Sintashta

Four-source European farmer tests:

  1. EBA + Austria LBK
  2. MBA + Austria LBK
  3. LBA + Austria LBK

Four-source BMAC tests:

  1. EBA + Gonur
  2. EBA + Dzharkutan
  3. EBA + Sapalli
  4. MBA + Gonur
  5. MBA + Dzharkutan
  6. MBA + Sapalli
  7. LBA + Gonur
  8. LBA + Dzharkutan
  9. LBA + Sapalli

Archived v54 raw outputs

The original 2024 analysis ran on AADR v54.1.p1 with per-chip targets and different source populations, including the AASI-rich farmer source whose composition effects are discussed in the Results. Its raw outputs are preserved unmodified:

  1. 23andMe - Russia Samara EBA Yamnaya
  2. AncestryDNA - Russia Samara EBA Yamnaya
  3. 23andMe - Czech Corded Ware
  4. AncestryDNA - Czech Corded Ware
  5. 23andMe - Russia Srubnaya Alakul
  6. AncestryDNA - Russia Srubnaya Alakul
  7. 23andMe - 3-way qpAdm Sintashta
  8. AncestryDNA - 3-way qpAdm Sintashta
  9. 23andMe - Russia MBA Poltavka
  10. 23andMe - Turkmenistan Gonur BA 1
  11. 23andMe - Mongolia EIA Slab Grave 1
  12. 23andMe - Kazakhstan Kumsay EBA
  13. 23andMe - BMAC
  14. AncestryDNA - BMAC
  15. 23andMe - 4-way qpAdm Gonur BA1
  16. AncestryDNA - 4-way qpAdm Gonur BA1
  17. 23andMe - 4-way qpAdm Gonur BA2
  18. AncestryDNA - 4-way qpAdm Gonur BA2
  19. 23andMe - 4-way qpAdm Geoksyur
  20. AncestryDNA - 4-way qpAdm Geoksyur

References

  1. Allen Ancient DNA Resource (AADR), version 66. Harvard Dataverse.
  2. Akbari, Ali, et al. "Ancient DNA reveals pervasive directional selection across West Eurasia." Nature 654 (2026): 419-428. Link to Article
  3. Allentoft, Morten E., et al. "Population genomics of Bronze Age Eurasia." Nature 522.7555 (2015): 167-172. Link to Article
  4. Anthony, David W. The Horse, the Wheel, and Language: How Bronze-Age Riders from the Eurasian Steppes Shaped the Modern World. Princeton University Press, 2007. Link to Book
  5. Basu, Analabha, et al. "Genomic reconstruction of the history of extant populations of India reveals five distinct ancestral components and a complex structure." Proceedings of the National Academy of Sciences 113.6 (2016): 1594-1599. Link to Article
  6. Broushaki, Farnaz, et al. "Early Neolithic genomes from the eastern Fertile Crescent." Science 353.6298 (2016): 499-503. Link to Article
  7. Damgaard, Peter de Barros, et al. "137 ancient human genomes from across the Eurasian steppes." Nature 557 (2018): 369-374. Link to Article
  8. de Barros Damgaard, Peter, et al. "The first horse herders and the impact of early Bronze Age steppe expansions into Asia." Science 360.6396 (2018): eaar7711. Link to Article
  9. Fournier, Romain, Alice Fulton, and David Reich. "A SNP panel for coanalysis of capture and shotgun ancient DNA data." Genome Research 36 (2026). Link to Article
  10. Haak, Wolfgang, et al. "Massive migration from the steppe was a source for Indo-European languages in Europe." Nature 522.7555 (2015): 207-211. Link to Article
  11. Harney, Éadaoin, et al. "Assessing the performance of qpAdm: a statistical tool for studying population admixture." Genetics 217.4 (2021): iyaa045. Link to Article
  12. Kerdoncuff, Élise, et al. "50,000 years of evolutionary history of India: Impact on health and disease variation." Cell 188 (2025). Link to Article
  13. Lazaridis, Iosif, et al. "The genetic origin of the Indo-Europeans." Nature 639 (2025): 132-142. Link to Article
  14. Maier, Robert, et al. "On the limits of fitting complex models of population history to f-statistics." eLife 12 (2023): e85492. Link to Article
  15. Narasimhan, Vagheesh M., et al. "The formation of human populations in South and Central Asia." Science 365.6457 (2019): eaat7487. Link to Article
  16. Papac, Luka, et al. "Dynamic changes in genomic and social structures in third-millennium BCE central Europe." Science Advances 7.35 (2021): eabi6941. Link to Article
  17. Reich, David, et al. "Reconstructing Indian population history." Nature 461 (2009): 489-494. Link to Article
  18. Shinde, Vasant, et al. "An ancient Harappan genome lacks ancestry from Steppe pastoralists or Iranian farmers." Cell 179.3 (2019): 729-735. Link to Article
  19. Silva, Marina, et al. "A genetic chronology for the Indian Subcontinent points to heavily sex-biased dispersals." BMC Evolutionary Biology 17 (2017): 88. Link to Article
  20. Underhill, Peter A., et al. "The phylogenetic and geographic structure of Y-chromosome haplogroup R1a." European Journal of Human Genetics 23 (2015): 124-131. Link to Article
  21. Wang, Chuan-Chao, et al. "Ancient human genome-wide data from a 3000-year interval in the Caucasus corresponds with eco-geographic regions." Nature Communications 10 (2019): 590. Link to Article