Strategies that do better than the crowd grow; the rest shrink. No agent reasons, no one optimizes — yet the population is dragged toward equilibrium by differential reproduction alone. Down the center, data flows: a payoff matrix and a starting mix go in, the replicator ODE integrates on the strategy simplex, and a proven trajectory comes out. The blue team builds and defends it; the red team tries to break it.
source P. D. Taylor & L. B. Jonker, Evolutionarily Stable Strategies and Game Dynamics, Math. Biosciences 40 (1978) 145–156 — semanticscholar.org · Taylor&Jonker·1978 (no open DOI — journal/year cited, AMBER). Rendered, not quoted.
One rule, no rationality. Let x be the population share vector on the simplex, A the payoff matrix, so strategy i’s fitness is fi = (Ax)i and the crowd average is φ = x·Ax. Then
dxi/dt = xi ( fi − φ )
Above average → grow. Below → shrink. The −φ is the whole trick: it is exactly what makes the shares keep summing to 1. Live, for the current state:
| i | share xi | fitness fi | fi−φ |
|---|
A symmetric Nash equilibrium is exactly a rest point of this flow: there every surviving strategy earns the same payoff fi=φ, so dx/dt=0 — equilibrium reached with no reasoning agent anywhere.
Selection does what deliberation was supposed to. This is the mean-field shadow of the-policy-gradient — a single learner ascending expected reward — and the dynamic that selects among the fixed points of the-nash-equilibrium. Each sphere is the next one’s premise.
The blue team’s live check, recomputed from the running engine: the simplex stays invariant, the interior Nash is a genuine rest point, and a dominated strategy dies. If red drops the −φ term, the invariance breaks and this badge turns red.
A symmetric matrix game on three strategies. Entry Aij is the payoff to a player using i against a partner using j; against a whole population x, strategy i earns (Ax)i. That is all the environment is — no beliefs, no lookahead. The current matrix:
| A | vs 1 | vs 2 | vs 3 |
|---|
Plus a starting mix x(0) on the simplex. Feed both to the engine below; the flow field does the rest.
Every step is Euler’s method on dxi/dt = xi(fi−φ), computed from A on the spot — never a stored curve.
Four properties, checked live on load and re-checked by the witness: (1) the simplex is invariant — shares stay ≥0 and sum to 1 (to 1e−6) at every step; (2) the interior Nash is a rest point (dx/dt=0 to 1e−9); (3) a strictly dominated strategy goes extinct; (4) near the ESS the trajectory converges to it.
The blue witness (left) confirms these live; the red team (right) tries to make them false.
The model is a mean field: an infinite, perfectly mixed population with no mutation, no noise, no structure. And Taylor & Jonker warn in the same paper that the discrete map can overshoot and destabilize an ESS the continuous flow attracts — the equation is a limit, not the finite truth.
“Replicator dynamics always converges to a Nash equilibrium.” Cut. RPS cycles indefinitely; only special games (e.g. an interior ESS) attract — run RPS in the panel and watch it refuse to settle.
“Every rest point is a Nash equilibrium.” Cut. Every simplex corner is a rest point; the badly-chosen ones are unstable. Nash ⊆ rest points, not the reverse.
“The discrete update behaves like the ODE.” Kept, corrected. Only for small dt — Taylor & Jonker show the discrete dynamic can overshoot where the continuous one is stable.
The red team’s move: delete the average-fitness subtraction, so dxi/dt = xi·fi with no −φ. The total mass then grows and the shares stop summing to 1 — the population leaves the simplex.
Drop the −φ term and mass is no longer conserved; the witness (window 7) recomputes the sum, sees it drift off 1, and turns red. Nothing is faked — the attack is real and it is caught.