Every intelligence runs the same loop
Pull the time scale back far enough and every history looks like one process: a variation is proposed, a consequence is measured, and whatever survived the measurement is kept. Four billion years of biology ran that loop. So did ten thousand years of human record. A model that can write to its own weights runs it too — at a different rate, and in a form that looks nothing like the first two.
What changes between those cases is not the loop. It is the quality of the judgment applied at the moment of keeping. Natural selection never invented anything; it decided what was retained. That is the part we work on.
Improvement can be written to three surfaces, and they behave differently. Data amplifies — in both directions, which is why the negative half matters as much as the positive one. The weights internalise what survives, permanently and expensively. Harness — the scaffolding bolted around a model — elicits capability the weights already hold, cheaply and reversibly. All three can evolve, and so can the environment they are measured against.
So we build the substrate that loop needs: corpora a model can actually learn from, environments that push back and grade every step of a run, and the expert judgment that decides what counts as correct. Then we publish how each one is constructed, because a benchmark whose methodology is private is a number, not a measurement.
The argument in full is in the essays; the work itself is under what we build.
