About
About these notes
What this river covers, how the notes are written, and what they leave out on purpose.
These are working notes about how language models get compared: what a comparison has to fix before it starts, how published numbers are produced, and which measurements survive a change of model version. The river runs newest first, and every note is dated because the subject moves.
How the notes are written
Each note describes method rather than standing. The aim is that a note stays usable after the particular systems in view have been replaced, which means the writing concentrates on procedure - what to hold constant, what to record, what a number can and cannot support - and stays away from the question of which system is currently ahead.
What is deliberately absent
There are no scores, rankings or verdicts on named systems here, and no figures presented as measurements. A published number without the harness that produced it cannot be checked by a reader, and an unverifiable number is worse than none: it carries the authority of measurement without the substance. Where a note needs an example it uses an unnamed candidate, because the point being made is about the procedure and not about the product.
Reading order
The notes are independent, but three of them carry most of the weight and are worth reading first: the note on fixing a task before choosing a model, the note on building a small private evaluation set, and the note on what a published benchmark score leaves out. The glossary collects the vocabulary the rest of the notes assume.