What has actually been measured
Two numbers are only comparable when they share a metric and a protocol version. Every figure on this site carries benchmark@version and a run date for exactly that reason — putting two differently-produced numbers in one table is the fastest way for a site like this to stop being worth reading.
No runs yet. The protocols below are written down first, deliberately, so the numbers cannot be shaped to fit a story after the fact. Nothing on this site is presented as measured until a run file exists to back it.
Protocols
| Version | Emits | Fixtures | Runs | Method |
|---|---|---|---|---|
| v1 | Tokens per page snapshot · Tool slots consumed · Resident schema tokens | fixtures/browser-snapshot-cost/v1 | 0 | unpublished |
For each candidate, install the server into a clean harness and record tool-count (tools registered) and schema-tokens (the serialised tool schemas, tokenised with o200k_base) before any call is made. Then drive the three fixture pages below: navigate, take one observation with the server's default observation tool at default settings, and tokenise the full tool result. Report the median of the three pages as snapshot-tokens. No model is invoked — this is characterisation, not behaviour, so the numbers are exact rather than distributional. Record failures to install as status:error with a note; a server that cannot be installed is a finding, not a blank cell.