Measurements

What has actually been measured

Two numbers are only comparable when they share a metric and a protocol version. Every figure on this site carries benchmark@version and a run date for exactly that reason — putting two differently-produced numbers in one table is the fastest way for a site like this to stop being worth reading.

No runs yet. The protocols below are written down first, deliberately, so the numbers cannot be shaped to fit a story after the fact. Nothing on this site is presented as measured until a run file exists to back it.

Protocols

Browser snapshot cost — What does one page observation cost, and what does the server charge before it does anything?
VersionEmitsFixturesRunsMethod
v1Tokens per page snapshot · Tool slots consumed · Resident schema tokensfixtures/browser-snapshot-cost/v10unpublished

For each candidate, install the server into a clean harness and record tool-count (tools registered) and schema-tokens (the serialised tool schemas, tokenised with o200k_base) before any call is made. Then drive the three fixture pages below: navigate, take one observation with the server's default observation tool at default settings, and tokenise the full tool result. Report the median of the three pages as snapshot-tokens. No model is invoked — this is characterisation, not behaviour, so the numbers are exact rather than distributional. Record failures to install as status:error with a note; a server that cannot be installed is a finding, not a blank cell.