logicsrc/docs/openstream/reports
Anthony Ettinger 9f42ce222a
docs: OpenStream benchmark reports, published per release (#150)
Adds a reports section to the OpenStream spec so its claims rest on a
reproducible measurement rather than an assertion. Each report is a run of
the envelope over a defined corpus on real hardware: proof that
decompression restores every byte, that an incompressible input costs only
the framing overhead, that a compressible one saves what it claims against
the complete wire size, and how long each codec takes.

- docs/openstream/reports/ holds a machine-readable <id>.json (canonical,
  with a versioned schema) and a rendered <id>.md per report, plus a README
  on the shape and on submitting one. The seed report is nixamp 0.17.1 over
  the synthetic corpus, labelled synthetic so no one reads a padded-fixture
  number as production.
- The site renders them at /docs/openstream/reports (index) and
  /docs/openstream/reports/<id> (one report), under the dynamic /docs/[slug]
  tree so the reports routes never shadow a spec's own doc page. A small
  lib/reports.ts reads the JSON at build time; REPORTED_SPECS keeps the
  route surface explicit. sitemap includes the index and every report.
- The spec doc gains a Benchmark reports section linking there, and repeats
  the honest caveats: OpenStream frames Zstandard and gzip rather than being
  a new algorithm, synthetic padding flatters a codec, an efficient real
  feed saves little, and round-trip exactness is the one pass/fail.

The report format is produced by `nixamp compression benchmark` (in the
nixamp repo); a release runs it and commits the two files here.


Claude-Session: https://claude.ai/code/session_01MxNif5tsYq4LczgG7aE8Jp

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
2026-09-12 06:23:46 -07:00
..
2026-09-12-nixamp-0.17.1-synthetic.json docs: OpenStream benchmark reports, published per release (#150) 2026-09-12 06:23:46 -07:00
2026-09-12-nixamp-0.17.1-synthetic.md docs: OpenStream benchmark reports, published per release (#150) 2026-09-12 06:23:46 -07:00
README.md docs: OpenStream benchmark reports, published per release (#150) 2026-09-12 06:23:46 -07:00

OpenStream benchmark reports

Each file here is a reproducible run of the OpenStream benchmark: proof, on real hardware with recorded versions, that the envelope restores every byte, that an incompressible input costs only the framing overhead, that a compressible one saves what it claims against the complete wire size, and how long each codec takes. The site renders them at /docs/openstream/reports.

A report is two files sharing one id:

  • <id>.json — the machine-readable report (canonical). Its schema field is the version of the shape below.
  • <id>.md — the same run rendered for reading, shown on the site.

The id is YYYY-MM-DD-<implementation>-<version>[-<corpus>], for example 2026-09-12-nixamp-0.17.1-synthetic. Append the corpus label when the run used the built-in synthetic corpus rather than authorized real samples, so no one mistakes a padded-fixture number for a production one.

Publishing one

The reference implementation emits both files:

nixamp compression benchmark --out .

writes openstream-report.json and openstream-report.md. Rename them to the id, drop them in this directory, and open a pull request. The command exits non-zero if any codec that applied failed to restore byte-for-byte, so a report that would not build is caught before it is committed.

A report is published with every release of a reference implementation. The release step runs the benchmark, commits the two files here, and the site picks them up on its next deploy. Prefer a run over authorized real samples (--corpus DIR); a synthetic run is acceptable as a floor but is labelled as such and is not evidence of production savings.

What the numbers mean, and do not

OpenStream is a framing envelope over Zstandard and gzip, not a new compression algorithm. A report measures those codecs at the block boundary, framed honestly. Read the caveats array in every report: a synthetic padded transport stream saves about its padding share and says more about the padding than the codec; an efficient real feed saves little; timings are for the one machine named in environment. The one number that is a pass/fail rather than a measurement is round-trip exactness, and it must hold for every applicable codec on every sample.