19. Benchmarks
This page renders every committed study in site/src/data/results/ — today that is one observational field study (a real production run, instrumented from durable logs). Controlled benchmarks — the three-arm runner below — join the list only when someone runs the runner on their own machine, and the renderer says so honestly until then. No number here is decorative: each cites its source, and the arms runner is reproducible on any machine.
2026-09-27 · qwen3.8-flash-next @ 65,536 (llama.cpp, 8GB VRAM) · observational field study, from durable session logs
Real multi-agent production session, instrumented post-hoc from durable session logs — not a controlled benchmark. Arms-based benchmarks run separately via the local runner; this is what the system actually faced in the wild.
| Metric | Value | Source |
|---|---|---|
| Sessions in the study | 10 parent + 9 sub-agents | ~/.dsh/sessions logs |
| LLM requests served | 402 | usage events |
| Compactions (all zero inference tokens) | 21 | compaction/summary events |
| Conversation tokens archived verbatim | 679,707 | shadowedTokenCount sum |
| Inference tokens spent compacting | 0 | deterministic summarize |
| Archive retrievals by the agents | 49 / 49 succeeded ~497 KB restored | tool/call↔tool/result join |
| Chapters written (registry = files on disk) | 30 / 30 | dsh_chapters.json vs .dsh-chapters/ |
| Turn/step events unbalanced, or error events | 0 | all 10 sessions closed clean |
| Peak steady-state cache hit (quiet windows) | 0.997–1.000 | usage series |
| Heaviest single session | 95 requests, 6 compactions one task threaded through 6 windows | child 6d2d93a6 |
| Stock-method equivalence estimate | ≈ 113 min 680K prefill tokens @ 100 tok/s + 21× summary generation | derived, clearly-labeled estimate |
The three arms
Section titled “The three arms”The runner compares the same long-conversation task under three memory regimes, at rising pressure horizons:
| Arm | Memory regime | Expect |
|---|---|---|
| Sliding-window truncation | keep last K tokens | recall collapses once the fact leaves the window |
| LLM summarization | compact by paraphrase | recall partial, cost = full-history prefill each compaction |
| dsh-chapters | verbatim archive + TOC + retrieval | recall exact on retrieval; compaction = 0 model tokens |
Metrics per arm × horizon: uncached tokens (the real cost), recall % (needle facts restored in the final answer), wall time (summed across the arm’s calls — the summarization tax shows exactly where it lives), and task completion. Results are machine/model-specific by design — that’s the point of running it on your local model.
Run them locally (never CI)
Section titled “Run them locally (never CI)”Benchmarks need a live model and a quiet machine, so they are not part of CI. From the repo root:
npm run build # first: the arms import the plugin's built lib/ modulesnpm run bench # one pass against the local llama server (default :8080)npm run bench -- --horizon 65536 --needles 12Each pass appends site/src/data/results/bench-<utc>.json; commit it and rebuild the site to publish. Methodology and the exact protocol live in benchmarks/README.md — including its documented honesty limits (arm C is a module-level emulation of the plugin, not the browser harness; needle grading is exact-substring and undercounts paraphrased-but-correct recall).