Skip to content

19. Benchmarks

This page renders every committed study in site/src/data/results/ — today that is one observational field study (a real production run, instrumented from durable logs). Controlled benchmarks — the three-arm runner below — join the list only when someone runs the runner on their own machine, and the renderer says so honestly until then. No number here is decorative: each cites its source, and the arms runner is reproducible on any machine.

2026-09-27 · qwen3.8-flash-next @ 65,536 (llama.cpp, 8GB VRAM) · observational field study, from durable session logs

Real multi-agent production session, instrumented post-hoc from durable session logs — not a controlled benchmark. Arms-based benchmarks run separately via the local runner; this is what the system actually faced in the wild.

MetricValueSource
Sessions in the study10 parent + 9 sub-agents~/.dsh/sessions logs
LLM requests served402usage events
Compactions (all zero inference tokens)21compaction/summary events
Conversation tokens archived verbatim679,707shadowedTokenCount sum
Inference tokens spent compacting0deterministic summarize
Archive retrievals by the agents49 / 49 succeeded ~497 KB restoredtool/call↔tool/result join
Chapters written (registry = files on disk)30 / 30dsh_chapters.json vs .dsh-chapters/
Turn/step events unbalanced, or error events0all 10 sessions closed clean
Peak steady-state cache hit (quiet windows)0.997–1.000usage series
Heaviest single session95 requests, 6 compactions one task threaded through 6 windowschild 6d2d93a6
Stock-method equivalence estimate≈ 113 min 680K prefill tokens @ 100 tok/s + 21× summary generationderived, clearly-labeled estimate

The runner compares the same long-conversation task under three memory regimes, at rising pressure horizons:

Arm Memory regime Expect
Sliding-window truncation keep last K tokens recall collapses once the fact leaves the window
LLM summarization compact by paraphrase recall partial, cost = full-history prefill each compaction
dsh-chapters verbatim archive + TOC + retrieval recall exact on retrieval; compaction = 0 model tokens

Metrics per arm × horizon: uncached tokens (the real cost), recall % (needle facts restored in the final answer), wall time (summed across the arm’s calls — the summarization tax shows exactly where it lives), and task completion. Results are machine/model-specific by design — that’s the point of running it on your local model.

Benchmarks need a live model and a quiet machine, so they are not part of CI. From the repo root:

Terminal window
npm run build # first: the arms import the plugin's built lib/ modules
npm run bench # one pass against the local llama server (default :8080)
npm run bench -- --horizon 65536 --needles 12

Each pass appends site/src/data/results/bench-<utc>.json; commit it and rebuild the site to publish. Methodology and the exact protocol live in benchmarks/README.md — including its documented honesty limits (arm C is a module-level emulation of the plugin, not the browser harness; needle grading is exact-substring and undercounts paraphrased-but-correct recall).

Next: 20. Field study — the TreeSeed run.