A single-page announcement post for the chutescoder project. Plain HTML/CSS/JS,
no build step, no external network calls (no CDN scripts, no web fonts, no
analytics). Everything the page displays comes from the JSON files in data/.
web/
index.html the page (structure + prose)
styles.css light + dark themes, one stylesheet
app.js chart, tables, filters, theme toggle (no deps)
build_data.py regenerates availability.json (+ ab_harness.json, opt-in)
data/
availability.json ← GENERATED. Per-model probe success rate.
binary_run.json ← the real chutescoder binary, RLM mode on (hand-transcribed)
smoke_run.json ← the pre-integration driver run on Kimi K3 (hand-transcribed)
ab_harness.json ← TOMBSTONE ONLY. No measurements. Not rendered.
results.json ← the standard-benchmark grid. Still empty on purpose.
public_scores.json ← the 130 published scores for Kimi K3 / GLM-5.2
harness_spread.json ← the three Terminal-Bench 2.1 runs in the hero chart
model_availability.jsonl raw probe log, copied from ../data/
render.yaml Render blueprint (static site, free plan)
Never put a number on this page that is not in one of the data/*.json files
with a source_url, or that has not actually been measured. An empty cell is
correct. A plausible-looking placeholder is not — this is public marketing
material and a single invented score poisons the whole page.
results.json ships with "cells": [], so every (benchmark × model × arm) cell
renders as pending. That is the intended state until real runs land.
The page used to say the model is offered exactly one tool and cite the
model_visible=["python"] log line for it. That line is emitted before
finalize_tool_router runs, so it reports an intermediate registry state. Under
the documented key [features.code_mode] direct_only_tool_namespaces, nine
tools go over the wire while the log still says one — the reviewer measured it
by pointing the binary at a recording HTTP server.
The substantive claim survived: on default config the request carries python
and nothing else. The page now says "on the default configuration", names the
escape hatch, and cites the recorded requests. The log field is renamed
exposure_after_collapse upstream. If you restore a one-tool claim anywhere,
cite a wire capture.
Same rule, applied to Windows: say the build succeeded on windows-latest and
43 of 44 tests passed, the one failure being our own POSIX-only test. Do not
write "Windows CI failed" — that reads as "it doesn't build on Windows", which is
false.
2026-08-06. Three rounds of review. The experiment supports no performance claim in either direction, and the page must not assert one.
Round one — headline "the harness wins" — withdrawn for two faults:
bench/ab_harness.pynever cleaned/tmpbetween runs, so the compaction task ran with ~480 leftover same-named files from earlier tasks on disk. The single failing trial that the whole accuracy result rested on went looking in them and said so in its own final message.- The arms were not comparable. RLM got a 5,438-byte prompt teaching batching
and filtering plus a post-compaction notice naming
ctx.grep; CLASSIC got 232 bytes and was not told compaction had happened. "The only variable is the tool interface" was false.
Rounds two and three — the clean re-run, headline "the harness loses" — also withdrawn, for three more:
- The compaction result was false. The write-up quoted the failing
baseline trial's own explanation ("the value was not preserved in my
summary") as evidence. The reviewer regenerated the corpus: that summary
contains
CACHE_TTL_BUDGET=837verbatim, in its first bullet list. Across all ten trials the baseline's summary retained the target fact 5/5. The mechanism under test never fired, sorecall-40supports nothing. Same error as round one: believing an agent's narration of its own behaviour. ctxwas used in 0 of 20 trials. Every recovery was a plain kernel-variable lookup. What is demonstrated is variable persistence across compaction — real, useful, more portable — not context-as-a-variable.- The negative result is partly an artefact of prompt length. On
needleboth arms did byte-identical work and RLM still used 2.3× the tokens; 100 % of that gap is system-prompt length, ~45–52 % on audit and recall.
Consequences for this folder:
data/ab_harness.jsonships as a tombstone only — retraction text, no measurements.build_data.pywrites the real one only under--emit-ab.app.jshas no A/B renderer. The removed code is described in a comment where it used to live. Do not resurrect it — it drew the invalid comparison.- Section 04 of the page is a retraction, not a placeholder. It names all five faults and carries zero measurements — including zero negative ones.
recall-5was dropped frombuild_data.pyentirely: its per-trial records no longer exist and its numbers survived only as hand-typed literals in this script.
Everything round two already had (n = 5 per cell, isolated corpus root per
trial, a CLASSIC prompt written with comparable care, a symmetric
post-compaction notice, a de-contaminated audit task), plus:
- an RLM system prompt trimmed to the baseline's length, so the token column measures the mechanism and not the prompt;
- a compaction task whose summariser demonstrably loses the target fact — verify that from the summary text before scoring anything;
- a snapshot of
messagestaken beforecompact()replaces it, so phase 1 is auditable; - transcripts for every cell, checked against every behavioural claim. No claim survives on the model's own account of itself.
Then:
cd web
cp ../data/model_availability.jsonl data/
python3 build_data.py --emit-ab…and write a new renderer. Keep the retraction visible above whatever replaces it — a reader who saw the withdrawn numbers is owed the correction more than a reader who did not is owed a clean page. A null result is an acceptable outcome and must be published as one.
To add a task, append an entry to TASKS in build_data.py with its report
filename, a blurb and a shape.
cd web
cp ../data/model_availability.jsonl data/
python3 build_data.pyThis one is safe and should be re-run whenever the probe log grows. The
GLM-5.2 blocked_reason in results.json deliberately contains no counts —
the numbers come from availability.json at render time, so they cannot go
stale. Keep it that way. Twice in review a hand-typed availability figure had
drifted from the log it cited; that is why the prose contains none.
The probe watcher was stopped at the last timestamp in the log, so
availability.json now carries closing figures rather than a snapshot of a
growing file ("watcher_stopped": true). Closing tally: Kimi K3 34/34,
GLM-5.1 34/34, GLM-5.2 11/34 — the watcher appended one more round after
docs/RESULTS.md was written, which is why that document says 33 probes. The
page renders whatever the file says.
data/binary_run.json — the real binary, bench/binary_ab.sh parser-bug,
rlm.enabled=true, provider chutes, Kimi K3. Re-derived cell by cell from the
session rollout for thread 019fd4dc-…. It deliberately carries no timing or
token comparison between the two arms: n = 1 each, and this page asserts no
performance result in either direction.
Read the artefact note in that file before citing it. bench/binary_ab.sh
used a fixed output directory and truncated it on every invocation, so
reports/binary_ab/ holds exactly one of the five attempts docs/RESULTS.md
§3.7b cites — the last, whose RLM arm hit a provider 403. It is not this run.
This run is auditable only from the session rollout, which lives outside the
repository. The script now writes one timestamped directory per invocation. This is the same mistake that lost recall-5's per-trial records, in
a second script, which is why the page states it rather than a commit message.
The third review pass also found that the earlier version of this file miscounted
its error cells and said the run verified with a host call when cell 10 used
subprocess. Both are corrected, and the timeline now walks all 11 cells —
including the four the file-shim bug corrupted.
data/smoke_run.json — the pre-integration driver run, hand-transcribed from
reports/smoke_kimi_k3/RESULT.md. There is no machine-readable source for it.
The review checked this one and its figures reproduce; it is n = 1 and the page
says so.
Editing data/results.json is the only thing you need to do. No HTML, no JS.
Fields other than benchmark / model / arm / state are optional; anything
missing renders as an em dash rather than a guess.
The grid itself is generated from benchmarks × models × benchmarks[].arms,
so the "N of M cells measured" meter and the pending list update themselves.
Two ways, and both render a red blocked pill rather than a grey pending one:
- Per model — set
"blocked": true,"blocked_reason": "…"and optionally"blocked_evidence": "data/…"on the entry inmodels. Every cell for that model becomes blocked, and the reason appears as a callout under the status meter. This is whatglm-5.2currently uses. - Per cell — a
cellsentry with"state": "blocked"and a"notes". Use this when only some cells are affected. A per-cell entry overrides the model-level flag.
Do not silently substitute a different model; that is exactly the failure
mode the benchmark plan is designed to avoid. To un-block, delete the flag —
the cells fall back to pending.
Once at least one cell is measured, the status pill switches from measurements in progress to partial results automatically. Also update, by hand:
"headline"— one honest sentence about what is now known."harness_commit"— the chutescoder sha the runs used. Until this is set the page says "not yet pinned"."status"—"pending"→"partial"→"complete"(informational).
When everything is measured, also soften the amber banner at the top of
index.html (<aside class="banner" id="status-banner">) — it is the one piece
of pending-state copy that lives in HTML rather than JSON.
Append to benchmarks / models in the same file. New rows appear
automatically. A model with "at_risk": true and an "at_risk_reason" gets a
hoverable warning marker on its pending cells.
data/public_scores.json is a copy of ../data/public_scores.json (the
research dataset, 130 records, schema
{model, benchmark, variant, score, unit, source_url, source_type, date, notes}).
Refresh it with:
cp ../data/public_scores.json data/public_scores.jsonThe table, the filters, the record count and the footer summary all derive from it — no other file needs touching.
data/harness_spread.json is separate on purpose: it holds the three
Terminal-Bench 2.1 runs behind the hero chart, including the Vals AI 80.90
row, which is not in public_scores.json (that dataset was compiled from the
Artificial Analysis API and the two lab model cards only). It was verified
directly against https://www.vals.ai/benchmarks/terminal-bench-2-1, and the
chart legend labels it as such. If you ever add it to public_scores.json, flip
"in_public_scores": true so the disclaimer disappears.
There is no build. But the page fetches JSON, so file:// will not work — serve
it over HTTP:
cd web
python3 -m http.server 8080
# → http://localhost:8080Check both themes (the ☀/☾ button in the header, or your OS setting) and a narrow viewport before shipping.
Hosted on Render as a static site on the free plan, in the "Florian S's Workspace" workspace.
| service | chutescoder-web |
| repo | https://github.com/chutesai/chutescoder-web |
| branch | main |
| publish path | . |
| build command | (none — echo "static site, no build") |
| auto-deploy | on |
Because auto-deploy is on, a push to main is the deploy:
cd web
git add -A && git commit -m "results: terminal-bench arm C, kimi-k3" && git pushRender picks it up within a few seconds and redeploys. To force one by hand, use
trigger_deploy in the Render MCP or the "Manual Deploy" button in the
dashboard.
Note the site directory is committed to its own repo (chutescoder-web), not
to chutesai/chutescode — the fork stays clean, and the marketing page can be
updated without touching the harness.
{ "benchmark": "terminal-bench", // must match a benchmarks[].id "model": "kimi-k3", // must match a models[].id "arm": "C", // must match an arms[].id, and be listed in benchmarks[].arms "state": "measured", // "measured" | "blocked" | "pending" "score": 71.4, // the benchmark's native metric — no rounding up "n": "12 / 89 (seed 1337)", "tokens": "4.1M in / 220k out", "cost_usd": 18.42, "run_id": "br_01J...", // bench-runner run id (recorded, not yet displayed) "notes": "preliminary — 10% subset" }