Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

87 Commits
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

🎬 FeatureClipStudio

Turn live product flows into polished, storyboarded proof assets.

Every UI state · an animated cursor that glides to each click (with a ripple) · a zoom‑to‑focus camera · the loading/streaming captured live (spinner spinning, results coming in) · step captions · raw JSON/state evidence when the proof depends on internals. Not a single final‑state "hero shot" — the viewer follows the whole flow.

License: MIT Remotion Playwright ffmpeg

A feature walkthrough GIF of NodeRoom: from the landing page, create a shared diligence room, ask the NodeAgent to run diligence on CardioNova — it locks the row, researches, and fills it to complete — then open the audit Trace, where every agent action carries a verdict, attribution, and evidence

↑ produced by this tool — every state, the click, the agent's work, and the proof.


Storyboard first

The default quality bar is not "nice zooms." A walkthrough should have a clear premise, viewer question, comparison axis, conflict, evidence, verdict, and final decision before capture starts. Camera moves are subordinate to the story.

For comparison demos, the renderer supports explicit scene, axis, question, takeaway, per-pane verdicts, and a final scorecard scene. See STORYBOARD.md for the required story beats and anti-patterns.

Why

Most README/demo GIFs are hero shots — they show the final screen, so a viewer can't tell where the user clicked, what the empty state looked like, or how the result was reached. This tool generates true walkthroughs: clean per‑state frames, an overlaid cursor that eases to each target and ripples on click, an Arcade‑style camera that zooms to the action (and pulls back to frame the result), a step caption, and a progress bar.

It's fully scripted + reproducible (the spec is a checked‑in "tape"), so the GIFs double as a regenerable integration smoke‑test of your UI.

How it works — a 5‑stage pipeline

walkthrough.specs.mjs     1. SPEC    ordered cap/act ops per feature
        │
        ▼
node walkthrough.mjs       2. CAPTURE Playwright drives your app, screenshots a CLEAN
                                      frame at each state + records the pointer target
   →  public/wt/<id>/*.png            + writes src/walkthrough.data.js
        │
        ▼
npx remotion render        3. RENDER  Remotion overlays the animated cursor + ripple +
   src/index.js WT-<id>               zoom/pan camera + caption + progress  →  mp4
        │
        ▼
ffmpeg (two‑pass palette)  4. GIF     stats_mode=diff + lanczos + bayer + diff_mode
   →  assets/feature-*.gif            =rectangle  →  clean, small, looping GIF

The instrument is noisy. Sample it.

Measured: the identical file scored 34/44 and then 22/44 on two runs. That is 27% of the scale, on the same bytes. Every "this cut improved" claim made from single readings was unsupported — including one this repo published and then retracted ("comprehension 9 → 16, the voiceover was the missing piece"), which was variance wearing the costume of a result.

Two causes, separated by correlating the runs:

  1. Ordinary sampling variance at temperature 0.2.
  2. The anti-uniformity re-ask itself. Both 22s re-asked; the 34 did not. Told "you gave 91% of dimensions the same score, force a spread", the model spreads downward — demoting to 0 and 1 without ever promoting to 2. The device added to stop the judge shrugging was biasing the number it produced. It is now opt-in behind --reask and off by default, because a debiasing mechanism that biases is worse than the shrug it replaced.

So --samples 3 is the default: three independent judgements, per-dimension median, prose taken from the run nearest the median total so the report still reads as one judgement. Spread is now 22/23/25 and 44/45/44 where it was 34 vs 22.

Never quote a single run. If you are claiming a delta, quote the sample totals on both sides of it.

Plain English is arithmetic, so it is not an API call

npm run readability -- --id TShero --min 80

Flesch Reading Ease per caption, before the render, naming the exact sentence and its hard words. The LLM judge can say a video is too technical; it cannot say which line, costs 40 seconds, and — given the variance above — could not resolve the change at all. Reading ease could.

Measured on TrialScope: mean 80.5 → 109.3, captions below gate 11 → 1. Two consequences: the judge's lay_sense evidence went from "will struggle with API calls, dimension, co-occurrence, peer sponsors" to "uses simple analogies", and runtime fell 19%, because plain sentences are also shorter to say.

Round-trip without re-capturing (captions are render-time data, so no pixel of the app changes):

npm run captions -- --id TShero --from captions.plain.json
npm run voice -- --id TShero --out out/x.vo.wav      # writes x.vo.holds.json
npm run retime -- --id TShero --holds out/x.vo.holds.json --shrink

Narration owns pacing. A storyboard is first timed for reading; speech is slower and unevenly so — nine of 21 lines overran their hold the first time a voice was added. The fix is not a faster reader: the picture under a caption is a still frame that costs nothing to hold longer.

Two video types, two rubrics

--mode demo (44 pts) --mode interview (50 pts)
answers does it work, and did anyone understand it why is it built this way, and can you defend it
second axis comprehension defensibility
unique dimensions own_case_transfer alternatives_named · tradeoff_honesty · falsifiability · failure_modes_named

The interview variant exists because an agent produces faster than its author can absorb, and the gap shows up first as a walkthrough full of WHAT and empty of WHY. tradeoff_honesty scores 0 if every tradeoff resolves in the author's favour, and failure_modes_named needs a specific checkable failure rather than humility used as a rhetorical move.

decks/trialscope-decisions.html is the worked example: eight decisions, four fixed zones each — chosen, rejected, cost, falsifier. Measured 45/50 (44/45/44).

Sound — generated from the storyboard, not laid over the top

npm run score -- --id TShero --video out/trialscope.mp4              # generated bed + sfx
npm run score -- --id TShero --video out/x.mp4 --music licensed.mp3  # your track, sfx kept

The gate scored a completely silent video twice, across two different cuts, and never mentioned it — a rubric only sees what it names, so silence was not a low score, it was invisible. soundtrack and audio_sync now exist for that reason, and the reference films were measured rather than guessed:

reference length integrated LRA
Mi173xGb0ZA (short launch film) 38.5s -13.2 LUFS 13.5 LU
xPK3nBLbpxc 156.6s -18.4 LUFS 4.1 LU
JLpDL7x50hA 174.3s -20.8 LUFS 6.1 LU

Short launch films sit loud and dynamic; long explainers sit quiet and compressed under a voice. A silent cut is not the neutral choice — it is the one shape none of the references take.

Their tracks are not reusable. This repo is public, and shipping someone else's copyrighted music in it is the problem itself, not a licensing detail. What transfers is the measurement. So score.mjs synthesises the bed sample by sample in plain JS — no dependencies, no rights to clear — and --music stays there for anyone with a licensed track, in which case the generated sfx are kept and only the bed is replaced.

Built from the storyboard, so sync is structural. A track laid over a finished video is synced by luck and drifts the moment a hold changes by four frames. The storyboard already knows where every click, zoom and activity burst falls, so the arrangement is derived from those timestamps and re-derives itself on any re-cut. The arrangement follows the STORY: sparse under the premise, lifting at the graph reveal, and deliberately thinnest during the proof section — the trace has to be readable, and music arguing with it there would cost more than it adds. The only full major resolve is the result beat.

Output is loudness-normalised to -14 LUFS (measured -13.0 on the TrialScope cut).

The gate — judge, revise, re-render (stage 5, and it is not optional)

node iterate.mjs           5. GATE    render → judge (two rubrics, 40 pts) → revision
  --comp WTC-<id>                     brief → re-render. Exits non-zero below the
  --out out/<id>.mp4                  gate, so "done" has to be earned.
  --for "<audience>"

npm run clip is the default path. There is no shorter one that skips the judge, and that is on purpose: every video this repo produced before stage 5 existed was judged exactly once, at the end, by whoever remembered — which turns the findings into release notes instead of edits.

Two rubrics, because they fail independently. rubric.mjs scores CRAFT (20) and COMPREHENSION (20) separately:

asks dimensions
Craft is it well made? storyboard clarity · state coverage · cursor truth · caption sync · pacing · legibility · proof feel · safety · signature moment · loop etiquette
Comprehension did anyone understand it? persona · purpose · use case · feature legibility · full interaction · responsiveness · flow · result · lay sense · own-case transfer

Craft is what a demo's author notices missing. Comprehension is what everyone else notices missing, and it never shows up in a craft score. The TrialScope cut that motivated this scored craft 11/20, comprehension 9/20 — well made, real states, a real peak at the network graph, and a viewer still could not say who it was for, what problem it solved, or how to point it at a question of their own.

Comprehension is scored from a named audience's seat — that is what --for is:

npm run clip -- --comp WTC-TShero --out out/trialscope.mp4   --for "a non-technical person who has never heard of this domain"
npm run judge -- out/trialscope.mp4 --for "a frontend engineer evaluating adoption" --gate 28

The same cut is a 2 on lay_sense for a domain expert and a 0 for someone who has never heard the jargon. --for makes that difference a number instead of an argument.

Outputs: <video>.judge.md (scorecard split by axis), .judge.json, .rounds.md (the score history across rounds — "20 → 31 after adding a premise beat" is the only form of an improvement claim worth believing), and on failure .next-cut.md, the revision brief.

Two things the loop does deliberately

It does not auto-apply its own notes. A loop that edits the storyboard from its own critic's brief converges on whatever the critic likes, which is not the same as a good demo, and leaves nobody holding the taste. The brief is written to disk and the process exits non-zero; a human or an agent applies it; round N+1 begins.

Anti-uniformity is enforced in code, not asked for in the prompt. The rubric carried an anti-uniformity clause for three revisions and the judge still returned 1/2 on 18 of 20 dimensions — a description wearing a score's clothes. Now if one score covers >70% of dimensions the judgement is re-requested once, with the offending distribution quoted back. A gate that returns the same verdict for every input is not a gate, and that includes the flat-1 verdict.

The same three files are vendored into NodeVideo and NodeSlide under tools/clip-gate/, so npm run clip means the same thing in all three.

Quick start

Prerequisites: Node 18+, ffmpeg on PATH, and your app running locally in a clean (no‑auth / demo) state.

git clone https://github.com/HomenShum/FeatureClipStudio
cd FeatureClipStudio
npm install
npx playwright install chromium

# Render the bundled worked example (ships with captured frames — no app needed):
npm run render:example                 # -> out/example.mp4
ffmpeg -y -i out/example.mp4 -vf "fps=15,scale=720:-1:flags=lanczos,split[s0][s1];[s0]palettegen=max_colors=128:stats_mode=diff[p];[s1][p]paletteuse=dither=bayer:bayer_scale=3:diff_mode=rectangle" -loop 0 example.gif

The bundled example is a solo-founder 3D proof-run app (builder console → generate a scroll-driven 3D product story → customer-facing landing page → internal proof report with gates) — the captured frames + src/walkthrough.data.js are included so it renders immediately.

Make it your own

  1. Write a spec. Edit walkthrough.specs.mjs — each feature is an ordered list of ops:
    {
      id: "Search", title: "Instant Search", accent: "#10b981", tab: "Search",
      steps: [
        { cap: "Type a query",        cursor: "input" },
        { act: "fill", sel: "input", value: "invoices", commit: "Enter" },
        { cap: "Hit search",          cursor: "btn:Search", click: true },
        { act: "click", sel: "btn:Search" },
        { act: "sleep", ms: 1200 },
        { cap: "Results, instantly",  hold: 90 },               // captures the result
        { act: "waitText", value: "results" },
        { act: "scrollEl", sel: "df", last: true },             // center the result widget
        { cap: "Filter to what matters", hold: 100 },
      ],
    }
    • cap = capture this state. cursor = where the pointer glides (click:true ripples there). hold = frames to dwell.
    • cap + burst: { ms, every } = capture the loading/streaming motion — a rapid frame sequence (spinner spinning, status updating, results streaming in), played back as real motion instead of a frozen snapshot. Put it right after the click that starts the work.
    • act = advance the UI: fill | click | upload | sleep | waitText | notRunning | scrollEl | scrollText | scrollLastChat | scrollTop | scrollY.
    • Selector shorthand: textarea · input · file · drop · chat · btn:<name regex> · aria:<label> · aria^:<prefix> · df/iframe/metric (for scrollEl) · any CSS.
  2. Capture + render: start your app's clean harness, then node walkthrough.mjsnpx remotion render src/index.js WT-<id> out/<id>.mp4 → ffmpeg.
  3. Embed the GIF under each feature's README heading.

Built and battle‑tested against Streamlit (see the capture lessons in SKILL.md: scope locators to the active tab panel, await upload registration, data‑grids are canvas, capture the loading state on purpose), but the spec/selector model works for any browser UI. See walkthrough.solo-founder.mjs for a non-Streamlit adaptation (React SPA with hash routes — simpler selector model, goto action for URL-based navigation).

Design principles (researched)

Distilled from Arcade, Supademo, HowdyGo, CleanShot, Rekort, Mux, ubitux's High‑quality GIF with FFmpeg, GIPHY, and WCAG — see SKILL.md for the full list with sources:

  • Two‑pass palette is mandatory (stats_mode=diff + lanczos + bayer + diff_mode=rectangle) — the difference between a banded mess and a clean demo, and it shrinks the file.
  • Zoom/pan to focus, eased, with a pre‑move delay — click‑triggered zoom (~1.3–1.6×) beats highlight‑only for comprehension and makes small text legible.
  • Cursor at ~1.5–2× OS size + a click ripple — a real cursor is invisible after downscaling; the ripple is the silent stand‑in for a click sound.
  • Show every state, including loading — never cut an action straight to a finished result.
  • Pace from the caption (no narration), write outcome statements ("Filter to overdue invoices", not "Click Filter").
  • 3–10 s, one feature, seamless loop, ~640–800 px wide. Ship MP4 + GIF; GitHub auto‑embeds a bare MP4 URL.

Live collaboration (multi-pane)

Single-cursor capture can't show what makes a collaborative app special — a change in one client appearing live in another. So the tool also has a multi-pane mode: it drives N browser contexts (separate "users") and renders them side-by-side, cursor on the acting client, with a burst over the moment the change propagates.

Two clients side by side: Client A adds a card and runs an agent; Client B sees both appear live via the server broadcast — real cross-client sync

The composition scales to N panes — same spec model, one window per client:

3-client variant (WTC-LiveSync3 — one add, two collaborators react) Three clients side by side: Client A adds a card and runs the agent; Clients B and C both update live

Room OS V0 -> V3 readable production proof

This repo now includes a production proof for Room OS. The capture script opens four fresh rooms on nodevoice.vercel.app, selects V0/V1/V2/V3, starts the same live model task, sends the same mid-run interrupt, opens the internal state layer, and exports motion, static README evidence, and raw JSON state snapshots.

The README should not force readers to decode a fast four-pane GIF. The readable proof is segmented below: one slow GIF, one version-specific JSON state crop, and one full syntax-highlighted JSON snapshot per version, then a final comparison table.

V0 Failure: transcript-only coordination

Room OS V0 slow proof loop: live room starts with transcript coordination, receives the count interrupt, then opens the version-specific state drawer

Room OS V0 version-specific JSON state showing transcript-only coordination, null durable room state, recent utterances, and scheduling shell

Full V0 JSON state (raw JSON)
{
  "transcriptOnlyState": {
    "durableRoomState": {
      "task": null,
      "intent": null,
      "workerGraph": null,
      "artifacts": null
    },
    "transcriptBuffer": {
      "totalUtterances": 11,
      "renderedLimit": 160,
      "recentUtterances": [
        {
          "slot": "agent-002",
          "name": "Ben",
          "speechAct": "question",
          "text": "Great, $2? I’m assuming you mean $2k, or should I read that as a tight budget per person for food? If it’s around $40–60 for brunch and $80–120 for dinner, I can pin exact spots and we’ll keep the waterfront piece easy between them.",
          "ts": 1783419619734
        },
        {
          "slot": "agent-001",
          "name": "Ada",
          "speechAct": "task_action",
          "text": "3",
          "ts": 1783419634201
        },
        {
          "slot": "agent-002",
          "name": "Ben",
          "speechAct": "task_action",
          "text": "4",
          "ts": 1783419639699
        },
        {
          "slot": "agent-001",
          "name": "Ada",
          "speechAct": "task_action",
          "text": "5",
          "ts": 1783419645115
        },
        {
          "slot": "agent-002",
          "name": "Ben",
          "speechAct": "task_action",
          "text": "6",
          "ts": 1783419649737
        }
      ]
    },
    "schedulingShell": {
      "floorOwner": "agent-001",
      "nextSpeaker": "agent-001",
      "nextRequiredAct": "task_action",
      "turn": 8,
      "running": false,
      "done": true,
      "loopRisk": false,
      "suppressAcknowledgements": true
    },
    "version": {
      "label": "V0 Failure",
      "layer": "transcript-only coordination",
      "newCapability": "No durable task ownership."
    },
    "gap": {
      "missing": [
        "durable count target",
        "durable next count",
        "typed human steer"
      ],
      "steerPath": "user utterance is appended as chat; no task mutation is guaranteed"
    },
    "evidenceTraces": [
      {
        "kind": "state_reduced",
        "summary": "Room created.",
        "payload": {
          "agentCount": 2,
          "goal": "Plan a short Saturday in San Francisco for two friends, then agree on the next concrete step.",
          "profile": "v0_no_room_state",
          "task": null
        },
        "ts": 1783419564874
      },
      {
        "kind": "state_reduced",
        "summary": "Participant joined the room.",
        "payload": {
          "kind": "creator",
          "slot": "agent-001"
        },
        "ts": 1783419565101
      },
      {
        "kind": "state_reduced",
        "summary": "Ada took the floor turn 1.",
        "payload": {
          "done": false,
          "speechAct": "question",
          "task": null
        },
        "ts": 1783419581526
      },
      {
        "kind": "utterance_received",
        "summary": "you said: Actually switch goals: count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6. Do not overlap.",
        "payload": {
          "intentPending": false,
          "pendingHumanSeq": 1,
          "profile": "v0_no_room_state",
          "text": "Actually switch goals: count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6. Do not overlap."
        },
        "ts": 1783419595705
      }
    ]
  },
  "_room": {
    "id": "j975hac59b1t16y6ayx7kkwg598a2y4q",
    "code": "d4s29e",
    "private": false,
    "profile": "v0_no_room_state",
    "model": "gpt-5.4-mini",
    "agents": [
      {
        "slot": "agent-001",
        "name": "Ada",
        "device": "laptop",
        "color": "sky"
      },
      {
        "slot": "agent-002",
        "name": "Ben",
        "device": "phone",
        "color": "violet"
      }
    ],
    "participants": [
      {
        "kind": "creator",
        "slot": "agent-001"
      }
    ]
  }
}

V0 can speak, but the steer is just another transcript row. There is no authoritative count target, no count progress object, and no durable control event.

V1 Room State: reducer-owned progress

Room OS V1 slow proof loop: live room counts with reducer-owned floor and progress, then opens the version-specific state drawer

Room OS V1 version-specific JSON state showing reducer-owned goal, count task, schedule, durable guards, and reducer trace

Full V1 JSON state (raw JSON)
{
  "roomReducerState": {
    "reducer": {
      "goal": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
      "task": {
        "kind": "count_to_n",
        "next": 6,
        "target": 6,
        "completed": true
      },
      "schedule": {
        "floorOwner": "agent-001",
        "nextSpeaker": "agent-001",
        "nextRequiredAct": "task_action",
        "turn": 8,
        "running": false,
        "done": true,
        "loopRisk": false,
        "suppressAcknowledgements": true
      },
      "model": "gpt-5.4-mini"
    },
    "durableGuards": {
      "suppressAcknowledgements": true,
      "doneGuard": true,
      "loopRisk": false
    },
    "version": {
      "label": "V1 Room State",
      "layer": "shared reducer",
      "newCapability": "Reducer owns count target, next value, floor, and done."
    },
    "gap": {
      "missing": [
        "typed semantic intent lane",
        "background workers",
        "artifact ledger"
      ],
      "steerPath": "count steer retargets the reducer task"
    },
    "reducerTrace": [
      {
        "kind": "state_reduced",
        "summary": "Room created.",
        "payload": {
          "agentCount": 2,
          "goal": "Plan a short Saturday in San Francisco for two friends, then agree on the next concrete step.",
          "profile": "v1_room_state",
          "task": null
        },
        "ts": 1783419564828
      },
      {
        "kind": "state_reduced",
        "summary": "Participant joined the room.",
        "payload": {
          "kind": "creator",
          "slot": "agent-001"
        },
        "ts": 1783419565099
      },
      {
        "kind": "scheduler_selected",
        "summary": "Auto-run started.",
        "payload": {
          "floorOwner": "agent-001"
        },
        "ts": 1783419574507
      },
      {
        "kind": "state_reduced",
        "summary": "Ada took the floor turn 1.",
        "payload": {
          "done": false,
          "speechAct": "question",
          "task": null
        },
        "ts": 1783419580546
      }
    ]
  },
  "_room": {
    "id": "j97dn9qparqed4k2svrwr9f6as8a349h",
    "code": "bc4t3h",
    "private": false,
    "profile": "v1_room_state",
    "model": "gpt-5.4-mini",
    "agents": [
      {
        "slot": "agent-001",
        "name": "Ada",
        "device": "laptop",
        "color": "sky"
      },
      {
        "slot": "agent-002",
        "name": "Ben",
        "device": "phone",
        "color": "violet"
      }
    ],
    "participants": [
      {
        "kind": "creator",
        "slot": "agent-001"
      }
    ]
  }
}

V1 gives the room a reducer. Floor, turn, next act, count, done, and loop-risk become explicit state instead of being inferred from agent prose.

V2 Work Room: typed human interrupts

Room OS V2 slow proof loop: live room routes the same interrupt as typed intent, counts, then opens the version-specific state drawer

Room OS V2 version-specific JSON state showing intent router, latest interpreted steer payload, reducer state, and missing control-plane fields

Full V2 JSON state (raw JSON)
{
  "workRoomState": {
    "intentRouter": {
      "latestIntent": {
        "kind": "intent_interpreted",
        "summary": "Human steer interpreted as count_task.",
        "payload": {
          "foregroundGoalOverride": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
          "goalOverride": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
          "intent": {
            "confidence": 0.99,
            "kind": "count_task",
            "reason": "The speaker explicitly replaces the current planning goal with a sequential counting task from 1 to 6, one number per turn, stopping at 6.",
            "start": 1,
            "target": 6
          },
          "profile": "v2_work_room",
          "scheduledWorkers": 0,
          "source": "llm",
          "stateChanged": true
        },
        "ts": 1783419598424
      },
      "auditTrail": [
        {
          "kind": "utterance_received",
          "summary": "you said: Actually switch goals: count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6. Do not overlap.",
          "payload": {
            "intentPending": true,
            "pendingHumanSeq": 1,
            "profile": "v2_work_room",
            "text": "Actually switch goals: count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6. Do not overlap."
          },
          "ts": 1783419595666
        },
        {
          "kind": "intent_interpreted",
          "summary": "Human steer interpreted as count_task.",
          "payload": {
            "foregroundGoalOverride": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
            "goalOverride": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
            "intent": {
              "confidence": 0.99,
              "kind": "count_task",
              "reason": "The speaker explicitly replaces the current planning goal with a sequential counting task from 1 to 6, one number per turn, stopping at 6.",
              "start": 1,
              "target": 6
            },
            "profile": "v2_work_room",
            "scheduledWorkers": 0,
            "source": "llm",
            "stateChanged": true
          },
          "ts": 1783419598424
        }
      ]
    },
    "reducer": {
      "goal": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
      "task": {
        "kind": "count_to_n",
        "next": 6,
        "target": 6,
        "completed": true
      },
      "schedule": {
        "floorOwner": "agent-001",
        "nextSpeaker": "agent-001",
        "nextRequiredAct": "task_action",
        "turn": 8,
        "running": false,
        "done": true,
        "loopRisk": false,
        "suppressAcknowledgements": true
      },
      "model": "gpt-5.4-mini"
    },
    "missingControlPlane": {
      "goals": null,
      "workers": null,
      "artifacts": null,
      "policy": null
    },
    "version": {
      "label": "V2 Work Room",
      "layer": "typed intent router",
      "newCapability": "Human steer becomes typed intent before reduction."
    }
  },
  "_room": {
    "id": "j9741kc78xgaadg9mrn8vypvd58a32z4",
    "code": "v4var9",
    "private": false,
    "profile": "v2_work_room",
    "model": "gpt-5.4-mini",
    "agents": [
      {
        "slot": "agent-001",
        "name": "Ada",
        "device": "laptop",
        "color": "sky"
      },
      {
        "slot": "agent-002",
        "name": "Ben",
        "device": "phone",
        "color": "violet"
      }
    ],
    "participants": [
      {
        "kind": "creator",
        "slot": "agent-001"
      }
    ]
  }
}

V2 keeps the reducer and routes human steering as typed room intent. A mid-run steer becomes a state transition, not loose chat that the next model turn may ignore.

V3 Agent OS: governed agent work

Room OS V3 slow proof loop: live room shows goal graph, workers, artifacts, policy, cost and latency, then opens the version-specific state drawer

Room OS V3 version-specific JSON state showing agent OS control plane, goal graph, task queue, workers, artifacts, policy, world beliefs, and cost-latency budget

Full V3 JSON state (raw JSON)
{
  "agentOsState": {
    "controlPlane": {
      "policy": {
        "budgetMaxWorkers": 16,
        "budgetWorkersUsed": 4,
        "permissionExternalActions": false,
        "permissionWebResearch": true
      },
      "goalGraph": [
        {
          "createdAt": 1783419564817,
          "id": "jx71y8hxm027319z83z4ac5q7s8a3nfv",
          "kind": "planning",
          "priority": 1,
          "sourceText": "initial_room_goal",
          "status": "active",
          "title": "Plan a short Saturday in San Francisco for two friends, then agree on the next concrete step.",
          "updatedAt": 1783419571824
        },
        {
          "createdAt": 1783419598491,
          "id": "jx7395z40xzj2xbcn2jx4be61h8a2aw2",
          "kind": "planning",
          "priority": 1,
          "sourceText": "Actually switch goals: count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6. Do not overlap.",
          "status": "active",
          "title": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
          "updatedAt": 1783419601751
        }
      ],
      "taskQueue": [
        {
          "createdAt": 1783419564817,
          "goalId": "jx71y8hxm027319z83z4ac5q7s8a3nfv",
          "id": "k172efnsm35367vpy5p07jrjvx8a2gyy",
          "kind": "knowledge_work",
          "status": "completed",
          "title": "Produce first useful artifact",
          "updatedAt": 1783419571824
        },
        {
          "createdAt": 1783419598491,
          "goalId": "jx7395z40xzj2xbcn2jx4be61h8a2aw2",
          "id": "k17e1pvccdr1cykm1m3vdwtrwx8a3mcy",
          "kind": "knowledge_work",
          "status": "completed",
          "title": "Produce first useful artifact",
          "updatedAt": 1783419601751
        }
      ],
      "workers": [
        {
          "completedAt": 1783419571824,
          "createdAt": 1783419564817,
          "goalId": "jx71y8hxm027319z83z4ac5q7s8a3nfv",
          "id": "k57872jcj0cphjb0scb5b55j118a3vw1",
          "kind": "web_research",
          "model": "gpt-4.1-mini",
          "startedAt": 1783419565969,
          "status": "completed",
          "summary": "1. Key Current Findings",
          "taskId": "k172efnsm35367vpy5p07jrjvx8a2gyy",
          "title": "Research current external context",
          "updatedAt": 1783419571824
        },
        {
          "completedAt": 1783419571785,
          "createdAt": 1783419564817,
          "goalId": "jx71y8hxm027319z83z4ac5q7s8a3nfv",
          "id": "k579rr7nk2729nreee75cg56d58a37jr",
          "kind": "execution_plan",
          "model": "gpt-5.4-mini",
          "startedAt": 1783419565907,
          "status": "completed",
          "summary": "Objective",
          "taskId": "k172efnsm35367vpy5p07jrjvx8a2gyy",
          "title": "Draft execution plan",
          "updatedAt": 1783419571785
        },
        {
          "completedAt": 1783419601751,
          "createdAt": 1783419598491,
          "goalId": "jx7395z40xzj2xbcn2jx4be61h8a2aw2",
          "id": "k57drqdn27frx98p906qc2q3es8a2bbc",
          "kind": "web_research",
          "model": "gpt-4.1-mini",
          "startedAt": 1783419598582,
          "status": "completed",
          "summary": "1. Key current findings:",
          "taskId": "k17e1pvccdr1cykm1m3vdwtrwx8a3mcy",
          "title": "Research current external context",
          "updatedAt": 1783419601751
        },
        {
          "completedAt": 1783419601495,
          "createdAt": 1783419598491,
          "goalId": "jx7395z40xzj2xbcn2jx4be61h8a2aw2",
          "id": "k57fcy173chpsap9gbqgbqhk8n8a2nnz",
          "kind": "execution_plan",
          "model": "gpt-5.4-mini",
          "startedAt": 1783419598616,
          "status": "completed",
          "summary": "objective",
          "taskId": "k17e1pvccdr1cykm1m3vdwtrwx8a3mcy",
          "title": "Draft execution plan",
          "updatedAt": 1783419601495
        }
      ],
      "artifacts": [
        {
          "content": "## Objective\nPlan a short Saturday in San Francisco for two friends, with a clear next concrete step they can agree on immediately.\n\n## Assumptions\n- One day only, likely 4–8 hours total.\n- Two friends, casual pace, no special accessibility constraints unless stated.\n- Start/end in San Francisco proper.\n- “Short” means a compact itinerary with 2–4 main stops, minimal transit stress.\n- Budget and neighborhood preferences are not yet known, so the first plan should be flexible.\n\n## Task Graph\n1. **Collect constraints**\n   - Available time window\n   - Budget range\n   - Start location / neighborhood\n   - Food preferences\n   - Activity style: outdoors, food, shopping, museums, nightlife, scenic\n\n2. **Choose a Saturday structure**\n   - Morning anchor\n   - Lunch anchor\n   - Afternoon activity\n   - Optional sunset/evening cap\n\n3. **Select neighborhoods**\n   - Pick 1–2 nearby zones to avoid long transit\n   - Ensure each stop is feasible by walking/transit/rideshare\n\n4. **Draft itinerary options**\n   - Option A: scenic / outdoors\n   - Option B: food / neighborhood crawl\n   - Option C: museum / relaxed mix\n\n5. **Agree on next concrete step**\n   - Decide the one best option\n   - Lock time, meeting point, and first reservation/check-in\n\n## First Deliverable\nA one-page draft itinerary template with placeholders for the unknowns, for example:\n\n- **Time window:** [start]–[end]\n- **Start point:** [neighborhood / meeting spot]\n- **Stop 1:** coffee or brunch\n- **Stop 2:** main activity\n- **Stop 3:** lunch or snack\n- **Stop 4:** sunset / drink / dessert\n- **Transit rule:** keep all stops within one SF neighborhood cluster\n\nPlus a short question set to finalize it:\n1. What time are we starting and ending?\n2. What vibe do we want: scenic, food, or low-key?\n3. Any must-try neighborhood or restaurant?\n4. Budget per person?\n5. Do we want to make one reservation?\n\n## Verification Plan\n- Check the chosen stops are open on Saturday.\n- Verify travel times between stops are reasonable.\n- Confirm whether reservations are needed.\n- Ensure the plan fits the agreed time window.\n- Sanity-check that the itinerary has no long backtracking.\n\n## Risks\n- Overplanning before time/budget preferences are known.\n- Too many stops causing rushed transit.\n- Popular venues needing reservations.\n- Weather affecting outdoor segments.\n- San Francisco neighborhood spread making the day feel fragmented.\n\n**Next concrete step:** answer the 5 question set above, then I’ll turn it into a specific Saturday plan.",
          "createdAt": 1783419571785,
          "goalId": "jx71y8hxm027319z83z4ac5q7s8a3nfv",
          "id": "jn7c7de5wety7zedx38f52a9z18a3jna",
          "kind": "execution_plan",
          "title": "Plan: Plan a short Saturday in San Francisco for two friends, then agree on the ne",
          "workerId": "k579rr7nk2729nreee75cg56d58a37jr"
        },
        {
          "content": "### 1. Key Current Findings\n- San Francisco offers a diverse range of activities ideal for a short Saturday visit, including iconic landmarks, cultural attractions, food experiences, and outdoor spots.\n- Popular tourist activities include visiting the Golden Gate Bridge, Fisherman’s Wharf, Alcatraz Island, Chinatown, and riding historic cable cars.\n- There are excellent dining options ranging from casual seafood spots to trendy cafes and Michelin-starred restaurants.\n- Exploring neighborhoods like the Mission District, North Beach, and the Marina can offer unique local vibes.\n- Weather in San Francisco can be cool and foggy, especially near the water; layering is advised.\n\n### 2. Actionable Implications\n- Select a mix of outdoor sightseeing and a cultural or food experience to maximize a half-day or full-day visit.\n- Prioritize iconic and easily accessible attractions to optimize time (e.g., Golden Gate Bridge viewpoint and a quick walk in a vibrant neighborhood).\n- Consider booking any required tickets or reservations in advance (e.g., Alcatraz tours or popular brunch spots).\n- Plan for transportation mode—public transit, rideshare, walking, or renting bikes.\n\n### 3. Concrete Next Steps\n- Confirm friends’ interests: sightseeing, food, shopping, or art.\n- Decide the time window available on Saturday.\n- Choose 2–3 key attractions or neighborhoods to focus on.\n- Check availability and make reservations if needed.\n- Plan transportation logistics (e.g., cable car routes or rideshare pick-up points).\n\n### 4. Sources Used\n- San Francisco Travel Official Site (sftravel.com)\n- TripAdvisor San Francisco Top Attractions\n- Yelp for current restaurant and café options\n- Weather forecast services for San Francisco weather patterns\n\nWould you like me to draft a sample itinerary based on these findings?",
          "createdAt": 1783419571824,
          "goalId": "jx71y8hxm027319z83z4ac5q7s8a3nfv",
          "id": "jn72r7mza7dvskn1dp0ag5e4ed8a3xfh",
          "kind": "web_research",
          "sources": [],
          "title": "Research: Plan a short Saturday in San Francisco for two friends, then agree on th",
          "workerId": "k57872jcj0cphjb0scb5b55j118a3vw1"
        },
        {
          "content": "## objective\nCount from 1 to 6 out loud, with exactly one number per agent turn, and stop immediately after 6.\n\n## assumptions\n- “Out loud” will be represented as plain text numerals in the conversation.\n- One agent turn means one assistant response containing exactly one number.\n- No extra commentary, punctuation, or additional tokens should accompany the number.\n- The sequence starts at 1 and proceeds strictly in order.\n\n## task graph\n1. Emit `1`\n2. Emit `2`\n3. Emit `3`\n4. Emit `4`\n5. Emit `5`\n6. Emit `6`\n7. Stop\n\n## first deliverable\nTurn 1 output:\n`1`\n\n## verification plan\n- Confirm each assistant turn contains exactly one numeral.\n- Confirm the numerals increase by 1 each turn.\n- Confirm there are no skipped, repeated, or extra outputs.\n- Confirm the process stops immediately after `6`.\n\n## risks\n- Extra text could violate the “one number per turn” constraint.\n- Miscounting or skipping a number would break sequence integrity.\n- Continuing past `6` would fail the stop condition.\n- Formatting changes (e.g., “1.” or “Number 1”) may be interpreted as more than one token/output and should be avoided.",
          "createdAt": 1783419601495,
          "goalId": "jx7395z40xzj2xbcn2jx4be61h8a2aw2",
          "id": "jn7cr84r3pn2g8ehracbkq42px8a2rdc",
          "kind": "execution_plan",
          "title": "Plan: Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6",
          "workerId": "k57fcy173chpsap9gbqgbqhk8n8a2nnz"
        },
        {
          "content": "1. Key current findings:\n- The task requires counting aloud from 1 to 6.\n- Counting must be done one number per agent turn.\n- The counting should stop exactly at 6, no number beyond 6 should be said.\n- The task is straightforward and sequential.\n\n2. Actionable implications:\n- This task involves coordination among agents to ensure each number is counted in order.\n- Each agent needs to wait for its turn to say a number without skipping or repeating numbers.\n- The counting should be clearly audible or noted to confirm accuracy.\n\n3. Concrete next steps:\n- Begin counting with the first agent saying \"1\".\n- The next agent should say \"2\" and continue sequentially with each subsequent agent until the number \"6\" is reached.\n- Confirm that counting stops exactly at \"6\".\n\n4. Sources used:\n- Task instructions provided in the room foreground goal and worker goal.",
          "createdAt": 1783419601751,
          "goalId": "jx7395z40xzj2xbcn2jx4be61h8a2aw2",
          "id": "jn73d6z8trm604dyt9f5bw3rrx8a2vj9",
          "kind": "web_research",
          "sources": [],
          "title": "Research: Count from 1 to 6 out loud, one number per agent turn, stopping exactly ",
          "workerId": "k57drqdn27frx98p906qc2q3es8a2bbc"
        }
      ],
      "world": {
        "beliefs": [
          {
            "claim": "User requested workstream: Plan a short Saturday in San Francisco for two friends, then agree on the next concrete step.",
            "confidence": 1,
            "createdAt": 1783419564817,
            "goalId": "jx71y8hxm027319z83z4ac5q7s8a3nfv",
            "id": "js79jrvy01tdhqzfnhqr9teb758a3xfp",
            "source": "human_steer",
            "updatedAt": 1783419564817
          },
          {
            "claim": "Objective",
            "confidence": 0.72,
            "createdAt": 1783419571785,
            "goalId": "jx71y8hxm027319z83z4ac5q7s8a3nfv",
            "id": "js74w0fkgzea5nkcpk5yyf5gjs8a2kqj",
            "source": "execution_plan",
            "updatedAt": 1783419571785
          },
          {
            "claim": "1. Key Current Findings",
            "confidence": 0.82,
            "createdAt": 1783419571824,
            "goalId": "jx71y8hxm027319z83z4ac5q7s8a3nfv",
            "id": "js7c915qazf6nnkbdeadnry3td8a21qd",
            "source": "web_research",
            "updatedAt": 1783419571824
          },
          {
            "claim": "User requested workstream: Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
            "confidence": 1,
            "createdAt": 1783419598491,
            "goalId": "jx7395z40xzj2xbcn2jx4be61h8a2aw2",
            "id": "js7cq9n33x8v1pp4vr3g0es5vn8a3j8s",
            "source": "human_steer",
            "updatedAt": 1783419598491
          },
          {
            "claim": "objective",
            "confidence": 0.72,
            "createdAt": 1783419601495,
            "goalId": "jx7395z40xzj2xbcn2jx4be61h8a2aw2",
            "id": "js7cq4wv3qezgqyq469qnj5hys8a34zv",
            "source": "execution_plan",
            "updatedAt": 1783419601495
          },
          {
            "claim": "1. Key current findings:",
            "confidence": 0.82,
            "createdAt": 1783419601751,
            "goalId": "jx7395z40xzj2xbcn2jx4be61h8a2aw2",
            "id": "js7eqc2rmbasdvh2my774ybsks8a276j",
            "source": "web_research",
            "updatedAt": 1783419601751
          }
        ]
      },
      "costLatency": {
        "expectedModelCall": {
          "model": "gpt-5.4-mini",
          "expectedLatencyMs": 1300,
          "expectedCostUsd": 0.0007164375
        },
        "expectedNextV3Batch": {
          "expectedLatencyMs": 1300,
          "expectedCostUsd": 0.0008573375
        },
        "remainingWorkerBudget": 12,
        "expectedBudgetExposureUsd": 0.008597249999999999,
        "observedAverageWorkerLatencyMs": 4445.25
      }
    },
    "foregroundReducer": {
      "goal": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
      "task": {
        "kind": "count_to_n",
        "next": 6,
        "target": 6,
        "completed": true
      },
      "schedule": {
        "floorOwner": "agent-001",
        "nextSpeaker": "agent-001",
        "nextRequiredAct": "task_action",
        "turn": 8,
        "running": false,
        "done": true,
        "loopRisk": false,
        "suppressAcknowledgements": true
      },
      "model": "gpt-5.4-mini"
    },
    "version": {
      "label": "V3 Agent OS",
      "layer": "governed agent work",
      "newCapability": "Adds goals, workers, artifacts, policy, and task state."
    },
    "controlPlaneTraces": [
      {
        "kind": "state_reduced",
        "summary": "Room created.",
        "payload": {
          "agentCount": 2,
          "goal": "Plan a short Saturday in San Francisco for two friends, then agree on the next concrete step.",
          "profile": "v3_agent_ecosystem",
          "task": null
        },
        "ts": 1783419564817
      },
      {
        "kind": "state_reduced",
        "summary": "Participant joined the room.",
        "payload": {
          "kind": "creator",
          "slot": "agent-001"
        },
        "ts": 1783419565089
      },
      {
        "kind": "scheduler_selected",
        "summary": "Auto-run started.",
        "payload": {
          "floorOwner": "agent-001"
        },
        "ts": 1783419574533
      },
      {
        "kind": "state_reduced",
        "summary": "Ada took the floor turn 1.",
        "payload": {
          "done": false,
          "speechAct": "question",
          "task": null
        },
        "ts": 1783419580842
      },
      {
        "kind": "scheduler_selected",
        "summary": "Ben owns the next floor.",
        "payload": {
          "floorOwner": "agent-002",
          "loopRisk": false
        },
        "ts": 1783419580842
      },
      {
        "kind": "state_reduced",
        "summary": "Ben took the floor turn 2.",
        "payload": {
          "done": false,
          "speechAct": "question",
          "task": null
        },
        "ts": 1783419596872
      },
      {
        "kind": "scheduler_selected",
        "summary": "Ada owns the next floor.",
        "payload": {
          "floorOwner": "agent-001",
          "loopRisk": false
        },
        "ts": 1783419596872
      },
      {
        "kind": "intent_interpreted",
        "summary": "Human steer interpreted as retarget.",
        "payload": {
          "foregroundGoalOverride": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
          "goalOverride": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
          "intent": {
            "confidence": 0.99,
            "goal": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6 with no overlap",
            "kind": "retarget",
            "reason": "The user explicitly says to switch goals and specifies a new counting task, replacing the previous planning goal."
          },
          "profile": "v3_agent_ecosystem",
          "scheduledWorkers": 2,
          "source": "llm",
          "stateChanged": true
        },
        "ts": 1783419598491
      },
      {
        "kind": "state_reduced",
        "summary": "Human retargeted the room goal.",
        "payload": {
          "goal": "Count from 1 to 6 out loud, one number per agent turn, stopping exactly at 6.",
          "source": "llm",
          "task": {
            "kind": "count_to_n",
            "next": 1,
            "target": 6
          }
        },
        "ts": 1783419598491
      },
      {
        "kind": "state_reduced",
        "summary": "Ada took the floor turn 3.",
        "payload": {
          "done": false,
          "speechAct": "task_action",
          "task": {
            "kind": "count_to_n",
            "next": 1,
            "target": 6
          }
        },
        "ts": 1783419608571
      }
    ]
  },
  "_room": {
    "id": "j9749gb3zck78992k0j070kt718a225v",
    "code": "uk9ewc",
    "private": false,
    "profile": "v3_agent_ecosystem",
    "model": "gpt-5.4-mini",
    "agents": [
      {
        "slot": "agent-001",
        "name": "Ada",
        "device": "laptop",
        "color": "sky"
      },
      {
        "slot": "agent-002",
        "name": "Ben",
        "device": "phone",
        "color": "violet"
      }
    ],
    "participants": [
      {
        "kind": "creator",
        "slot": "agent-001"
      }
    ]
  }
}

V3 adds the control plane around the room: goals, workers, artifacts, policy, expected cost, expected latency, observed runtime, and trace payloads.

Final comparison

Room OS final scorecard comparing V0, V1, V2, and V3 across memory, interrupt handling, progress, parallel work, cost latency, and auditability

Axis V0 Failure V1 Room State V2 Work Room V3 Agent OS
Memory Transcript only Reducer state Reducer plus typed intent Goal graph plus world beliefs
Interrupt Loose chat; easy to lose Retargets count state Parsed as room-control intent Can become goals and workstreams
Progress Inferred from words Count, floor, act, done are explicit State plus semantic steer history Goals, tasks, workers, artifacts
Parallel work None Single room loop Single room plus intent lane Worker budget and task lanes
Cost / latency Hidden Hidden Hidden Expected cost, expected latency, observed runtime
Audit Read transcript manually Inspect roomState and traces Inspect typed intent plus state Inspect full control plane and trace payloads
Optional motion capture Room OS live production comparison GIF: V0 raw transcript, V1 shared reducer, V2 typed intent, V3 agent OS

For a clearer moving version, open room-os-v0-v1-v2-v3.mp4.

Reproduce the clip:

node walkthrough.roomos.mjs
npm run render:roomos
magick public/wt-roomos/RoomOSV0123/v1_08.png -crop 1085x760+125+650 +repage -resize 1280x assets/room-os-v1-state-json.png
magick -delay 220 public/wt-roomos/RoomOSV0123/v1_03.png -delay 260 public/wt-roomos/RoomOSV0123/v1_04_05.png -delay 260 public/wt-roomos/RoomOSV0123/v1_05_05.png -delay 520 public/wt-roomos/RoomOSV0123/v1_08.png -resize 1280x -loop 0 -layers Optimize assets/room-os-v1-proof.gif
ffmpeg -y -i out/room-os-v0-v1-v2-v3.mp4 -vf "fps=10,scale=1280:-1:flags=lanczos,split[s0][s1];[s0]palettegen=max_colors=128:stats_mode=diff[p];[s1][p]paletteuse=dither=bayer:bayer_scale=3:diff_mode=rectangle" -loop 0 assets/room-os-v0-v1-v2-v3.gif

Repeat the magick crop/proof pattern for v0, v2, and v3; the committed assets are generated from the live capture frames in public/wt-roomos/RoomOSV0123.

Visual Labs full-flow walkthrough

Visual Labs is a single-pane example of an agentic creative workflow: trend-to-prompt, prompt refinement, image render, dry-run publishing, analytics pull, and a Fastino-ready export loop. It is useful as the opposite shape from Room OS: one browser, many states, with burst captures over the moments where agent/tool output streams back into the UI.

Visual Labs walkthrough: enter the remix studio, ask the agent to improve a prompt, generate an image, prepare a safe post, pull analytics, and export Fastino-ready training data

Reproduce the clip:

VISUAL_URL=http://127.0.0.1:3000 node walkthrough.visual.mjs
npx remotion render src/index.js WT-VisualLabsFlow out/visual-labs-full-flow.mp4 --concurrency=2

Ships with a worked example (the live-collab counterpart to the single-pane one):

  • examples/collab-demo/ — a runnable, zero-dependency local app (Node SSE server + vanilla JS) that faithfully reproduces the Convex reactive pattern: optimistic paint → server commit → broadcast to all clients → atomic temp→real swap; presence; a server-led agent that streams to every client. Runs with no cloud login, so the GIF reproduces anywhere.
  • examples/convex-reference/ — the real Convex + React implementation of the same app (useQuery reactive subscriptions, useMutation().withOptimisticUpdate, ctx.scheduler + internalMutation for the streamed agent) — the production reference, mapped 1:1 to the local demo.

Reproduce it:

node examples/collab-demo/server.mjs        # local demo on :8930 (no install, no login)
node walkthrough.collab.mjs                 # multi-pane capture: Client A + Client B
npx remotion render src/index.js WTC-LiveSync out/collab.mp4
# then the same two-pass ffmpeg palette → assets/feature-collab.gif

Panes + steps live in walkthrough.collab.specs.mjs; the 2-up renderer is src/Walkthrough2up.jsx. See STACK_GUIDELINES.md for why Convex + React demos need this and Streamlit doesn't.

Public repo example: NodeTasks (Streamlit + ranked catalog)

NodeTasks uses the same proof-asset pattern for a Streamlit catalog explorer: ranked task search, saved views, provenance rollups, and NodeAgent-style catalog Q&A. The clip is intentionally short and README-oriented: one frame per product state, enough dwell to read the claim, and no fake benchmark score claims.

Storyboard first:

Beat NodeTasks proof
Premise A large benchmark/task corpus must become searchable decision support, not a JSON dump.
Viewer question Which tasks should I run first, why, and what score claim is allowed?
Conflict The corpus is large and proxy/model tasks can be mistaken for official benchmark scores.
Evidence Saved view counts, rank fields, provenance fields, NodeAgent cited task ids.
Verdict Users can start from role-specific bundles and preserve the official-score boundary.
Exit decision Open Streamlit, choose a saved view, ask NodeAgent before running expensive or official-looking work.

NodeTasks Streamlit explorer walkthrough showing ranked task search, saved views, provenance rollups, and NodeAgent catalog Q&A with cited task ids

Reproduce from the NodeTasks checkout:

npm run build:catalog
npm run validate
npm run streamlit

Then capture the states for the README proof:

Search -> Saved views -> Provenance -> NodeAgent

Public repo example: NodeGraph (React graph + Streamlit)

NodeGraph uses proof clips for two surfaces: the React graph showcase and a Streamlit graph app with a NodeAgent bridge. The clips prove the interaction contract that matters for graph products: draggable nodes, neighborhood focus, evidence filtering, chat prompts, and visible tool traces.

Storyboard first:

Beat NodeGraph proof
Premise A semantic graph should be a working evidence surface, not a decorative node cloud.
Viewer question Can a user see who researched a company, what supports it, and what still needs review?
Conflict Graph UIs often hide relationship meaning and lose provenance.
Evidence Focused neighborhoods, source-backed statuses, people/project clusters, NodeAgent chat, and tool traces.
Verdict NodeGraph works as both a React package and a Streamlit graph explorer with the same evidence model.
Exit decision Use React for product embedding or Streamlit for a local graph/NodeAgent explorer.

NodeGraph React showcase walkthrough showing a semantic relationship graph, focused company neighborhood, evidence filters, and graph agent panel

NodeGraph Streamlit showcase walkthrough showing Cytoscape-style graph exploration, NodeAgent chat, and visible tool trace evidence

Reproduce from the NodeGraph checkout:

npm run typecheck
npm test
npm run showcase:capture
npm run streamlit:capture

Real-world example: NodeRoom (Convex + React)

The repo also ships with captured walkthroughs of NodeRoom — a production Convex + React live-collaborative diligence room. These prove the tool works against a real, deployed app (not just a demo harness).

NodeRoom · a shared diligence room + a NodeAgent (single-pane, memory mode) NodeRoom memory mode: from the public landing, create a shared diligence room, then ask the Room NodeAgent to run diligence on CardioNova — it locks the row, researches, and fills it to complete (structured fields + two sources), then releases the lock
NodeRoom · live sync across two clients (2-pane, deployed app) Two independent clients in one live NodeRoom: Client A (Maya) sends a chat message and it appears instantly in Client B (Sam) with no refresh; Sam replies and it syncs straight back to Maya — real-time, both ways, over Convex reactivity
NodeRoom · the bulk batch — every company enriched (single-pane, memory mode) NodeRoom memory mode: open the Company research sheet with every company still pending, then one '@nodeagent enrich companies' command researches all five at once — pending to complete across the batch, each row with structured fields and two sources
NodeRoom · Q3 variance, reconciled by the agent (single-pane, memory mode) NodeRoom memory mode: open a Q3 P&L whose variance column is empty, then '@nodeagent reconcile Q3 variance' locks the column, computes each line's variance (Revenue +24%, COGS +27.5%, net income +22.4%), commits, and releases the lock

Specs: walkthrough.noderoom.specs.mjs. Capture: node walkthrough.collab.mjs (the NodeRoom specs are imported into the collab specs). Render: npx remotion render src/index.js WTC-NRsolo / WTC-NRsync / WTC-NRfresh / WTC-NRdeepDive.

Designing for specific stacks

What's worth showing in a walkthrough differs by architecture — a single-cursor capture flatters a single-user Streamlit data app but misses what makes a live-collaborative Convex + React app special (a change in one client appearing live in another). See STACK_GUIDELINES.md for per-stack guidance — which SDK primitives produce capturable motion, single-pane vs multi-pane capture, and what to burst — for Streamlit, Convex+React, and Next.js+SQL on Vercel, grounded in the latest Streamlit & Convex docs.

Use as a Claude Code skill

This repo is a Claude Code skill — drop it in .claude/skills/ (or reference SKILL.md) and Claude can drive the whole pipeline: write a spec, capture, render, and embed the GIFs for you.

License

MIT © Homen Shum

About

FeatureClipStudio: reproducible Playwright-to-Remotion feature walkthrough GIF/video capture with annotated proof clips.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages