A fully typesafe, saga-style script framework for TypeScript, with a live terminal UI.
- Scripts are a collection of phases.
- Phases are a collection of steps.
- Steps have a
handlerand an optionalrollbackthat fires when a later step fails. - Whatever a handler returns is merged into a typed context that every later step can read.
- A step can
cleankeys it is done with, and arollbackdeclares the keys it needs. - Zero runtime dependencies. Live rendering on a TTY, plain line output in CI.
const result = await new Script<{ service: string }>({ name: "deploy" })
.addPhase("Validation")
.addStep({
name: "resolve commit",
handler: async ({ input, status }) => {
status("querying git");
return { sha: await git.head(input.service) };
},
})
.addPhase("Release")
.addStep({
name: "upload artifact",
handler: async ({ ctx, progress }) => {
// ^ ctx.sha is typed here
const bar = progress({ total: 100, label: "uploading" });
return { uploadId: await upload(ctx.sha, bar) };
},
rollbackKeys: ["sha"],
rollback: async ({ ctx, output }) => cdn.delete(output.uploadId, ctx.sha),
// ^ { sha: string } ^ typed as { uploadId: string }
})
.run({ service: "api" });
if (result.ok) result.ctx.uploadId; // string ▌ deploy
▌ Build, upload and release a service
✔ Validation 739ms
✔ resolve commit 222ms
✔ check permissions 517ms
⠸ Build 17.7s
✔ install dependencies (428 packages) 14.4s
⠸ compile bundle › transform 3.3s
▕██████████████░░░░░░░░░░▏ 58% 412/710 src/router.ts
✔ typecheck
⠸ transform
○ minify
○ write manifest
○ Release
○ upload artifact
○ shift traffic
On failure:
─────────────────────────────────────────────────────────────────────────────
✖ Failed at Release › shift traffic after 21.4s
health check failed: 3/5 pods unhealthy
↺ Rolled back 2 steps
Run npm run example and npm run example:fail to see both live.
The context type parameter is inferred and never written by hand. Each addStep
whose handler returns an object widens it:
new Script<{ id: string }>() // Ctx = {}
.addStep({ handler: () => ({ token: "x" }) }) // Ctx = { token: string }
.addStep({ handler: ({ ctx }) => ({ n: ctx.token.length }) })
// Ctx = { token: string; n: number }Reaching for a key a previous step did not produce is a compile error, as is
reading the wrong shape in a rollback. A handler that returns nothing leaves
the context unchanged.
If a step produces — or receives — data that later steps do not need, listing
those keys in clean deletes them from the context at runtime and removes
them from the type that flows on:
new Script<{ password: string; email: string }>({ name: "signup" })
.addStep({
name: "read input",
handler: ({ input }) => ({ ...input }),
})
.addStep({
name: "hash",
handler: async ({ ctx }) => ({ hash: await hashIt(ctx.password) }),
clean: ["password"], // not needed past this point
})
.addStep({
name: "persist",
handler: ({ ctx }) => {
ctx.hash; // string
ctx.email; // string
ctx.password; // compile error — cleaned away
},
});Rules:
- Keys are checked against the context as it stands entering that step — an
unknown key is a compile error, and the editor autocompletes the valid ones.
A step cannot clean a key it produces itself; that key does not exist yet
when
cleanis read. - A key reserved by any step's
rollbackKeyscan never be cleaned, by that step or any later one. Trying to throwsStepDefinitionErrorfromaddStep, while the script is being built, before anything runs. cleanis part of the declared shape, not of the work, so it applies even when the step is skipped bywhen— runtime and types never disagree.- A step's own
outputis kept separately and is never cleaned, so itsrollbackstill sees everything the handler returned.
In doesn't have to be written by hand either. defineInput takes any
Standard Schema-compliant validator — Zod,
Valibot, ArkType, Effect Schema, or a hand-rolled one — infers In from its
output type, and validates whatever is passed to run() before any phase
executes, throwing SchemaValidationError on failure:
const deploy = new Script({ name: "deploy" })
.defineInput(z.object({ service: z.string() }))
.addStep({
name: "resolve commit",
handler: async ({ input }) => ({ sha: await git.head(input.service) }),
// ^ input.service: string
});Call it first, right after construction — any addStep/addPhase added
before it still sees the previous In.
When a step fails, every step that already succeeded is compensated in reverse order. The step that failed is not compensated — it never completed.
rollback option |
behaviour |
|---|---|
"all" (default) |
unwind every completed step in the script |
"phase" |
unwind only the phase that failed |
"none" |
leave state as-is |
A rollback that throws is recorded in result.rollbacks and surfaced in the
summary; it never masks the original error, and the remaining rollbacks still
run.
A rollback always gets its own step's output in full. It gets no context at
all by default — it has to ask, via rollbackKeys. Asking does two things: it
narrows ctx inside the rollback to exactly those keys, and it reserves them,
so neither that step nor any later one can clean them away. Whatever the
rollback declared is therefore guaranteed to still be there if it ever runs.
new Script<{ accountId: string; amount: number }>({ name: "charge" })
.addStep({ name: "read input", handler: ({ input }) => ({ ...input }) })
.addStep({
name: "reserve funds",
handler: async ({ ctx }) => ({ reservationId: await reserve(ctx.accountId, ctx.amount) }),
rollbackKeys: ["accountId", "reservationId"],
rollback: async ({ ctx }) => {
// ctx is { accountId: string; reservationId: string } — and nothing else
await release(ctx.accountId, ctx.reservationId);
},
})
.addStep({
name: "charge card",
handler: () => {
throw new Error("payment gateway timeout");
},
});rollbackKeys accepts the step's incoming context keys and its own output keys;
anything else is a compile error. Cleaning a reserved key is a compile error
too, and throws StepDefinitionError at build time:
script
.addStep({ name: "one", handler: () => {}, rollbackKeys: ["a"], rollback: async () => {} })
.addStep({ name: "two", handler: () => {}, clean: ["a"] });
// ^ StepDefinitionError: reserved by rollbackKeysEvery handler receives one object:
input |
the value passed to run() |
ctx |
everything earlier steps produced, fully typed |
signal |
aborts on timeout, Ctrl-C, or an external signal |
attempt |
1-based, useful with retry |
log / info / warn / error / success |
a line, placed per logPlacement (default: permanent, above the live frame) |
status(text) |
transient one-liner beside the step, cleared when it settles |
note(text) |
annotates the step title — (400 rows found) — and stays |
progress({ total, label }) |
a progress bar → update / increment / setTotal / done |
task(label) |
one nested checklist item → succeed / fail / skip |
tasks([...] as const) |
a whole checklist, keyed for typed lookup |
rollback receives input, signal, phase, step, log, status, note
and progress, plus output (that step's own return value) and error (what
triggered the unwind). Its ctx holds only the keys named in rollbackKeys.
Both write next to the step name, and they differ in lifetime. status is
scratch space for a step in flight — it is wiped the moment the step settles.
note is part of the title: it is written once and stays on the finished line,
which is where a count, a size, or an id belongs.
.addStep({
name: "query database",
handler: async ({ status, note }) => {
status("scanning"); // ⠸ query database › scanning
const rows = await db.all(sql);
note(`${rows.length} rows found`);
return { rows };
},
})
// ✔ query database (400 rows found) 1.2sCalling note again replaces the text; note("") removes it. Both renderers
show it — the live frame in the title, the CI renderer on the step's line.
By default, log/info/warn/error/success write a permanent line above
the whole frame — the complete record for the run, in order. logPlacement
moves them somewhere more local instead:
new Script({ name: "checkout", logPlacement: "step" })logPlacement |
where it lands |
|---|---|
"scrollback" (default) |
a permanent line above the frame — the full run, in order |
"step" |
nested under the step that logged it, beside its tasks and progress bar |
"bottom" |
a rolling tail of the most recent lines, below the whole phase tree |
✖ Main 320ms
✔ reserve inventory 320ms
• checked warehouse A ← "step": nested where it happened
• checked warehouse B
✖ boom 0ms
boom
"step" and "bottom" both keep only the most recent few lines — they trade
completeness for locality, which is exactly what makes a rollback's log()
land next to the step it is undoing instead of scrolling past at the top.
Reach for "scrollback" when the log is the thing you need to still have once
the run is over. Only the live renderer honours this: the plain/CI renderer is
already sequential, so every placement just prints each line immediately, in
order, which is already both complete and in place.
Run npm run example:log-placement to see the same script three times, once
per placement.
.addStep({
name: "publish",
description: "push the tag",
when: ({ input, ctx }) => input.env === "production",
retry: { attempts: 3, delayMs: (attempt) => attempt * 250, retryIf: isTransient },
timeoutMs: 30_000,
clean: ["draftId"], // drop from the context once this step settles
rollbackKeys: ["tagName"], // the context keys `rollback` needs, and reserves
handler,
rollback,
})when returning false marks the step skipped — its rollback never runs.
Phases accept when too, which skips all their steps.
A phase or a step can reuse what it produced on an earlier run. Point it at a store and it stops repeating work:
import { Script, fileStore } from "@michaelrwalker/stagehand";
const cache = fileStore("./.stagehand-cache.json");
new Script<Input>({ name: "deploy" })
.addPhase("Build", { cache })
.addStep({ name: "install", handler: … })
.addStep({ name: "compile", handler: … })
.addPhase("Release")
.addStep({ name: "warm CDN", cache, handler: … })On a hit the work is skipped and the stored value is merged into the context — downstream steps see the same keys, at the same types, either way.
What gets stored. A phase stores its delta: the keys its steps
contributed, minus anything they cleaned. Keys from earlier phases are never
part of it, so a hit can't overwrite a value this run just computed. A step
stores exactly what its handler returned.
There is no key. The store is the identity. Say when an entry stops being
good with stale, or delete the file:
.addPhase("Build", {
cache: {
store: cache,
stale: ({ value, ctx, ageMs }) => ageMs > 3_600_000 || value.sha !== ctx.sha,
schema: BuildSchema, // optional; a value that no longer fits is a miss
},
})stale and schema both see the context as it stands when the phase is
reached, so ctx is typed with no annotation. value arrives as unknown —
annotate the parameter, or hand over a schema and let it narrow. A stored
value that fails its schema is treated as a miss, not an error, which is
what stops a cache written by an older version of the script from feeding the
wrong shape into the context.
Rollback. Cached work never ran, so its rollback can't fire during a
later unwind — the side effect belongs to the run that wrote the entry and was
never undone. For the same reason a failure downstream leaves entries alone.
Per run, without touching the script:
await deploy.run(input, { cache: argv.includes("--no-cache") ? "off" : "on" });"on" (default) reads and writes · "off" ignores caching entirely ·
"refresh" runs everything and overwrites · "read-only" uses hits but never
writes.
Every handler and rollback gets a cache handle, typed to the cached phases
declared before it:
.addPhase("Fetch", { cache: { store } })
.addStep({ name: "pull orders", handler: … }) // → { orders }
.addPhase("Adjust")
.addStep({
name: "restock",
handler: async ({ ctx, cache }) => {
await cache.clear("Fetch"); // next run refetches
await cache.write("Fetch", { orders: ctx.orders }, { keepAge: true });
const age = await cache.ageOf("Fetch");
},
})Slot names are checked, not stringly-typed. An uncached phase isn't a slot at
all, and as renames a mounted routine's slots along with its phases — so the
moment you write .use(pullChannel, { as: "Amazon" }), every
cache.clear("Fetch") in the script goes red and offers "Amazon / Fetch"
instead. That's the failure a hand-written string can't catch: it would keep
compiling and silently address nothing.
Invalidation flows backwards only — a later step can drop an earlier phase's entry, not the reverse.
write restamps the entry as written now unless you pass keepAge. Reach
for it when you are correcting a value rather than refreshing it: restamping
silently buys another full TTL for data that is exactly as old as it was a
moment ago, which is the opposite of what you want when the cache is there to
keep you inside a rate limit.
Values are typed too. A slot with a schema reads back as the schema's
output, because that is the one thing actually checked against the stored
JSON. Without a schema the phase's own delta stands in, which is convenient and
true right up until an entry outlives the code that wrote it. For that case
read returns the value behind a guard:
CacheShapeError: Read "orders.0.date_fixed" from cache slot "Fetch", but the
stored value has no such property. The entry was probably written before this
script's shape changed — drop it with cache.clear("Fetch"), or declare a schema
on the slot to have mismatches rejected on read.
Only reads are trapped; assigning a new property is how you'd patch an entry.
Missing array indices, JSON.stringify, await and iteration all behave
normally. Pass { raw: true } to opt out — do that before putting the value
into the context, or the guard travels with it into later steps.
Stores are three methods, so Redis or S3 is a dozen lines:
interface CacheStore {
read(slot: string): Awaitable<CachedEntry | undefined>;
write(slot: string, entry: CachedEntry): Awaitable<void>;
clear(slot: string): Awaitable<void>;
}slot is derived from the phase or step name so several phases can share one
store; scripts never write it. fileStore(path) keeps every slot in one JSON
file — a missing, unreadable or corrupt file all read as no cache at all, so
deleting it is always a valid reset. memoryStore() lasts for the process.
A store that throws — unreadable file, full disk, a stale predicate with a bug
— logs a warning and does the work. Caching never decides whether a run passes.
new Script<Input>({
name: "deploy",
description: "Build, upload and release a service",
rollback: "all", // "all" | "phase" | "none"
throwOnError: false, // true to throw instead of returning ok: false
plain: undefined, // force the CI renderer; default auto-detects a TTY
silent: false,
handleSignals: true, // Ctrl-C aborts and compensates; a second one gives up
logPlacement: "scrollback", // "scrollback" | "step" | "bottom"
})type RunResult<Ctx> =
| { ok: true; status: "success"; ctx: Ctx; durationMs: number; steps: StepReport[] }
| { ok: false; status: "failed" | "aborted"; error: unknown;
failedAt: { phase: string; step: string } | null;
ctx: Partial<Ctx>; // what did get produced
durationMs: number; steps: StepReport[]; rollbacks: RollbackReport[] }stepFor binds the input and context a step expects while keeping inference:
export const verifyId = stepFor<Input, { user: User }>()({
name: "verify id",
handler: async ({ ctx }) => ({ verified: ctx.user.id }),
rollbackKeys: ["verified"],
rollback: async ({ ctx }) => unverify(ctx.verified),
});rollbackKeys narrows the rollback here exactly as it does inline, and the keys
stay reserved once the step is handed to addStep.
addStep is where the requirement is enforced: TypeScript checks that the
script has actually produced Ctx by that point. Dropping a step that needs
{ user: User } into a script that has not loaded a user yet is a compile
error, not a runtime undefined.
Writing the next step in a chain means restating what the previous one returns.
WithStepFor<Step, Rest> computes that instead — the step's output merged over
whatever else the context already holds — and nests for longer chains:
export const loadRules = stepFor<
Input,
WithStepFor<typeof logIntoDb, { conditions: Array<{ name: string }> }>
>()({
name: "load rules",
handler: ({ ctx }) => ({ rules: ctx.conn.query(ctx.conditions) }),
});
// longer chains nest, innermost first
WithStepFor<typeof second, WithStepFor<typeof first, { conditions: string[] }>>routineFor does for whole phases what stepFor does for one step. Declare a
fragment once, mount it wherever it fits:
// routines/channel.ts
export const pullChannel = routineFor<{ channel: string; since: string }>()(
"pull channel",
(script) =>
script
.addPhase("Fetch", { cache: { store: cache, stale: … } })
.addStep({ name: "authenticate", handler: … })
.addStep({ name: "pull orders", handler: … })
.addPhase("Normalize")
.addStep({ name: "dedupe and total", handler: … }),
);new Script<Input>({ name: "channel sync" })
.use(pullChannel)
.addPhase("Validation")
.addStep({ name: "check totals", handler: ({ ctx }) => … }) // ctx.orderCount, ctx.gross
// months later, a different job entirely
new Script<ReportInput>({ name: "monthly report" })
.use(pullChannel)
.addPhase("Report")
.addStep({ name: "build summary", handler: ({ ctx }) => … })The two parameters on routineFor<In, Ctx>() are the minimum the fragment
needs — the input it reads, and the context it expects to already exist. Both
are checked at the use() call:
Property 'hash' is missing in type '{}' but required in type '{ hash: string; }'.
A plain Script mounts too, which is what keeps a pipeline runnable on its
own — pipeline.run(input) today, .use(pipeline) tomorrow. Its own
ScriptOptions (rollback mode, logPlacement, silent) are ignored in favour
of the host's; only its phases and steps come across.
Everything travels with the fragment. Phase-level when and cache,
step-level retry, clean, rollback — a spliced rollback compensates during
the host's unwind, and keys it reserved through rollbackKeys stay reserved in
the host, so a later clean there still can't delete them.
That combines well with caching: mount a routine whose phase caches, and the second script to mount it gets the data the first one pulled.
› Fetch
⊙ authenticate cached
⊙ pull orders cached
› Report
✔ build summary (avg 24.50) 263ms
A mount can feed its fragment something other than the host's input — which is what lets one routine serve two storefronts, each with its own credentials:
new Script<{ since: string; amazonKey: string; shopifyKey: string }>({ name: "all channels" })
.use(pullChannel, {
as: "Amazon",
input: ({ input }) => ({ channel: "amazon", since: input.since, apiKey: input.amazonKey }),
})
.use(pullChannel, {
as: "Shopify",
input: ({ input }) => ({ channel: "shopify", since: input.since, apiKey: input.shopifyKey }),
})input takes a fixed value or a function of { input, ctx }, and may be
async. It is resolved once per mount, when the mount is reached, and every
phase and step of that fragment sees the result — including its when, its
cache, and its rollbacks during an unwind.
With a mapper the host no longer has to match the routine's input shape at all; it only has to produce it. TypeScript checks the return value:
Property 'apiKey' is missing in type '{ channel: string; }' but required in type 'ChannelInput'.
Without one, the host's own input must satisfy the routine, as stepFor steps
require.
Mounting twice needs as, because phase names have to stay unique:
.use(pullChannel, { as: "Amazon" }) // → "Amazon / Fetch", "Amazon / Normalize"
.use(pullChannel, { as: "Shopify" })Without it you get a DuplicateNameError at build time. That rule is not
routine-specific — two phases anywhere in a script, or two steps within one
phase, are refused. Names label the frame, are what outline() reports, and
identify a cache entry; two units sharing a name share a slot, which surfaces
as wrong data rather than as a failure.
A routine owns whole phases, so addStep directly after use() is refused
too — open a phase of your own first, rather than quietly appending to a
fragment someone else's script also mounts.
Ten runnable scripts, each aimed at a different part of the API. Most take a flag to switch between the happy path and the interesting one.
deploy.ts |
The tour: phases, progress bars, checklists, retry, note, clean, rollback. npm run example, npm run example:fail |
checkout.ts |
Saga semantics end to end — three mutations, three compensations, a secret dropped with clean, and logPlacement: "step" nesting each rollback's log under the step it undoes. -- --ok for the happy path |
resilience.ts |
retry with backoff, retryIf refusing to retry a 401, timeoutMs on a hung step, and an external AbortSignal. -- --timeout, -- --fatal, -- --cancel |
intake.ts |
defineInput with a hand-rolled Standard Schema, plus the full handler UI surface. -- --bad to watch validation reject the run, -- --dirty for a failing checklist item |
modular.ts |
stepFor steps living in steps/tenancy.ts, shared by two unrelated scripts — the imported rollback compensates in both. -- --fail |
log-placement.ts |
The same script run three times, once per logPlacement value, so the difference is visible directly |
cache.ts |
A cached Build phase and a cached step. Run it twice to watch the second run skip both. -- --fresh to rebuild and overwrite |
invalidate.ts |
context.cache from inside a step: correct an entry without refetching, drop one that is beyond saving, and watch the shape guard catch an entry written before the code changed. -- --patch, -- --drop, -- --drift |
routine.ts |
One reusable fragment in routines/channel.ts, mounted by a sync script and, months later, a report that maps its own input onto the routine's — and gets the pull from cache. -- --multi mounts it twice, one storefront and API key per mount |
validate.ts |
The smallest thing that runs: two steps, rollback: "none" |
npm run example:checkout -- --ok
npm run example:resilience -- --cancel
npm run example:intake -- --bad
npm run example:modular -- --fail
npm run example:invalidate -- --drift
npm run example:routine -- --multi
npm run example:log-placement
npm run build # emit dist/
npm test # node:test suite
npm run typecheck # library + examples + tests + scripts
npm run preflight # the release gate (add -- --allow-dirty while working)
npm run example # the happy path, live
npm run example:fail
npm run release # release-it: version, tag, publish
scripts/preflight.ts is the gate, and it is written
with this library — so every release exercises phases, tasks, note, clean,
when and retry on real work:
› Quality
✔ typecheck 1.2s
✔ tests (120 passing in 9 files) 5.1s
› Package
✔ build 404ms
✔ entry point resolves 0ms
✔ tarball contents (47 files, 66kB) 180ms
› Registry
✔ version is unpublished 212ms
It checks the working tree is clean, typechecks all four projects separately so
a failure names the one that broke, runs the suite, then imports dist/index.js
the way a consumer would and asserts the public surface is actually there.
npm pack --dry-run has to contain the entry point, its declarations, LICENSE
and README — and must not leak src/, test/, examples/ or scripts/.
Finally it asks npm whether the version is already published.
release-it runs it from after:bump, which is the first point at which
package.json holds the version being released. A failure aborts before anything
is committed, tagged or published.
One config per directory, because editors resolve a file's config by finding
the nearest tsconfig.json — a non-standard filename like tsconfig.foo.json
is invisible to them, so the IDE and the CLI disagree.
| Config | Covers | Why it differs |
|---|---|---|
tsconfig.json |
src |
The library. noEmit, strict. |
tsconfig.build.json |
src |
The only config that emits dist/. |
examples/tsconfig.json |
examples |
Imports the built dist/; adds allowImportingTsExtensions. |
test/tsconfig.json |
test |
Same, without the flag. |
Only examples/ needs that flag, and only because modular.ts imports a
sibling: Node's type stripping does not rewrite import extensions, so the
import has to say ./steps/tenancy.ts rather than ./steps/tenancy.js. The
flag cannot go in the root config — tsconfig.build.json extends it and sets
noEmit: false, and TypeScript rejects allowImportingTsExtensions unless
noEmit or emitDeclarationOnly is set (TS5096).
The live renderer repaints an in-place region and writes logs permanently above
it. It falls back to line-per-event output when stdout is not a TTY, when
NO_COLOR is set, or when TERM=dumb — so CI logs stay greppable.
The frame is always made to fit the viewport, and this is load-bearing
rather than cosmetic. Repainting works by moving the cursor up N lines, so a
frame taller than the terminal scrolls its own anchor off screen and each
repaint appends a copy instead of overwriting. renderBody applies four
progressively stronger reductions until the frame fits:
- everything, in full;
- finished phases collapse to a one-line summary;
- only the running phase keeps its steps, and drops detail lines;
- the running phase's steps are windowed around the one executing.
The executing step is never windowed out, and long lines are truncated rather than wrapped. A terminal resize invalidates the recorded geometry, so the renderer abandons the old region and redraws below it. The closing frame is exempt — nothing repaints after it, so it keeps every step.