Skip to content

Repository files navigation

GraphARC

GraphARC

PyPI Downloads Python CI License

The admission gate for agent graphs, built on LangGraph.

A model proposes a graph of work; a deterministic checker admits it or refuses it with reasons; only admitted graphs execute, under budgets, onto one replayable JSONL trace. No step runs unadmitted, the worst case is priced before a run and billed per node after, and the dashboard cannot disagree with the audit trail because they are the same file. The edges are documented in Limits.

A nine-node incident-investigation graph running live in the browser: triage fans out into four parallel evidence pulls, they join at correlate, then hypothesize, verify and report — each node amber while it runs and green with its own token bill when done.

One question in, a governed graph out, live in the browser. (mp4)

Install

pip install grapharc     # Python >= 3.12
grapharc demo stage0     # costs nothing, needs no key

Backends are extras: grapharc[openrouter], [openai], [ollama], [server], [all]. The default backend drives the claude CLI on your PATH — a Claude subscription, no API key.

Quick start

grapharc start                    # guided tour
grapharc init                     # scaffold registry.py + grapharc.toml
grapharc plan "look into the outage" --model ollama/qwen3:8b   # propose -> admit -> save
grapharc go                       # execute the saved plan
grapharc plan "..." --scripted    # free rehearsal, no AI involved
grapharc serve --live-root .grapharc/runs   # live browser view of every run
grapharc replay <trace> <run-id>  # reconstruct a run from its trace

A terminal running grapharc plan: round 1 is rejected with edge_denied and never executes, round 2 replans and runs, then viz draws the graph and metrics prints the per-node bill — all from the one trace file.

Free and deterministic — --scripted runs the registry's own stand-in planner, so this reproduces exactly on any checkout. Three more recordings, and what is and is not staged in each, in docs/demo/.

Building a graph directly:

from grapharc import GraphARC, GraphARCState, Budget
from grapharc.runtime.graph import START, END

class State(GraphARCState):
    question: str
    answer: str = ""

def answer(state: State) -> dict:
    return {"answer": f"42 (asked: {state.question})"}

g = GraphARC(State, name="demo", budget=Budget(max_iterations=10))
g.add_node("answer", answer, writes={"answer"})   # undeclared writes raise
g.add_edge(START, "answer")
g.add_edge("answer", END)
print(g.compile().invoke({"question": "meaning of life"}))
# {'question': 'meaning of life', 'answer': '42 (asked: meaning of life)'}

Typed state, declared writes, a budget — none optional. The rest of the surface (agents, tools, memory, sessions, policy documents, cost attribution, OTel) is in the deep dive and the cookbook.

The admission gate

You cannot pre-author a graph for "investigate this incident" — the shape is discovered while working. So the graph is proposed at runtime, and a deterministic checker stands between proposing and running: registry, policy, remaining budget, depth, acyclicity, all on every proposal. A rejection is structured feedback the planner replans against; work discovered mid-run re-enters the same gate. Watch it refuse (free, scripted, the first proposal names a policy-denied kind):

grapharc plan "investigate the checkout outage" --scripted --go
goal      : investigate the checkout outage
model     : scripted stand-in (--scripted)
registry  : grapharc.examples.plan_incident:build_registry
kinds     : deploy, patch, triage, verify
policy    : grapharc.examples.plan_incident:build_registry default (deny -> deploy, otherwise allow)  [registry-default]
config    : no grapharc.toml (flags and defaults only)

stopped   : goal_met  (the goal check was satisfied)
rounds    : 2 of max 8
   round 1: rejected  nodes=2 executed=False  rejected: edge_denied
   round 2: admitted  nodes=3 executed=True

state     : goal='investigate the checkout outage' notes=['triage ran', 'patch ran', 'verify ran']

Round 1 wanted to deploy and never executed. Round 2 went through the same checker and ran. The full assembly, runnable as written:

from pydantic import BaseModel

from grapharc.harness.permissions import Decision
from grapharc.planner import (
    AdmissionChecker, CostEstimate, EdgePolicy, EdgeRule, GovernedLoop,
    LoopLimits, Materializer, NodeRegistry, NodeSpec, PlannerNode,
)
from grapharc.runtime.budget import Budget
from grapharc.testing import ScriptedChatModel


class State(BaseModel):
    found: str = ""
    fixed: str = ""


def factory(spec):                              # bodies come from HERE, never a proposal
    def body(state):
        return {"found": "cause"} if spec.name == "search" else {"fixed": "patch"}
    return body


registry = NodeRegistry([                        # the kinds a planner may propose
    NodeSpec(name="search", factory=factory, worst_case=CostEstimate(tokens=500)),
    NodeSpec(name="edit",   factory=factory, worst_case=CostEstimate(tokens=2000)),
    NodeSpec(name="deploy", factory=factory),
]).freeze()
policy = EdgePolicy(rules=(                      # deny -> ask -> allow, unmatched is deny
    EdgeRule(action=Decision.DENY, target="deploy"),
    EdgeRule(action=Decision.ALLOW),
))

plan = '{"nodes": [{"name": "%s"}], "edges": [{"source": "__start__", "target": "%s"}]}'
loop = GovernedLoop(
    planner=PlannerNode(
        ScriptedChatModel(responses=[plan % ("deploy", "deploy"), plan % ("edit", "edit")]),
        catalog=registry.catalog(),
    ),
    checker=AdmissionChecker(registry=registry, edge_policy=policy),
    materializer=Materializer(
        registry=registry, state_schema=State,
        writes={"search": {"found"}, "edit": {"fixed"}, "deploy": set()},
    ),
    budget=Budget(max_tokens=100_000),
    limits=LoopLimits(max_rounds=8),
    goal_reached=lambda s: bool(s.fixed),
)
result = loop.run("find and fix the bug", State())

print(result.stop.value)
for record in result.rounds:
    print(record.round, record.admission.status.value, record.executed)
print([r.code for r in result.rejections()])

Output:

goal_met
1 rejected False
2 admitted True
['edge_denied']

Both blocks are executed by tests/test_readme.py against every commit, so this page cannot drift from the code.

What a policy actually does to an agent

Not "blocks the bad call at the last moment" — it decides what the planner is able to propose. The planner is told the policy before it plans, so a denied edge changes the shape of the graph rather than producing a refusal to retry.

A terminal: a failing test, then a policy denying any edge into apply_change, then a plan under that policy proposing three read-only nodes whose rationale says it cannot reach apply_change, then the amended policy, then the same goal producing a five-node mutating graph, then the execution, the diff, and the test passing.

That is GraphARC fixing a bug in GraphARC — a real one, from this project's backlog, in a copy of this repository. Under deny *->apply_change the planner proposes three read-only nodes and says so itself:

Since apply_change cannot be reached by an edge, this round investigates the torn trace bug and writes findings to notes for a human to act on.

mutating: false. A person then amends the rule, and the same goal on the same model returns a five-node graph containing apply_change, mutating: true. Claude Code executes it and the red test goes green. The recording, and exactly what in it is staged (the pacing) and what is not (everything else), is in docs/demo/.

Supervised Claude Code, from Slack

One message asks for work; the answer is the graph it intends to run — nodes, their governed kinds, edges, worst-case cost — and two buttons. Nothing executes until a human presses one.

A Slack thread: the bot replies with the proposed three-node graph and Approve / Deny buttons, the plan is approved, the nodes then run one by one, and the closing frame reads the trace back — approval_request, approval_response, then the first node's start.

/grapharc plan "explain what flaky.py does and why it is not reproducible" --go \
    --model claude-cli --registry grapharc.stdlib:build_registry

--go means plan and execute. From Slack the gate appends --approve to it unconditionally, so the run parks before its first node — a message from anyone in the workspace can propose work, and only a person can start it. The click is bound to the plan's fingerprint, so a button on a superseded proposal is refused rather than honoured. Setup and the full rule list are in the Slack cookbook; what is and is not mocked in that recording is in docs/demo/.

Where it sits

GraphARC Claude Code OpenClaw raw LangGraph
Shape Governed multi-node graph runtime Interactive single-agent loop Personal AI assistant gateway Graph mechanism library
Who authorizes work Deterministic gate, pre-execution, with reasons A human, live, per action Configuration and allowlists Nobody — convention
Cost control Worst-case admission + per-node bill, fail-closed Usage visibility Spend settings None built in
Audit One replayable JSONL trace Session transcripts Logs Checkpoints (state, not why)

Different jobs, not competitors — GraphARC's default backend drives the Claude CLI. See benchmarks for measured comparisons against third-party agents on the same tasks: success, cost, wall time, and policy violations, with raw logs committed.

Limits

The edges are documented, not denied — the full list with mechanisms is in the deep dive.

  • Admission authorises a node's kind; its arguments only where the kind declares an args_schema, and a schema bounds their shape, not what a factory lets them reach.
  • The in-process sandbox is defense in depth; ContainerExecutor is the real boundary. run_command children are unconfined.
  • The HTTP API does not yet use the durable session layer.
  • On the Claude CLI backend an agent node is delegated, not governed: by default it runs under an allowlist mapped from the node's own tools, but enforcement there is Claude Code's, and the bypass tier — explicit opt-in — has no checks at all.
  • Policy documents govern planning; the tool plane still reads CLI flags.
  • The MCP gate binds the MCP surface, not the host: an agent with its own file tools in the run directory could forge the approval decision. The trust boundary is the working directory, as it is for the Slack workspace.

Version 0.1.5 · changelog · roadmap · website · MIT

About

An end to end implementation of Graph Engineering as proposed by Andrew NG and Peter Steinberger

Topics

Resources

Contributing

Stars

46 stars

Watchers

3 watching

Forks

Releases

Packages

Contributors

Languages