operational-rigor + skill-vetting: trust-grant breadth — judge the effective grant expansion, never the syntax - #229
Merged
Conversation
…tive grant expansion, never the syntax Canonical rule in operational-rigor §2, between the instruction-files install gate and the activation-gated-payload rule; bare-pointer mirror (Over-broad trust grant) in skill-vetting §2. Two-round dual-blind design gate (gpt-5.6-luna + gpt-5.6-sol, both at max effort): r1 = luna FIX-1 + sol FIX-2 -> both findings reproduced and fixed -> r2 = PROCEED x2, zero findings. Landed blocks byte-identical to the r2-reviewed fenced blocks; declared adaptations: owner-settled provenance entry + placement only. Static discrimination controls W1-W8 + W2b 9/9. One new `unprobed` marker (canonical only) -> standing #115 queue; behavioral transmission/effectiveness unprobed. Evidence package under reviews/2026-08-30-trust-grant-breadth/. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_0132RthrKSsMkywcEwtkXhkx
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
One new canonical rule in
operational-rigor§2 — trust-grant breadth: a trust or allow rule is judged by its effective grant expansion, never its syntax — placed between the instruction-files install-gate block and the activation-gated-payload rule, plus a bare-pointer mirror (Over-broad trust grant) inskill-vetting§2, plus the evidence package. Two canonical files, pure additions; 9 evidence files. Commit0f2e344onmain-tip5824f30.Invariant (headline): the effective granted capability set must stay within what was actually vetted, or what an independent trusted policy explicitly authorizes as a class.
Why (adjudicated PARTIAL-GAP)
Existing doctrine catches config WRITES (config self-propagation), authorization RHETORIC (authorization-default flip), and tool REGISTRATION (MCP rule) — but no rule evaluated whether a persistent trust/allow/approval entry's EFFECTIVE GRANT SET exceeds the actually-vetted set, and a candidate that merely instructs the human to paste a broad grant intercepted nothing. External threat evidence is cited as shape (ATR-2026-02192; AWS Kiro AWS-2025-019), not first-hand reproduced.
The rule (owner-locked semantics)
Review
Design gate, two rounds, dual-blind (mutually blind, isolated dirs; identity from tool banners in all four runs:
gpt-5.6-luna/gpt-5.6-sol, reasoning effort max). r1 = luna FIX-1 + sol FIX-2: convergent finding — the mirror restated the operative criterion and clearance fragments (an incomplete second authority) — plus sol's headline finding ("silently" wrongly implied disclosure could matter; the independently-authorized limb was missing). Both reproduced first-hand and fixed. r2 = 2/2 PROCEED, zero findings, all ten mandatory axes line-anchored by both reviewers (canonical/mirror separation and the authorized-as-a-class verdict language explicitly confirmed by both).Landing fidelity
Landed blocks byte-identical to the r2-reviewed fenced blocks (machine-proven: each block exactly once; full diff pure additions; per-file added-line multisets equal exactly the reviewed blocks plus the owner-settled Provenance entry). DECLARED-LANDING-ADAPTATIONS: provenance entry + placement only — the single inline
unprobedmarker was already inside the r2-reviewed bytes. Neighbors byte-unchanged (instruction-files bullet, activation-gated-payload rule, skill-vetting's config-self-propagation / authorization-default-flip / MCP bullets, self-vouching + activation pointer). Static discrimination controls 9/9 (W1–W8 + W2b; W2b proves the carve-out records class authorization, never individual vetting).checks.pygreen, invisible-Unicode sweep included. Details inreviews/2026-08-30-trust-grant-breadth/FINAL-WORDING-RECONSTRUCTION.md.Dup-check (orientation, owner-ratified PARTIAL-GAP; r2 reviewers confirmed axes 1 and 5): the adjacent homes — operational-rigor §2's install-gate family and skill-vetting §2's config-self-propagation / authorization-default-flip / MCP-auto-registration checks — each catch the write, the claimed authority, or the registration act; none evaluates effective grant breadth. Not found under those reads; orthogonality and gap closure reviewer-confirmed.
Probe debt
Behavioral transmission/effectiveness remains unprobed. Exactly one new
unprobedmarker (canonical rule only; the mirror carries none); the future behavioral probe joins the standing #115 queue.The three locked sibling candidates (compiled-bytecode/decoy-source divergence; visible homoglyph deception; approval-fatigue reviewer-targeting) remain LOCKED — zero of their content here.
🤖 Generated with Claude Code
https://claude.ai/code/session_0132RthrKSsMkywcEwtkXhkx