docs(science): Phase A closure report — validity program A1–A7 - #64
Merged
pixelstrade-dev merged 2 commits intoAug 23, 2026
Merged
Conversation
… answers, Phase B gates Maps the external audit's validity criticisms to the merged Phase A deliverables (PRs #57-#63), states what is now honest to claim, and lists the empirical Phase B gates (Run 002 with >=3 judge families, corpus v1 per the A6 sizing, human annotation of PIGA intent spaces, external replication) that alone can move the validity score further. Under orchestrator fact-check; corrections will follow if any claim fails verification. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y7wiMwFa2D4P8zD9RhCgsv
- construct-card provenance corrected: #57 shipped 7 cards, the 8th (PIGA) came with #63 - Run 002 owner action made operationally complete: BOTH OPENWEIGHT_API_KEY and OPENWEIGHT_BASE_URL as Secrets plus the exact model identifier in the run config — Run 001 recorded this judge as skipped for a missing env var - 'published anchors' (plural) narrowed to the one published anchor (Shrout-Fleiss 1979) plus hand-computed anchors - 'a lucky silent guess still scores 0' corrected to 'earns zero behavior credit' (the composite retains the coverage channel, ~0-12) - adversarial-review sentence now states the reviews are session history and only their outcomes (committed correction notes) are repo-verifiable Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Y7wiMwFa2D4P8zD9RhCgsv
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes the loop on Phase A of the validity program
docs/validity-program-phase-a.mdis the closure record for Phase A (PRs #57–#63): a table mapping each criticism from the external audit ("scientific validity of the measurements: 2/10") to the merged answer, what it is now honest to claim, and — stated plainly — what still caps the score: no further code moves validity much from here; the four remaining gates are empirical (Run 002 with ≥ 3 judge families, corpus v1 per the A6 sizing, human annotation of PIGA intent spaces, external replication).The report itself went through the orchestrator's fact-check (~45 claims verified against committed artifacts). Findings applied before this PR:
OPENWEIGHT_API_KEYandOPENWEIGHT_BASE_URLas Secrets plus the exact model identifier — Run 001 recorded that judge as skipped for the missing env var);Docs-only change; no code paths touched.
🤖 Generated with Claude Code
https://claude.ai/code/session_01Y7wiMwFa2D4P8zD9RhCgsv
Generated by Claude Code