docs(plans): verifier-independence plan - #590
Conversation
…agy/grok panel Nine tickets, six waves: A (replay harness, autopsy, binder row-label repair #331/#418/#146), B (per-stage exclusive timings, breakdown run), C (design and implementation of geometry-addressed cell correspondence with an abstain rule, corpus re-run). Frozen evidence: ~/Data/socr/ladder-run2-2026-09-04 (run 2 at main@f434019, 1 lifted / 6 held, 3 native-label-blocked). #155 parked.
|
ⓘ Qodo reviews are paused because your trial has ended. Ask your workspace admin to add credits to resume reviews. Manage billing |
|
Important
This repository does not receive automatic reviews because it has fewer than 10 stars. ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Team Run ID: Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
r-uben
left a comment
There was a problem hiding this comment.
Verdict: MERGE
Docs-only plan. Spot-checked the cited seams on current main (92c1527): binder owns its own rows (binding.py:70-71, _native_rows:559, _assign_bands:332, _row_label_and_bbox:322, bind:1276 — not native_rows.py); _disprove_one really self-witnesses via item.native_bbox (adjudication.py:260-268); bands_from_rules:122 / _horizontal_rules:169 match; orchestrator clamp/lift/held call sites (:5223 / :5245 / :5292) and _adjudicate_clamped_table / _render_adjudication_crop / _apply_binding_adjudication_meta land where claimed. Target label count matches the run-2 log (doc02 p3×3 + doc02 p4×4 + doc04 p3×1 = 8).
Gates are the right shape for this repo: A2 hard-gates on frozen replay (3/3 false clamps cleared + control fixtures against false accepts), not on lifted and not on a live ACCEPTED quota; C3 is a report; abstain never falls back to either lane's box; B1 keeps timings out of fingerprint / fragment markdown and restores on resume; hermetic patches + pin-a-difference called out; Wave 1 file split (A1 import-only of _render_adjudication_crop, B1 owns orchestrator) is real.
Panel call on C3 (Codex over agy's ≥10/18 quota) is correct — live OCR is not a deterministic gate.
Non-blocking — fix before Wave 1 / Wave 3 dispatch, not merge-blocking:
- STATUS base SHA is stale. Board says
main @ b7323f7(#587); this PR's base and currentmainare92c1527(#588, digest-correction log). Refresh the base-state line so/plan nextdoes not start from a dead tip. - A2 label gate formula. Done-when says both "fix only A1b's native-side classes" and "8/8 unresolved target labels reconstructed". If A1b reclassifies any of the 8 off
shredded_label/bbox_truncated, pin explicitly: reconstruct gate = native-side subset only;absent_text/otherlogged on their own line and out of that denominator (the "never subtracted" line is easy to read the wrong way under pressure). - B2 vs the 8.0 min/page figure. Run-2's own log records the venv-pointer confound (three digests). B2 should state it establishes a fresh timed baseline under one digest + B1, not that it explains the confounded 8.0.
- STATUS "Next action" still says commit/open the PR — already done. Point it at Wave 1 dispatch.
Bot noise on the PR (Qodo paused, CodeRabbit trigger) ignored.
No code. No merge from the desk.
Revise panel (gitty demand)Rubén wants another revise pass on this plan. Please attack fidelity (does the plan match the tree?), coverage (are the real holes ticketed?), gating (are done-when checks pinable / non-vacuous?), and dispatchability (is wave 1 = A1 ∥ B1 still the right first cut?). Known soft notes from desk MERGE on
Streams: A (binding.py row labels / replay harness / autopsy), B (per-stage wall-clock), C (geometry cell correspondence + abstain; C3 report-not-quota). Parked: #155, #585, latency opt, booktabs. @claude — coverage + gating; walk the ticket graph against I’ll fold verified findings into the plan and push a revision. Do not merge. |
|
@claude Please revise this verifier-independence plan for coverage and gating. Walk A1/A1b/A2/B1/B2/C* against the tree ( |
|
bugbot run |
|
@codex review Revise pass on this docs-only plan. Axes: fidelity to the tree, dispatchability of wave 1 (A1 ∥ B1), and whether A2/C3 gates are non-vacuous. Soft notes to verify: STATUS base SHA vs main, A2 8/8 vs A1b demotion, B2 as fresh baseline not confounded run-2 minutes. |
|
Claude finished @r-uben's task in 5m 57s —— View job Revise panel — coverage + gating pass (claude seat)
Spot-checked every cited symbol and line range against current Confirming the desk's three soft notes — all still open:
Two new coverage findings (not in the desk review):
Not re-litigated: stream C's abstain rule, Verdict: coverage and gating both hold at the ticket-graph level; the two new findings (#4, #5) are non-blocking for Wave 1 (A1/B1 don't touch this) but should land before A2 is dispatched in Wave 3, alongside the desk's three open items. No code, no merge from this seat. |
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
|
/gemini review |
|
Claude finished @r-uben's task in 4m 10s —— View job Revise pass — coverage + gating (@claude)
Verdict: no new blockers. All four of the desk's soft notes are still open (not yet folded in) and I confirm all four independently. One additional fidelity nit for C2b (wave 5, not merge-blocking). Independent tree walkRe-verified against
Soft notes from the desk — all confirmed still open
One new finding (wave 5, not merge-blocking)C2b's file/target annotation for Wave 1 dispatchability: still correctA1 ( No code changes made — this is a review-only pass, docs-only PR as before. Nothing here blocks Wave 1; items 1–4 and the C2b nit are cheap edits before their respective waves dispatch (1/2/4 before Wave 1/2, the C2b nit before Wave 5). |
There was a problem hiding this comment.
✅ Bugbot reviewed your changes and found no new issues!
Comment @cursor review or bugbot run to trigger another review on this PR
Reviewed by Cursor Bugbot for commit 63c4e71. Configure here.
r-uben
left a comment
There was a problem hiding this comment.
Verdict: REQUEST_CHANGES (fold before Wave 1 / before /plan next; docs-only — do not merge until revised)
Revise pass on 63c4e71 vs current main (92c1527). Axes: fidelity, coverage, gating, dispatchability. Soft notes from prior MERGE — all KEEP (one strengthened).
Soft notes (verify → keep)
-
STATUS base SHA / next-action — KEEP. Board still says
main @ b7323f7(#587). Tip is92c1527(#588, digest-correction log). Next action still says commit/open the PR — already done. Point at Wave 1 dispatch and refresh the base-state line so/plan nextdoes not start from a dead tip. -
A2 8/8 vs A1b demotion — KEEP / strengthen. Done-when says both "fix only A1b's native-side classes" and "8/8 unresolved target labels reconstructed". Pin explicitly: reconstruct gate = native-side subset only (
shredded_label/bbox_truncated);absent_text/other/neighbour_capturelogged on their own line and out of that denominator. -
B2 fresh baseline — KEEP / strengthen. Run-2's own log (now on
mainvia #588) records the three-digest venv-pointer confound. STATUS still presents8.0 min/pageas an uncaveated measured baseline. B2 must say it establishes a fresh timed baseline under one digest + B1, not that it explains the confounded 8.0.
Fidelity
Seams still match the tree: binder owns rows (binding.py:70-71, _native_rows:559, _assign_bands:332, _row_label_and_bbox:322, bind:1276 — not native_rows.py); _disprove_one self-witnesses item.native_bbox (adjudication.py:260-268); bands_from_rules:122 / _horizontal_rules:169 match; clamp/lift/held / crop / meta call sites land; fingerprint carries socr_source_digest (:1251); restore at :8901. Flat tests/, scripts slot in pyproject.toml — fine.
Minor citation drift only: Fixed-inputs _flush_page_sidecar span (:8272-8453) is stale — def is at :8147 on this tip. Refresh on fold; not a retarget.
Example/target mismatch: A2's problem text quotes Sample 1988:1–2019:12 … (doc05/doc07 class (c) in the run-2 log) but the 3/3 gate is doc02 p3 / doc02 p4 / doc04 p3 only. Either drop Sample from the A2 blurb or ticket those loci.
Coverage
Real hole: plan parks doc05/doc07 under #585 ("sibling-LaTeX and lane items") but the run-2 log keeps class (c) native row-label items there too (doc05: 2, doc07: 1) alongside (b)/(d). A2's 8/8 does not cover them. Fold one explicit line: park those class-(c) items as out-of-wave (mixed hold; #585/#331 follow-up), or add them to A1b's census so they are not invisible.
12 class-(c) items in the log vs 8 target labels — the delta must be named, not implied.
Streams A/B/C scope and parked set (#155, #585, latency opt, booktabs) otherwise still right. C3 report-not-quota (Codex over agy's live quota) still correct.
Gating
Shape is still right for this repo: A2 hard-gates frozen replay (3/3 false clamps cleared + control fixtures), not lifted, not a live ACCEPTED quota; C3 is a report; abstain never falls back to either lane's box; B1 keeps timings out of fingerprint / fragment markdown and restores on resume; hermetic patches + pin-a-difference called out.
Vacuous risks to pin on fold (non-blocking if stated):
- A1 "reproduces 1 lifted / 6 held on unchanged tree" — good smoke; does not prove the harness rebinds (add one hermetic fixture that forces a bind delta).
- A2 8/8 — see demotion rule above.
- B2 — must not treat confounded 8.0 as the thing explained.
Dispatchability
Wave 1 = A1 ∥ B1 still the right first cut. File split is real: A1 = new benchmark/replay_binding.py + scripts entry + hermetic tests + read-only import of _render_adjudication_crop; B1 owns orchestrator.py / state.py / manifest.py / cli.py. C2b's depends-on B1 is file-serialization on orchestrator, not a semantic dep — fine as written.
Do not dispatch Wave 1 until STATUS base SHA + next-action + B2 confound wording + A2 demotion rule + doc05/doc07 class-(c) park line are folded.
Bot noise (Qodo paused, CodeRabbit trigger) ignored. Claude/Codex/Bugbot revise seats still in flight — fold their verified hits in a follow-up push; this is the desk seat.
No code. No merge from the desk.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 63c4e711d0
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| **Done when:** `~/venvs/socr/bin/socr-replay-binding ~/Data/socr/ladder-run2-2026-09-04` | ||
| prints 7 rows and reproduces run-2's `1 lifted / 6 held` from the frozen sidecars on the | ||
| unchanged tree; `~/venvs/socr/bin/pytest tests/test_replay_binding.py -q` exits 0 with |
There was a problem hiding this comment.
Verify fresh binding results instead of replaying recorded statuses
The corpus acceptance check can reproduce 1 lifted / 6 held simply by reading the recorded status that A1 is explicitly told to emit; bind() itself produces no disposition, and this ticket never requires the fresh contradiction signatures or labels to match the recorded items. A harness that extracts the wrong table region, passes the wrong markdown, or returns an empty binding can therefore satisfy this done condition while being unusable as A2's regression oracle. Require an explicit comparison between each fresh binding result and the frozen item/signature set on the unchanged tree.
Useful? React with 👍 / 👎.
| **Done when:** on frozen replay, **3/3 target tables cleared of false native-label | ||
| clamps**; **8/8 unresolved target labels reconstructed against the rendered source** | ||
| (A1b's `labels.json`); **zero false accepts** on the row-swap and neighbouring-label |
There was a problem hiding this comment.
Derive A2's target count from the A1b autopsy
If A1b classifies even one of the eight failures as absent_text, neighbour_capture, or other, A2 is instructed to fix only shredded_label and bbox_truncated cases but is still required to reconstruct all 8 labels and clear all 3 tables. That makes the ticket impossible to complete precisely when the preceding autopsy disproves the current attribution. Define the A2 denominator from A1b's repairable native-side set, while retaining the excluded cases as explicit non-regression expectations.
Useful? React with 👍 / 👎.
| **Done when:** the log exists with all of the above; **the hard gate is A2's frozen-replay | ||
| 3/3 (already met by then), not this count.** The derivable expectation is 10/18 ACCEPTED | ||
| (7 baseline + 3 targets) *if* nothing else moves; the log reports the actual number and, | ||
| where it differs, names the class that moved it. |
There was a problem hiding this comment.
Gate the geometry verifier on frozen real disputes
A2's 3/3 gate runs before C2b and clears the target contradictions, so it cannot validate the geometry-addressed disproof introduced by C2b. The remaining C3 requirement only asks for a report, which means C2b may abstain on every real corpus contradiction—or otherwise fail to produce useful correspondence—and the entire plan still completes. Add a deterministic frozen-replay assertion for known geometry-addressable disputes and expected abstentions after C2b; this avoids a nondeterministic acceptance quota while actually gating the verifier-independence change.
Useful? React with 👍 / 👎.
Second panel (Codex/Claude/Cursor/Bugbot; Gemini silent) updates: - Base SHA 92c1527 (#588); caveat confounded 8.0 min/page baseline - Wave 1 dispatch as next action (A1 ∥ B1) - A1 explicit fresh bind() vs frozen sidecar comparison gate - A2 native-side-only reconstruct denominator; target-table examples - B2 fresh timed baseline under one digest - C2b file split (ContradictionItem in adjudication.py) + frozen-replay gate - Park doc05/doc07 class-(c) native labels out of wave Co-authored-by: Rubén Fernández-Fuertes <r-uben@users.noreply.github.com>
|
Revision 2 pushed (
Wave 1 (A1 ∥ B1) unchanged. |
|
Claude finished @r-uben's task in 3m 13s —— View job Revise pass — verify revision-2 fold (@claude)
Note: the local checkout only has 7 of 8 confirmed folded correctly
1 of 8 not actually landed: the 12th class-(c) item is still unallocatedThe changelog says "A1b censuses all 12 and locates the unallocated one," but the ticket text doesn't do this yet.
Suggested fold: add one line to A1b's Not re-checked this passStream C1/C2a substance, hermeticity patches, byte-identity/resume-gate citations — these held on the prior two passes and nothing in this revision touched them. Verdict: one real, concrete coverage gap (doc00 p4-t0 / the 12th class-(c) item); everything else in the changelog checks out against the tree. Non-blocking for Wave 1 (A1 ∥ B1 don't touch A1b), but should land before Wave 2 dispatches A1b. No code, no merge from this seat. |
…of the rev-2 fold On 6473b2a (the concurrent fold of the same panel): A1b censuses all 12 class-(c) items and defines A2's denominator N explicitly; B2 asserts one socr_source_digest before tabulating; C2b's frozen-replay gate is a per-item prediction recorded before implementation and fails if every item abstains; Fixed inputs name ContradictedCell / RowLabelContradiction vs ContradictionItem.
|
Rev 3 ( |
…1/A2/B2/C2b gates A1 compares multisets with duplicate counts and asserts the exact expected delta with sidecar bytes unchanged; A2 freezes A1b's classifications before repair and revises 3/3 explicitly if invalidated; B2 checks the resolved package path and that every sidecar digest equals the checkout's intended digest (sidecars agreeing with each other does not catch a repointed venv); C2b's prediction is a separately committed artifact referenced by SHA, with a feasibility checkpoint instead of a forced geometry-addressed item.
|
Rev 4 (
|
…eckout with explicit PYTHONPATH
r-uben
left a comment
There was a problem hiding this comment.
Verdict: MERGE
Re-review of tip ff799b5 (rev 5) vs desk REQUEST_CHANGES on 63c4e71 (5118389648). Docs-only; prior blockers folded.
Prior REQUEST_CHANGES — cleared
- STATUS base SHA / next-action →
92c1527(#588); next action = Wave 1 A1 ∥ B1. - A2 demotion / N/N → denominator = A1b native-side only; classifications frozen before repair; silent shrink forbidden.
- B2 vs confounded 8.0 → Goal + STATUS caveat the three-digest mid-run; B2 is a fresh one-digest baseline (rev 5: pinned checkout + explicit
PYTHONPATH+ resolved-path check). - doc05/doc07 class-(c) park → explicit parked line; A1b censuses all 12.
- Sample /
_flush_page_sidecarfidelity → Sample dropped from A2; cite refreshed to:8147. - A1 vacuity → multiset (
native_token/model_token/kind, duplicate counts) + exact-delta fixture + sidecar bytes unchanged. - C2b →
ContradictedCell/RowLabelContradictionin binding; prediction as separately committed SHA-referenced artifact; feasibility checkpoint (no forced geometry item). - C3 → inherits B2 run discipline.
Wave 1 file split still real. C3 report-not-quota still right.
Soft — fold before Wave 2 (A1b) / Wave 3 (A2), not merge-blocking
- A2 sibling-module framing.
#331/#418/#146already land inreconstruct.py/header_repair.py(test_gh331_stub_labels.pyimportsrowize_from_word_list, notbinding). Binder has its own rowizer (binding.py:70-71). Say explicitly: A2 fixes an independent recurrence inbinding.py; usereconstruct.py's proven fix as design reference — do not re-touchreconstruct.py. - A2
Files:list.test_tr1_rowizer.py/test_gh331_stub_labels.py/test_table_header_gh146.pydo not importbinding— label them must-stay-green scope guards, separate from the edit list, or Wave-disjointness convention breaks. - 12th class-(c). A1b says "find it"; Claude's arithmetic pins doc00 p4-t0 (log: contradiction items not captured). Name that candidate in A1b so it cannot stay invisible the way doc05/doc07 almost did.
Bot noise ignored. No code. No merge from the desk.
Summary
Ticket graph for making the free native lane a sound witness of the model lane. Drafted from a Claude ↔ Codex conversation, attacked by a three-seat panel (Codex fidelity, agy coverage, grok gating) — 24 findings verified against the tree and folded in; Codex approved the revision.
binding.py, which builds its own rows.Baseline: run 2 (
docs/log/2026-09-04_ladder-corpus-rerun.md) — 7 adjudicated tables, 1 lifted / 6 held, 3 blocked by native row labels.Test plan
uvx ruff@0.16.0 format --check .clean.Summary by cubic
Adds a docs-only, nine-ticket, six-wave plan for making the free native lane a sound witness of the model lane: stream A repairs the native row-label reference in
binding.py, stream B instruments per-stage exclusive wall-clock, and stream C moves the disproof crop to geometry-derived cell addresses with an abstain rule.bind()against the frozen sidecar as a multiset with duplicate counts; A1b censuses all 12 class-(c) items and freezes its classifications as A2's native-side-only denominator.PYTHONPATH; B2 asserts every sidecar'ssocr_source_digestequals the intended digest via the resolved package path before tabulating.ruff formatclean; wave 1 (A1 ∥ B1) can dispatch on merge.Written for commit ff799b5. Summary will update on new commits.
Note
Low Risk
Documentation-only; no production code, sidecar schema, or CI behavior changes until follow-up tickets land.
Overview
Adds
docs/plans/verifier-independence/with a nine-ticket, six-wave execution plan (no runtime or test changes in this PR).STATUS.mdrecords plan stage, frozen run-2 baseline (7 adjudicated tables, 1 lifted / 6 held), a ticket board with dependencies, parallel dispatch waves, and next steps (open PR, then wave 1: A1 ∥ B1).TICKETS.mdis the full spec: goal and panel-verified repo facts (e.g. row repair targetsbinding.py, notnative_rows.py; norow_bands_from_rulesyet), then per-ticket problem / files / done-when criteria across three streams — A frozen replay-binding + autopsy + binder row-label fixes, B exclusive per-stagetimings_sand a corpus breakdown log, C geometry-based cell correspondence, abstain semantics, and a ladder re-run reported (not gated) vs A2’s frozen-replay gate. Parked work and panel disagreement on C3 gating are documented explicitly.Reviewed by Cursor Bugbot for commit 63c4e71. Bugbot is set up for automated code reviews on this repo. Configure here.