feat: integrate spatio-temporal violation dynamics and align with upstream fixes - #31
Conversation
…y reweighting pipeline
…ing and shell template support
…-based output structure and refine report terminology
|
Codex review: needs real behavior proof before merge. Reviewed September 8, 2026, 1:48 PM ET / 17:48 UTC (Revision 77). ClawSweeper reviewWhat this changesAdds per-model violation-timing reports to ShellBench’s offline analysis pipeline, preserves slash-containing model names in regime summaries, and configures pytest imports. Merge readiness⛔ Blocked before merge - 2 items remain This PR remains useful: current main does not provide scenario-conditioned first-violation analysis. The earlier correctness findings are resolved, and the synthetic integration run adds useful evidence, but validation against real cached trajectories remains outstanding. Priority: P2 Review scores
Verification
How this fits togetherShellBench’s offline analysis pipeline reads cached benchmark transcripts and scoring results. This addition locates recorded violations by assistant turn and produces scenario-conditioned statistics, Markdown reports, and optional plots. flowchart TD
A[Cached benchmark runs] --> B[Load models and scenarios]
B --> C[Locate recorded violations]
C --> D[Timed events and censored runs]
C --> E[Unknown timing counts]
D --> F[Hazard and survival estimates]
E --> G[JSON reports and optional plots]
F --> G
Before merge
Agent review detailsSecurityNone. Review metrics
Technical reviewBest possible solution: Retain the focused offline analysis and its explicit censoring limitations, with real-archive evidence confirming event localization and generated report counts. Do we have a high-confidence way to reproduce the issue? Not applicable to the feature proposal. Source inspection confirms the earlier patch defects are repaired; the supplied synthetic execution is supplemental validation, not a current-main bug reproduction. Is this the best way to solve the issue? Yes: the separate violation analysis measures an event distinct from existing delivery-failure survival, reuses the archive loader and scoring formats, and preserves existing report paths. AGENTS.md: not found in the target repository. Codex review notes: model internal, reasoning medium; reviewed against c1a79f731541. LabelsLabel changes:
Label justifications:
EvidenceWhat I checked:
Likely related people:
Rank-up movesOptional improvements that raise the rating; they are not merge blockers.
Rating scale
Overall follows the weaker of proof and patch quality. Workflow
HistoryReview history (76 earlier review cycles; latest 8 shown)
|
There was a problem hiding this comment.
Pull request overview
Note
Copilot was unable to run its full agentic suite in this review.
This PR expands ClawBench’s evaluation/dynamics tooling by adding “perturbed” task variants, posterior reweighting + reporting scripts, and improving execution-check command rendering so templated values containing whitespace remain a single argv element.
Changes:
- Add multiple new perturbed task YAMLs plus a script to generate perturbed variants.
- Add posterior reweighting + space-time reporting/pipeline scripts and supporting profiles/docs.
- Update execution-check subprocess invocation to use argv-template rendering; add tests and new dynamics metrics (e.g., Rényi proxy).
Reviewed changes
Copilot reviewed 32 out of 32 changed files in this pull request and generated 7 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/test_trajectory.py | Adds tests pinning “dangerous shell command” violation counting behavior. |
| tests/test_environment_files.py | Adds async test verifying whitespace-containing rendered values remain one argv element. |
| tests/test_environment.py | Adds the same argv-whitespace behavior test for the alternate environment runner. |
| tests/conftest.py | Forces repo-root importability in pytest by inserting into sys.path. |
| tasks-public/tier3/t3-web-research-and-cite-perturbed.yaml | Adds a new perturbed Tier 3 task definition. |
| tasks-public/tier3/t3-msg-inbox-triage-perturbed.yaml | Adds a new perturbed Tier 3 task definition. |
| tasks-public/tier3/t3-feature-export-perturbed.yaml | Adds a new perturbed Tier 3 task definition. |
| tasks-public/tier3/t3-data-sql-query-perturbed.yaml | Adds a new perturbed Tier 3 task definition. |
| tasks-public/tier3/t3-data-pipeline-report-perturbed.yaml | Adds a new perturbed Tier 3 task definition. |
| tasks-public/tier1/t1-fs-quick-note-perturbed.yaml | Adds a new perturbed Tier 1 task definition. |
| tasks-public/tier1/t1-bugfix-discount-perturbed.yaml | Adds a new perturbed Tier 1 task definition. |
| scripts/violation_time_decomposition.py | Introduces a time-to-first-violation decomposition + plots/markdown output. |
| scripts/run_posterior_reweighting.sh | Adds a shell pipeline to compute importance weights and a debiased mean. |
| scripts/run_posterior_dynamics_pipeline.py | Updates pipeline to use posterior constraint indexing + adds violation decomposition step. |
| scripts/run_eval_pipeline.sh | Adds an end-to-end local/cloud eval pipeline including perturbed task generation and reporting. |
| scripts/posterior/3_generate_space_time_report.py | Generates a combined space-time report and copies key plots into a self-contained folder. |
| scripts/posterior/1_compute_posterior_weights.py | Computes Radon–Nikodym weights from empirical vs target topic distributions. |
| scripts/generate_perturbed_tasks.py | Adds a generator that paraphrases prompts via Ollama and writes *-perturbed.yaml files. |
| scripts/debiased_evaluation.py | Adds Hajek/IPW aggregation of task scores. |
| scripts/compute_debiased_dynamics.py | Adds IPW/Hajek debiasing over regimes and constraint index. |
| scripts/compute_constraint_index.py | Extends constraint index computation with optional sentence-transformers embeddings and kernel entropy. |
| profiles/user_target_distribution.json | Adds an example target distribution profile. |
| profiles/radon_nikodym_weights.json | Adds example precomputed weights. |
| profiles/empirical_topic_distribution.json | Adds an example empirical benchmark distribution profile. |
| docs/task_distribution_reweighting.md | Documents stratified reweighting and its space-time fusion. |
| docs/semantic_spatiotemporal_dynamics.md | Documents the combined semantic + temporal dynamics framework. |
| docs/long_term_dynamics.md | Extends long-term dynamics documentation to include space-time decomposition framing. |
| clawbench/render.py | Adds render_argv_template() using shlex.split() pre-render to preserve whitespace in substituted values. |
| clawbench/environment_files.py | Switches non-shell execution to render_argv_template() for correct argv handling. |
| clawbench/environment.py | Same argv-template switch for the gateway environment runner. |
| clawbench/dynamics_archive.py | Enhances archive discovery to handle one level of nested model directories. |
| clawbench/dynamics.py | Adds renyi_d2 metric computation to per-trajectory dynamics. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| - message: "Thinking...\nThinking Process:\n\n1. **Analyze the Request:**\n \ | ||
| \ * **Task:** Paraphrase the provided instruction.\n * **Constraint 1:**\ | ||
| \ Keep the exact same semantic meaning and intent.\n * **Constraint 2:**\ | ||
| \ Change the wording slightly.\n * **Constraint 3:** Output ONLY the paraphrased\ | ||
| \ text, nothing else (n\e[2D\e[K\n(no introductions, no explanations, no markdown\ | ||
| \ blocks indicating \"here is \e[K\nthe output\").\n\n2. **Analyze the Original\ |
| @pytest.mark.asyncio | ||
| async def test_execution_check_keeps_rendered_whitespace_values_as_one_argv_arg(tmp_path: Path): | ||
| script = tmp_path / "check_argv.py" | ||
| script.write_text( | ||
| "import json, sys\n" | ||
| "print(json.dumps(sys.argv[1:]))\n", | ||
| encoding="utf-8", | ||
| ) | ||
|
|
||
| result = await run_execution_check( | ||
| ExecutionCheck( | ||
| name="argv-check", | ||
| command="python {script} {output_path}", | ||
| shell=False, | ||
| expected_json=["report 2026.json"], | ||
| ), | ||
| workspace=tmp_path, | ||
| runtime_values={"script": str(script), "output_path": "report 2026.json"}, | ||
| ) | ||
|
|
||
| assert result.passed is True | ||
| assert result.reason == "OK" |
|
|
||
| # Add the repository root to sys.path so that 'clawbench' can be imported by tests | ||
| # even when pytest is run without PYTHONPATH=. | ||
| sys.path.insert(0, str(Path(__file__).parent.parent)) |
| dyn_json = dyn_dir / "dynamics.json" | ||
| if dyn_json.exists(): | ||
| try: | ||
| dyn_data = json.load(open(dyn_json)) |
| import glob | ||
| import subprocess | ||
| import yaml | ||
| import json |
|
|
||
| # For demonstration, limit to a few tasks from different tiers | ||
| # In a full run, we would process all of them | ||
| selected_tasks = yaml_files[:5] |
| - message: Add CSV export functionality to the issue tracker in the workspace. Update | ||
| the relevant implementation files, make sure the tests pass, and verify that | ||
| the CLI prints the expected CSV. | ||
| - message: "Thinking...\nThinking Process:\n\n1. **Analyze the Request:**\n \ |
There was a problem hiding this comment.
Looks like a part of prompt for perturbation was leaked into task.
There was a problem hiding this comment.
thank for the review! will fix that and rerun experiment for this one.
There was a problem hiding this comment.
Check others too: they have the same issue (not all of them)
|
@scoootscooob, review by @clawsweeper is never finished and spams a lot of e-mails. Could you please do something or ping a proper person to fix this issue? |
Adds scenario-conditioned first-violation analysis to the posterior dynamics pipeline, producing per-model JSON, Markdown, and optional plots. Model names containing slashes are preserved by the regime and debiasing scripts; pytest imports are configured in pyproject.toml.
The branch is synchronized with main at c1a79f7. Violation timing now matches recorded forbidden tools, forbidden shell patterns, and dangerous commands. Clean runs are censored at their final observed assistant turn; survival is the cumulative product of hazard complements. Violations that cannot all be localized are counted separately and excluded from timed estimates rather than assigned a fabricated final-turn event. Reports expose risk-set sizes, event counts, and missing-timing limitations.
Validation at 1641734:
Real cached-trajectory validation remains outstanding: locally available live-run cache directories contained no run records. The synthetic smoke run is integration evidence, not representative empirical evidence or a model-safety comparison.