docs(triage): #132 parameter-description evolution — investigated, null - #141
Merged
Conversation
A measure-first kill-gate (precise vs. name-only param descriptions, capable + weak models) found zero separation: the description text doesn't move the agent's value-selection — it's inferred from the task + param name. Closed #132 as investigated->null rather than building the evolution/parameters/ subsystem; ~$1 spent. The probe is a local spike (uncommitted per convention); the finding + numbers live in the review log.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes triage item #132 as investigated → null rather than building the proposed
evolution/parameters/subsystem.A measure-first kill-gate compared precise vs. name-only parameter descriptions on the case most favorable to signal — the hermes
terminalbackground/notify_on_completeconvention, carried entirely in description prose — scoring the agent's supplied argument values:gpt-5.4-nano, reps=3)The non-saturated weak model had room to benefit from the guidance and didn't — the description text does not move value-selection; the agent infers it from the task + param name. So a parameter-description evolver has nothing to optimize toward. ~$1 spent vs. a six-file build.
Docs-only: updates
docs/upstream_pr_triage.md(#132 row, review log, snapshot). The probe is a local spike underspikes/param_probe/(uncommitted per convention; reproducible viaprobe.py).