Fix guided generation: don't leak grammar terminal token on GPU path - #183
Draft
stikves wants to merge 2 commits into
Draft
Fix guided generation: don't leak grammar terminal token on GPU path#183stikves wants to merge 2 commits into
stikves wants to merge 2 commits into
Conversation
The GPU constrained decode loop (runConstrainedCompletion) yielded each
sampled token to the consumer *before* asking the grammar whether that
token terminated it. When the grammar terminates on a stop token (e.g.
<|endoftext|>, 151643), that token had already been streamed and got
decoded into the structured output, producing trailing text like:
{...}<|endoftext|>
This is the GPU analogue of the CPU bug fixed in apple#117, which only patched
ConstrainedDecodingStrategy. The GPU path added later in apple#170 reintroduced
the same class of bug and was never covered by apple#117's fix.
Defer the yield: a sampled token is now emitted only after the grammar has
accepted it and we've confirmed the acceptance did not terminate. This
matches the accept -> isTerminated -> yield ordering already used by the
test mock (MockConstrainedEngine).
Verified with llm-runner --json-schema against Qwen3-0.6B (greedy):
before: {...}<|endoftext|>
after: {...}
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The GPU constrained decode loop (runConstrainedCompletion) yielded each sampled token to the consumer before asking the grammar whether that token terminated it. When the grammar terminates on a stop token (e.g. <|endoftext|>, 151643), that token had already been streamed and got decoded into the structured output, producing trailing text like:
This is the GPU analogue of the CPU bug fixed in #117, which only patched ConstrainedDecodingStrategy. The GPU path added later in #170 reintroduced the same class of bug and was never covered by #117's fix.
Defer the yield: a sampled token is now emitted only after the grammar has accepted it and we've confirmed the acceptance did not terminate. This matches the accept -> isTerminated -> yield ordering already used by the test mock (MockConstrainedEngine).
Verified with llm-runner --json-schema against Qwen3-0.6B (greedy):
before: {...}<|endoftext|>
after: {...}