Skip to content

Fix guided generation: don't leak grammar terminal token on GPU path - #183

Draft
stikves wants to merge 2 commits into
apple:mainfrom
cgreening:fix/gpu-constrained-terminal-token-leak
Draft

Fix guided generation: don't leak grammar terminal token on GPU path#183
stikves wants to merge 2 commits into
apple:mainfrom
cgreening:fix/gpu-constrained-terminal-token-leak

Conversation

@stikves

@stikves stikves commented Aug 19, 2026

Copy link
Copy Markdown
Contributor

The GPU constrained decode loop (runConstrainedCompletion) yielded each sampled token to the consumer before asking the grammar whether that token terminated it. When the grammar terminates on a stop token (e.g. <|endoftext|>, 151643), that token had already been streamed and got decoded into the structured output, producing trailing text like:

{...}<|endoftext|>

This is the GPU analogue of the CPU bug fixed in #117, which only patched ConstrainedDecodingStrategy. The GPU path added later in #170 reintroduced the same class of bug and was never covered by #117's fix.

Defer the yield: a sampled token is now emitted only after the grammar has accepted it and we've confirmed the acceptance did not terminate. This matches the accept -> isTerminated -> yield ordering already used by the test mock (MockConstrainedEngine).

Verified with llm-runner --json-schema against Qwen3-0.6B (greedy):
before: {...}<|endoftext|>
after: {...}

Chris Greening and others added 2 commits August 19, 2026 15:59
The GPU constrained decode loop (runConstrainedCompletion) yielded each
sampled token to the consumer *before* asking the grammar whether that
token terminated it. When the grammar terminates on a stop token (e.g.
<|endoftext|>, 151643), that token had already been streamed and got
decoded into the structured output, producing trailing text like:

    {...}<|endoftext|>

This is the GPU analogue of the CPU bug fixed in apple#117, which only patched
ConstrainedDecodingStrategy. The GPU path added later in apple#170 reintroduced
the same class of bug and was never covered by apple#117's fix.

Defer the yield: a sampled token is now emitted only after the grammar has
accepted it and we've confirmed the acceptance did not terminate. This
matches the accept -> isTerminated -> yield ordering already used by the
test mock (MockConstrainedEngine).

Verified with llm-runner --json-schema against Qwen3-0.6B (greedy):
  before: {...}<|endoftext|>
  after:  {...}
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants