fix(qwen3.8): wire-format detect nvfp4 artifact profile - #107
Open
koloved wants to merge 1 commit into
Open
Conversation
Qwen3.8-27B NVFP4 ships in two wire formats: the upstream FP8 profile (row-scaled FP8 vocabulary/attention endpoints, 21 GB artifact) and the community W8+NVFP4 profile (W8 vocabulary endpoints, Qwen3.6 NVFP4 layer layout, 18.3 GB artifact). Resolve the weights profile from the artifact's text/token_embedding descriptor instead of always selecting the FP8 profile. The W8+NVFP4 profile comes from Ostfralla/Qwen3.8-27B-NVFP4-NInfer (https://huggingface.co/Ostfralla/Qwen3.8-27B-NVFP4-NInfer) and runs significantly faster on an RTX 5090 while leaving more VRAM, so it supports a larger context (262144 tokens at full speed) than the 21 GB upstream FP8 artifact, which cannot reach those speeds on the same hardware. Existing artifacts are unaffected: the upstream FP8 model still resolves to the FP8 profile (Qwen38Nvfp4) because its token_embedding is row-scaled FP8; only W8-embedded community artifacts are routed to Qwen36Nvfp4.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: b03557e82a
ℹ️ About Codex in GitHub
Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".
This was referenced Aug 28, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Qwen3.8-27B NVFP4 ships in two wire formats: the upstream FP8 profile (row-scaled FP8 vocabulary/attention endpoints, 21 GB artifact) and the community W8+NVFP4 profile (W8 vocabulary endpoints, Qwen3.6 NVFP4 layer layout, 18.3 GB artifact). Resolve the weights profile from the artifact's text/token_embedding descriptor instead of always selecting the FP8 profile.
The W8+NVFP4 profile comes from Ostfralla/Qwen3.8-27B-NVFP4-NInfer (https://huggingface.co/Ostfralla/Qwen3.8-27B-NVFP4-NInfer) and runs significantly faster on an RTX 5090 while leaving more VRAM, so it supports a larger context (262144 tokens at full speed) than the 21 GB upstream FP8 artifact, which cannot reach those speeds on the same hardware.
Existing artifacts are unaffected: the upstream FP8 model still resolves to the FP8 profile (Qwen38Nvfp4) because its token_embedding is row-scaled FP8; only W8-embedded community artifacts are routed to Qwen36Nvfp4.