Skip to content

fix(qwen3.8): wire-format detect nvfp4 artifact profile - #107

Open
koloved wants to merge 1 commit into
Neroued:masterfrom
koloved:master
Open

fix(qwen3.8): wire-format detect nvfp4 artifact profile#107
koloved wants to merge 1 commit into
Neroued:masterfrom
koloved:master

Conversation

@koloved

@koloved koloved commented Aug 28, 2026

Copy link
Copy Markdown

Qwen3.8-27B NVFP4 ships in two wire formats: the upstream FP8 profile (row-scaled FP8 vocabulary/attention endpoints, 21 GB artifact) and the community W8+NVFP4 profile (W8 vocabulary endpoints, Qwen3.6 NVFP4 layer layout, 18.3 GB artifact). Resolve the weights profile from the artifact's text/token_embedding descriptor instead of always selecting the FP8 profile.

The W8+NVFP4 profile comes from Ostfralla/Qwen3.8-27B-NVFP4-NInfer (https://huggingface.co/Ostfralla/Qwen3.8-27B-NVFP4-NInfer) and runs significantly faster on an RTX 5090 while leaving more VRAM, so it supports a larger context (262144 tokens at full speed) than the 21 GB upstream FP8 artifact, which cannot reach those speeds on the same hardware.

Existing artifacts are unaffected: the upstream FP8 model still resolves to the FP8 profile (Qwen38Nvfp4) because its token_embedding is row-scaled FP8; only W8-embedded community artifacts are routed to Qwen36Nvfp4.

Qwen3.8-27B NVFP4 ships in two wire formats: the upstream FP8 profile
(row-scaled FP8 vocabulary/attention endpoints, 21 GB artifact) and the
community W8+NVFP4 profile (W8 vocabulary endpoints, Qwen3.6 NVFP4 layer
layout, 18.3 GB artifact). Resolve the weights profile from the artifact's
text/token_embedding descriptor instead of always selecting the FP8 profile.

The W8+NVFP4 profile comes from Ostfralla/Qwen3.8-27B-NVFP4-NInfer
(https://huggingface.co/Ostfralla/Qwen3.8-27B-NVFP4-NInfer) and runs
significantly faster on an RTX 5090 while leaving more VRAM, so it supports
a larger context (262144 tokens at full speed) than the 21 GB upstream
FP8 artifact, which cannot reach those speeds on the same hardware.

Existing artifacts are unaffected: the upstream FP8 model still resolves to
the FP8 profile (Qwen38Nvfp4) because its token_embedding is row-scaled FP8;
only W8-embedded community artifacts are routed to Qwen36Nvfp4.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: b03557e82a

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment thread src/targets/qwen3_6_27b/impl/package.cpp
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant