Skip to content

Pull requests: xLLM-AI/xllm

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

Fix MTP TPOT latency accounting
#2131 opened Aug 5, 2026 by Rellion-926 Loading…
9 of 17 tasks
feat: add mlu linear prefix cache support.
#2130 opened Aug 5, 2026 by phantomlei3 Collaborator Loading…
9 of 17 tasks
refactor: consolidate KV cache transfer around MoonCake.
#2129 opened Aug 5, 2026 by JimHsiung Collaborator Loading…
17 tasks
feat: add reliable runtime lifecycle for EPLB expert rebalancing.
#2127 opened Aug 5, 2026 by DongheJin Collaborator Loading…
17 tasks
feat: support DeepSeek V4 0731 reasoning effort handling and template.
#2126 opened Aug 5, 2026 by chenchuw886 Contributor Loading…
9 of 17 tasks
feat: add decode context parallel (DCP) for Qwen3.5 GQA dense.
#2125 opened Aug 5, 2026 by Enguikong Contributor Draft
6 of 13 tasks
feat: add optimized MUSA kernels.
#2123 opened Aug 5, 2026 by fay85 Collaborator Loading…
9 of 17 tasks
feat: support DeepSeek-V3.2 Python MTP ACLGraph.
#2121 opened Aug 4, 2026 by yt-zeng Collaborator Loading…
1 of 7 tasks
feat: add MUSA graph executor.
#2117 opened Aug 4, 2026 by fay85 Collaborator Loading…
7 of 17 tasks
feat: add json object constrained decoding.
#2110 opened Aug 3, 2026 by DragonFive Collaborator Loading…
17 tasks
feat: add json object constrained decoding.
#2102 opened Aug 3, 2026 by DragonFive Collaborator Loading…
7 of 17 tasks
feat: add VLM MTP backend support for speculative decoding.
#2099 opened Aug 3, 2026 by phantomlei3 Collaborator Loading…
9 of 17 tasks
feat: support host prefix cache offload for DeepSeek V4.
#2096 opened Aug 1, 2026 by Kang-Meng Collaborator Loading…
8 of 17 tasks
feat: add pd prefill short request first scheduling.
#2092 opened Jul 31, 2026 by immengzi Loading…
10 of 17 tasks
feat: add Qwen3.5 Python model cuda impl.
#2090 opened Jul 30, 2026 by zhang-minchao Collaborator Draft
7 of 17 tasks
build: optimize native and tilelang incremental builds.
#2068 opened Jul 29, 2026 by jarvis666666 Loading…
10 of 17 tasks
bugfix: prevent incompatible npu graph replays.
#2065 opened Jul 28, 2026 by shifengmin Collaborator Loading…
7 of 17 tasks
feat(lora): prefix cache adapter isolation (PR 3/3)
#2047 opened Jul 27, 2026 by cchh05 Collaborator Loading…
3 of 5 tasks
feat: support decode context parallel.
#2021 opened Jul 23, 2026 by phantomlei3 Collaborator Draft
17 tasks
feat(lora): request path adapter_id propagation + chat routing (2/2)
#2015 opened Jul 23, 2026 by cchh05 Collaborator Loading…
feat(lora): multi-tenant LoRA framework + attention/MLP wire-up (1/2)
#2014 opened Jul 23, 2026 by cchh05 Collaborator Loading…
feat: add mlu host kv cache support.
#2012 opened Jul 23, 2026 by phantomlei3 Collaborator Draft
17 tasks
ProTip! no:milestone will show everything without a milestone.