Skip to content

Pull requests: ggml-org/llama.cpp

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

tests : remove unnecessary sync in test-save-load-state testing Everything test related
#26166 opened Jul 27, 2026 by ggerganov Member Loading…
[Model/VLA] Support MiniCPM-RobotManip build Compilation issues documentation Improvements or additions to documentation mtmd Related to multimodal functionality (video/image/audio) server
#26164 opened Jul 27, 2026 by tc-mb Contributor Draft
opencl: bugfix increment ref_count in ggml_backend_opencl_init() ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend
#26162 opened Jul 27, 2026 by akleine Contributor Loading…
cuda: fuse MoE expert weighting and reduction (+3% to +9% prefill, MoE models only) CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#26160 opened Jul 27, 2026 by coder543 Contributor Draft
cuda: compact Blackwell NVFP4 MoE work scheduling (+10% to +15% prefill, NVFP4-only) CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#26159 opened Jul 27, 2026 by coder543 Contributor Draft
ggml: add support for MXFP8 CPU conversion ggml changes relating to the ggml tensor library for machine learning
#26157 opened Jul 27, 2026 by michaelw9999 Contributor Loading…
mtmd: support multi-row batching for deepseek-ocr mtmd Related to multimodal functionality (video/image/audio)
#26154 opened Jul 26, 2026 by ngxson Collaborator Loading…
skills: add ggml-test skill documentation Improvements or additions to documentation
#26153 opened Jul 26, 2026 by ngxson Collaborator Loading…
Make the WebGPU backend pass test-llama-archs tests ggml changes relating to the ggml tensor library for machine learning model Model specific testing Everything test related WebGPU
#26146 opened Jul 26, 2026 by fairydreaming Collaborator Draft
ggml-cuda : fall back to cuBLAS when no MMQ tile size fits in shared memory CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#26141 opened Jul 26, 2026 by KakaruHayate Loading…
feat: QFX16/QFX32 lossless quantization (2.05x compression, zero quality loss) Apple Metal https://en.wikipedia.org/wiki/Metal_(API) examples ggml changes relating to the ggml tensor library for machine learning
#26136 opened Jul 26, 2026 by theQarchitect Loading…
ggml-webgpu: simplify flash_attn shaders and refactor several WGSL shaders ggml changes relating to the ggml tensor library for machine learning WebGPU
#26134 opened Jul 26, 2026 by yomaytk Member Draft
ggml-cpu : fix conv_transpose_2d for multiple batches ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#26132 opened Jul 26, 2026 by tekinertekin Contributor Loading…
server : expose per-device memory usage on /metrics and GET /memory documentation Improvements or additions to documentation mtmd Related to multimodal functionality (video/image/audio) server
#26130 opened Jul 26, 2026 by tobocop2 Loading…
vulkan: extend topk_moe fusion to support sqrt(softplus) ggml changes relating to the ggml tensor library for machine learning Vulkan Issues specific to the Vulkan backend
#26124 opened Jul 25, 2026 by jeffbolznv Contributor Loading…
common: Add CLI > ENV > models-presets > INI precedence documentation Improvements or additions to documentation
#26118 opened Jul 25, 2026 by jcmdln Loading…
Fix speculative models failing to load when running llama-server
#26114 opened Jul 25, 2026 by sheldonrobinson Contributor Loading…
hexagon: enable quantized matmul for multi-sequence inputs ggml changes relating to the ggml tensor library for machine learning Hexagon
#26113 opened Jul 25, 2026 by w1049 Contributor Loading…
CUDA : add warp-per-row WKV7 kernel for single-token decode CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#26111 opened Jul 25, 2026 by 123123213weqw Loading…
Feature/e2k support ggml changes relating to the ggml tensor library for machine learning
#26107 opened Jul 25, 2026 by png-tech Loading…
sycl: fix classification of iGPUs ggml changes relating to the ggml tensor library for machine learning SYCL https://en.wikipedia.org/wiki/SYCL - GPU programming language
#26105 opened Jul 25, 2026 by KyleHagy Contributor Loading…
ggml-cpu: skip unsupported ARM ISA variants under GGML_CPU_ALL_VARIANTS ggml changes relating to the ggml tensor library for machine learning
#26103 opened Jul 25, 2026 by ninihuang2026 Loading…
ProTip! Follow long discussions with comments:>50.