-
Notifications
You must be signed in to change notification settings - Fork 21k
Pull requests: ggml-org/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
tests : remove unnecessary sync in test-save-load-state
testing
Everything test related
#26166
opened Jul 27, 2026 by
ggerganov
Member
Loading…
common: fix explicit -md precedence over draft sidecar resolution
#26165
opened Jul 27, 2026 by
ServeurpersoCom
Contributor
Loading…
[Model/VLA] Support MiniCPM-RobotManip
build
Compilation issues
documentation
Improvements or additions to documentation
mtmd
Related to multimodal functionality (video/image/audio)
server
opencl: bugfix increment ref_count in ggml_backend_opencl_init()
ggml
changes relating to the ggml tensor library for machine learning
OpenCL
Issues specific to the OpenCL backend
#26162
opened Jul 27, 2026 by
akleine
Contributor
Loading…
ggml: add support for MXFP8 CPU
conversion
ggml
changes relating to the ggml tensor library for machine learning
#26157
opened Jul 27, 2026 by
michaelw9999
Contributor
Loading…
mtmd: support multi-row batching for deepseek-ocr
mtmd
Related to multimodal functionality (video/image/audio)
#26154
opened Jul 26, 2026 by
ngxson
Collaborator
Loading…
skills: add ggml-test skill
documentation
Improvements or additions to documentation
#26153
opened Jul 26, 2026 by
ngxson
Collaborator
Loading…
Make the WebGPU backend pass test-llama-archs tests
ggml
changes relating to the ggml tensor library for machine learning
model
Model specific
testing
Everything test related
WebGPU
#26146
opened Jul 26, 2026 by
fairydreaming
Collaborator
•
Draft
Responses API: Convert text.format JSON schema into the corresponding Chat Completions response_format
server
testing
Everything test related
#26145
opened Jul 26, 2026 by
boondocklabs
Contributor
Loading…
ggml-cuda : fall back to cuBLAS when no MMQ tile size fits in shared memory
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
#26141
opened Jul 26, 2026 by
KakaruHayate
Loading…
feat: QFX16/QFX32 lossless quantization (2.05x compression, zero quality loss)
Apple Metal
https://en.wikipedia.org/wiki/Metal_(API)
examples
ggml
changes relating to the ggml tensor library for machine learning
#26136
opened Jul 26, 2026 by
theQarchitect
Loading…
ggml-webgpu: simplify flash_attn shaders and refactor several WGSL shaders
ggml
changes relating to the ggml tensor library for machine learning
WebGPU
ggml-cpu : fix conv_transpose_2d for multiple batches
ggml
changes relating to the ggml tensor library for machine learning
testing
Everything test related
#26132
opened Jul 26, 2026 by
tekinertekin
Contributor
Loading…
server : expose per-device memory usage on /metrics and GET /memory
documentation
Improvements or additions to documentation
mtmd
Related to multimodal functionality (video/image/audio)
server
#26130
opened Jul 26, 2026 by
tobocop2
Loading…
vulkan: extend topk_moe fusion to support sqrt(softplus)
ggml
changes relating to the ggml tensor library for machine learning
Vulkan
Issues specific to the Vulkan backend
#26124
opened Jul 25, 2026 by
jeffbolznv
Contributor
Loading…
common: Add CLI > ENV > models-presets > INI precedence
documentation
Improvements or additions to documentation
#26118
opened Jul 25, 2026 by
jcmdln
Loading…
Fix speculative models failing to load when running llama-server
#26114
opened Jul 25, 2026 by
sheldonrobinson
Contributor
Loading…
hexagon: enable quantized matmul for multi-sequence inputs
ggml
changes relating to the ggml tensor library for machine learning
Hexagon
#26113
opened Jul 25, 2026 by
w1049
Contributor
Loading…
CUDA : add warp-per-row WKV7 kernel for single-token decode
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
testing
Everything test related
#26111
opened Jul 25, 2026 by
123123213weqw
Loading…
Feature/e2k support
ggml
changes relating to the ggml tensor library for machine learning
#26107
opened Jul 25, 2026 by
png-tech
Loading…
sycl: fix classification of iGPUs
ggml
changes relating to the ggml tensor library for machine learning
SYCL
https://en.wikipedia.org/wiki/SYCL - GPU programming language
#26105
opened Jul 25, 2026 by
KyleHagy
Contributor
Loading…
cmake : fall back to empty UI on stale assets when provisioning is disabled
server/ui
testing
Everything test related
#26104
opened Jul 25, 2026 by
wuisabel-gif
•
Draft
1 task done
ggml-cpu: skip unsupported ARM ISA variants under GGML_CPU_ALL_VARIANTS
ggml
changes relating to the ggml tensor library for machine learning
#26103
opened Jul 25, 2026 by
ninihuang2026
Loading…
Previous Next
ProTip!
Follow long discussions with comments:>50.