-
Notifications
You must be signed in to change notification settings - Fork 22.3k
Pull requests: ggml-org/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Blackwell IQ Failures issue Fixes
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
#27902
opened Aug 28, 2026 by
matteius
Loading…
vulkan : fix target-env for non-_cm2-named shaders needing vulkan1.3
ggml
changes relating to the ggml tensor library for machine learning
Vulkan
Issues specific to the Vulkan backend
#27900
opened Aug 28, 2026 by
skywalk1411
•
Draft
speculative: fix combined draft-mtp + external draft (-md) init crash
server
#27897
opened Aug 28, 2026 by
simongonzalezdc
•
Draft
cuda: enable GGML_CUDA_FA_ALL_QUANTS by default
ggml
changes relating to the ggml tensor library for machine learning
#27895
opened Aug 28, 2026 by
47Hunter47
•
Draft
2 of 4 tasks
ggml: avoid KleidiAI init mutex on dispatch (#27078)
ggml
changes relating to the ggml tensor library for machine learning
#27891
opened Aug 28, 2026 by
ac-mmi
Loading…
ci : check for missing autoreleasepools
devops
improvements to build systems and github actions
#27884
opened Aug 28, 2026 by
nikwen
Member
•
2/2
Loading…
metal : fix more leaks due to missing autoreleasepools
Apple Metal
https://en.wikipedia.org/wiki/Metal_(API)
ggml
changes relating to the ggml tensor library for machine learning
merge ready
A maintainer can use this label to indicate that they consider the changes final and ready to merge.
#27883
opened Aug 28, 2026 by
nikwen
Member
•
1/2
Loading…
bench: add --tensor-read-lazy
documentation
Improvements or additions to documentation
examples
#27881
opened Aug 28, 2026 by
ngxson
Collaborator
Loading…
Qwen4exp correctness fixes
Apple Metal
https://en.wikipedia.org/wiki/Metal_(API)
ggml
changes relating to the ggml tensor library for machine learning
model
Model specific
testing
Everything test related
ggml-cuda: fix divergent barrier in f16 flash attention
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
#27870
opened Aug 28, 2026 by
siavashnorouzi
Loading…
mtmd: reject zero granite4-vision downsample sides
mtmd
Related to multimodal functionality (video/image/audio)
spec : reset ngram-cache state between requests
testing
Everything test related
#27866
opened Aug 28, 2026 by
Beatrice0377
Loading…
llama: GPU-resident LRU cache for host-offloaded MoE expert weights
ggml
changes relating to the ggml tensor library for machine learning
#27861
opened Aug 28, 2026 by
csantiago78
•
Draft
fix assertion DFlash2 with Model specific
--split-mode tensor on CUDA
model
#27858
opened Aug 28, 2026 by
art-den
Loading…
ggml-backend: check return value of ggml_gallocr_reserve_n in alloc_splits
ggml
changes relating to the ggml tensor library for machine learning
#27855
opened Aug 28, 2026 by
kritikagarg
Loading…
ggml-cpu: tiled mul_mat for k-quants
ggml
changes relating to the ggml tensor library for machine learning
testing
Everything test related
#27851
opened Aug 28, 2026 by
jbooth
Loading…
sycl: split long rows in TOP_K instead of one workgroup per row
ggml
changes relating to the ggml tensor library for machine learning
SYCL
https://en.wikipedia.org/wiki/SYCL - GPU programming language
#27847
opened Aug 28, 2026 by
Titaniumtown
Contributor
Loading…
ggml-cuda: hip: add missing AMD GCN MMQ config
CUDA
Related to the CUDA backend
ggml
changes relating to the ggml tensor library for machine learning
testing
Everything test related
#27841
opened Aug 28, 2026 by
thelittlefireman
Loading…
clamp video fps to the value of new --video-max-tokens cli flag
mtmd
Related to multimodal functionality (video/image/audio)
server
#27838
opened Aug 28, 2026 by
DipCrai
Loading…
qwen4exp : add NextN/MTP draft head (--spec-type draft-mtp) for Qwen3.8-Flash-Next
conversion
model
Model specific
#27836
opened Aug 27, 2026 by
rmonsurate
•
Draft
1 task done
Previous Next
ProTip!
Mix and match filters to narrow down what you’re looking for.