forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 83
Pull requests: PrismML-Eng/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
cuda: MoE prefill — fused SwiGLU epilogue and fused weighted reduction
CUDA
ggml
#107
opened Aug 3, 2026 by
bri-prism
Loading…
5 tasks done
cuda: enable the Hopper wgmma Q1_0/Q2_0 prefill path by default
CUDA
ggml
#106
opened Aug 2, 2026 by
bri-prism
Loading…
4 of 5 tasks
Refactor CUDA workflow and remove ROCm steps
devops
documentation
Improvements or additions to documentation
#105
opened Aug 1, 2026 by
MasterCool4389
Loading…
POWER10/11 MMA acceleration for all quantized formats (silicon-validated)
ggml
#100
opened Jul 23, 2026 by
mavin2009
Loading…
speculative: window dspark drafter staging to its trained position range
documentation
Improvements or additions to documentation
server
testing
#98
opened Jul 22, 2026 by
thadreber-web
Loading…
tools/ui: Respect configured MCP transport instead of always using Streamable HTTP
server/ui
#97
opened Jul 21, 2026 by
joydolma
Loading…
Gb10 cuda graph fix
CUDA
documentation
Improvements or additions to documentation
ggml
#96
opened Jul 20, 2026 by
sumergoconicio
Loading…
ggml-cpu: add opt-in Q2_0 VNNI64 four-row decode
documentation
Improvements or additions to documentation
ggml
#95
opened Jul 20, 2026 by
chris-lee-mc
•
Draft
gb10-blackwell: env-gated Blackwell int8 MMA for Q1_0/Q2_0 weights on DGX Spark
CUDA
documentation
Improvements or additions to documentation
ggml
#79
opened Jul 16, 2026 by
sumergoconicio
Loading…
ggml-cpu: enable Q2_0 VNNI kernel on AVX-VNNI-only CPUs
ggml
#76
opened Jul 15, 2026 by
gondoi
Loading…
opencl: Q1_0 support first attempt
ggml
OpenCL
#25
opened Apr 15, 2026 by
khosravipasha
Collaborator
•
Draft
ProTip!
Filter pull requests by the default branch with base:prism.