llama.cpp

mirror of https://github.com/ggerganov/llama.cpp synced 2026-05-02 12:22:01 +02:00

History

Masashi Yoshimura 6da7168312 ggml-webgpu: Add fused RMS_NORM + MUL (#21983 ) * fused rms_norm_mul + mul * Add GGML_WEBGPU_DISABLE_FUSION for being able to disable kernel fusion. * Decouple num_fused_ops from webgpu_context; misc cleanup * Fix eps handling and remove disable_fusion. * Fix not to use c++20 initializers.		2026-04-22 10:51:40 -07:00
..
cmake	ggml: backend-agnostic tensor parallelism (experimental) (#19378 )	2026-04-09 16:42:19 +02:00
include	CUDA: manage NCCL communicators in context (#21891 )	2026-04-15 15:58:40 +02:00
src	ggml-webgpu: Add fused RMS_NORM + MUL (#21983 )	2026-04-22 10:51:40 -07:00
.gitignore	vulkan : cmake integration (#8119 )	2024-07-13 18:12:39 +02:00
CMakeLists.txt	ggml : bump version to 0.10.0 (ggml/1463)	2026-04-21 11:04:21 +03:00