ik_llama.cpp

History

Kawrakow fc06bc9d27 Enable CUDA graphs for MoE models + GPT-OSS support (#689 ) * gmp-oss: common * gpt-oss: attnetion sinks, swiglu_oai * gpt-oss: WIP llama Model loads and runs (CPU only), but PPL is much to high (~1500 for 1st batch vs ~200 in mainline). Is it because of SWA, because of vocab, or did I introduce a bug somewhere? * gpt-oss: CPU seems to be working It was the SWA thta was missing in the previous commit. There are issues with EOG tokens, so this still needs to be added. * CUDA: ADD_ID Just a copy from mainline * gpt-oss: Seems to be working on CUDA * gpt-oss: add sinks to the attn-vec kernels * CUDA: add head size of 64 to new mma Haven't turned it on yet, but observe slightly better PP and slightly worse TG performance with that. * gpt-oss: add ability to use -fmoe (only CUDA for now) * Move row sums to the write place * Add sinks to iqk flash attention * gpt_oss: Implement -fmoe on the CPU * Simdify swiglu_oai Turning it off for now as performance becomes more variable, so perhaps I'm running into thermal trottling imore often because of making the CPU work too hard. * llama: factor out model loader * Builds successfully * It runs, but mmap does not work * Fix llama_mmap so mmap works * Minor * Fix CUDA after latest changes * Attempt to use CUDA graphs with MoE models - not working * CUDA graphs WIP - still not working * CUDA graphs - seems to be working Likely not all MLA variants are working. I no longer remember why I added the q8_0 cpy that transposes the tensor, but if really needed, this is now missing. Also missing is q6_0. * Make q8_0 cache work for DeepSeek models with CUDA graphs * cuda: cpy for q6_0 * Fix llama_mmap on non-Linux platforms * Adding forgotten file * Iterating on Windows build failures * cuda: re-add q8_0 -> q8_0 transpose so mla = 2 can be used with CUDA graphs and q8_0 cache. * Disable graphs without -fmoe * Minor * Turn graphs on by default --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>		2025-08-15 09:18:07 +03:00
..
cmake	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
base64.hpp	llava : expose as a shared library for downstream projects (#3613 )	2023-11-07 00:36:23 +03:00
build-info.cpp.in	build : link against build info instead of compiling against it (#3879 )	2023-11-02 08:50:16 +02:00
chat-parser.cpp	Fix for Deepseek r1 parsing (#676 )	2025-08-08 13:56:44 +03:00
chat-parser.h	Enable CUDA graphs for MoE models + GPT-OSS support (#689 )	2025-08-15 09:18:07 +03:00
chat-template.hpp	add jinja template support (#677 )	2025-08-09 12:50:30 +00:00
chat.cpp	Enable CUDA graphs for MoE models + GPT-OSS support (#689 )	2025-08-15 09:18:07 +03:00
chat.h	Enable CUDA graphs for MoE models + GPT-OSS support (#689 )	2025-08-15 09:18:07 +03:00
CMakeLists.txt	add jinja template support (#677 )	2025-08-09 12:50:30 +00:00
common.cpp	add jinja template support (#677 )	2025-08-09 12:50:30 +00:00
common.h	add jinja template support (#677 )	2025-08-09 12:50:30 +00:00
console.cpp	check C++ code with -Wmissing-declarations (#3184 )	2023-09-15 15:38:27 -04:00
console.h	gguf : new file format with flexible meta data (beta) (#2398 )	2023-08-21 23:07:43 +03:00
grammar-parser.cpp	Added support for . (any character) token in grammar engine. (#6467 )	2024-06-06 06:08:52 -07:00
grammar-parser.h	gguf : new file format with flexible meta data (beta) (#2398 )	2023-08-21 23:07:43 +03:00
json-partial.cpp	Function calling support for Kimi-K2 (#628 )	2025-07-23 18:11:42 +02:00
json-partial.h	Function calling support for Kimi-K2 (#628 )	2025-07-23 18:11:42 +02:00
json-schema-to-grammar.cpp	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
json-schema-to-grammar.h	JSON: [key] -> .at(key), assert() -> GGML_ASSERT (#7143 )	2024-05-08 21:53:08 +02:00
json.hpp	json-schema-to-grammar improvements (+ added to server) (#5978 )	2024-03-21 11:50:43 +00:00
log.h	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
minja.hpp	add jinja template support (#677 )	2025-08-09 12:50:30 +00:00
ngram-cache.cpp	Fixed lookup compilation issues on Windows (#6273 )	2024-03-24 14:21:17 +01:00
ngram-cache.h	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
regex-partial.cpp	Function calling support for Kimi-K2 (#628 )	2025-07-23 18:11:42 +02:00
regex-partial.h	Function calling support for Kimi-K2 (#628 )	2025-07-23 18:11:42 +02:00
sampling.cpp	Do not crash when there is no DRY sampler (#578 )	2025-07-03 15:26:52 +02:00
sampling.h	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
stb_image.h	examples: support LLaVA v1.5 (multimodal model) (#3436 )	2023-10-12 18:23:18 +03:00
train.cpp	train : change default FA argument (#7528 )	2024-05-25 15:22:35 +03:00
train.h	sync : ggml (backend v2) (#3912 )	2023-11-13 14:16:23 +02:00