llama.cpp

mirror of https://github.com/ggerganov/llama.cpp synced 2026-04-28 18:30:16 +02:00

History

Adrien Gallouët 5dd102539b server : ignore --alias when using --models-preset (#21380 ) I'm not sure what the purpose of keeping `--alias` was when using `--models-preset`, but the result is really weird, as shown in the following logs: $ build/bin/llama-server --models-preset preset.ini --alias "Gemma 4 E4B UD Q8_K_XL" ... init: using 31 threads for HTTP server srv load_models: Loaded 2 cached model presets srv load_models: Loaded 1 custom model presets from preset.ini main: failed to initialize router models: alias 'Gemma 4 E4B UD Q8_K_XL' for model 'angt/test-split-model-stories260K:F32' conflicts with existing model name So I propose to simply ignore `--alias` too in this case. With this commit, the server starts in routing mode correctly. Signed-off-by: Adrien Gallouët <angt@huggingface.co>		2026-04-10 17:42:56 +02:00
..
batched-bench	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
cli	server: save and clear idle slots on new task (`--clear-idle`) (#20993 )	2026-04-03 19:02:27 +02:00
completion	server: save and clear idle slots on new task (`--clear-idle`) (#20993 )	2026-04-03 19:02:27 +02:00
cvector-generator	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
export-lora	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
fit-params	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
gguf-split	gguf-split : clarify operation of gguf-split (#19749 )	2026-03-25 13:12:50 +02:00
imatrix	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
llama-bench	ggml: backend-agnostic tensor parallelism (experimental) (#19378 )	2026-04-09 16:42:19 +02:00
mtmd	mtmd: support dots.ocr (#17575 )	2026-04-09 12:16:38 +02:00
parser	common/parser: fix call ID detection (Mistral parser mostly) + atomicity for tag-json parsers (#21230 )	2026-04-03 17:51:52 +02:00
perplexity	ggml: backend-agnostic tensor parallelism (experimental) (#19378 )	2026-04-09 16:42:19 +02:00
quantize	ggml: add Q1_0 1-bit quantization support (CPU) (#21273 )	2026-04-06 20:55:21 +02:00
results	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
rpc	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
server	server : ignore --alias when using --models-preset (#21380 )	2026-04-10 17:42:56 +02:00
tokenize	Fix locale-dependent float printing in GGUF metadata (#17331 )	2026-03-04 09:30:40 +01:00
tts	common : move up common_init() and fix Windows UTF-8 logs (#21176 )	2026-03-31 12:53:41 +02:00
CMakeLists.txt	llama: end-to-end tests (#19802 )	2026-03-08 12:30:21 +01:00