ik_llama.cpp

History

firecoperana 52adcf1e90 Update grammar (#1023 ) * grammar : fix JSON Schema for string regex with top-level alt. (#9903) Prior to this commit, using a JSON Schema containing a string with `pattern` regular expression that uses top-level alternation (e.g. `"pattern": "^A\|B\|C\|D$"`) would result in invalid JSON output from the constrained sampling grammar, because it ended up creating a grammar rule like this for the string: ``` thing ::= "\"" "A" \| "B" \| "C" \| "D" "\"" space ``` Note that this rule will only match a starting quote for the "A" case, and will only match an ending quote for the "D" case, so this rule will always produce invalid JSON when used for sampling (that is, the JSON will always be lacking the starting quote, the ending quote, or both). This was fixed in a simple way by adding parentheses to the generated rule (for all string pattern rules, to keep it simple), such that the new generated rule looks like this (correct): ``` thing ::= "\"" ("A" \| "B" \| "C" \| "D") "\"" space ``` * grammars : add English-only grammar (#10612) * grammar : handle maxItems == 0 in JSON schema (#13117) Co-authored-by: Richard Lyons <frob@cloudstaff.com> * grammar-parser : fix possible null-deref (#9004) Fixes: https://bugs.chromium.org/p/oss-fuzz/issues/detail?id=70680 Signed-off-by: David Korczynski <david@adalogics.com> * llama : fix typo in llama-grammar.h [no ci] (#11816) * * server: fix "--grammar-file" parameter (#12285) * common : use std::string_view now that we target c++17 (#14319) * json : support `enum` values within `allOf` (#15830) * grammar : use int64_t to avoid int overflows in int schema to grammar conversion logic (#16626) * grammar : support array references in json schema (#16792) * grammar : support array references in json schema * Update json-schema-to-grammar.cpp Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> * grammar : improve regex when naming ref derived rules * grammar : replace non-conformant definitions array with anyOf test case --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> # Conflicts: # tests/test-json-schema-to-grammar.cpp * merge fix * llama : minor grammar refactor (#10897) * llama: fix error on bad grammar (#12628) * grammar : fix integer overflow (#17381) * Fix DoS / integer overflow * Remove optional, use INT64_MAX instead as placeholder value (it's technically -1, so it fits :) * White space * Actually, since it's unsigned, use UINT64_MAX # Conflicts: # src/llama-grammar.cpp * grammar: fix regression caused by #17381 (#17412) * grammar: fix regression caused by #17381 * more readable # Conflicts: # src/llama-grammar.cpp * Merge Fix * Fix warnings --------- Signed-off-by: David Korczynski <david@adalogics.com> Co-authored-by: Joe Eli McIlvain <joe.eli.mac@gmail.com> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> Co-authored-by: frob <rick+github@frob.com.au> Co-authored-by: Richard Lyons <frob@cloudstaff.com> Co-authored-by: DavidKorczynski <david@adalogics.com> Co-authored-by: Daniel Bevenius <daniel.bevenius@gmail.com> Co-authored-by: firecoperana <firecoperana> Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com> Co-authored-by: Aldehir Rojas <hello@alde.dev> Co-authored-by: Olivier Chafik <olivier.chafik@gmail.com> Co-authored-by: Piotr Wilkin (ilintar) <piotr.wilkin@syndatis.com> Co-authored-by: Xuan-Son Nguyen <son@huggingface.co> Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>		2025-11-30 18:45:38 +01:00
..
baby-llama	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
batched	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
batched-bench	MoE fix for R4 quants (#170 )	2025-01-12 13:19:14 +02:00
batched.swift	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
benchmark	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
convert-llama2c-to-ggml	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
cvector-generator	CUDA: set compute parameters via command line arguments (#910 )	2025-11-07 07:11:23 +02:00
deprecation-warning	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
embedding	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
eval-callback	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
export-lora	Merge vulkan code from mainline up to commit of 6/28/2025 (#563 )	2025-07-02 08:49:42 +02:00
gbnf-validator	Update grammar (#1023 )	2025-11-30 18:45:38 +01:00
gguf	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
gguf-hash	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
gguf-split	gguf-split : update (#444 )	2025-05-23 08:07:42 +03:00
gritlm	llama : allow pooled embeddings on any model (#7477 )	2024-06-21 08:38:22 +03:00
imatrix	Fix imatrix calculation for MLA models (#411 )	2025-05-13 17:53:38 +03:00
infill	Tool calls support from mainline (#723 )	2025-09-01 08:38:49 +03:00
jeopardy	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
llama-bench	Fix llama-bench mla parameter (#1016 )	2025-11-27 09:33:30 +01:00
llama.android	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
llama.swiftui	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
llava	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
lookahead	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
lookup	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
main	Port mdmd from mainline + Qwen2/2.5-VL support (#798 )	2025-09-27 08:45:29 +02:00
main-cmake-pkg	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
mtmd	Update mtmd to improve accuracy of M-RoPE (#993 )	2025-11-29 07:27:15 +01:00
parallel	Tool calls support from mainline (#723 )	2025-09-01 08:38:49 +03:00
passkey	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
perplexity	More informative PPL readout line (#914 )	2025-11-07 16:41:24 +02:00
quantize	Allow quantization of ffn_gate_inp (#896 )	2025-11-05 10:44:32 +02:00
quantize-stats	Disable experimental code that causes issues with MSVC (#707 )	2025-08-19 18:09:49 +03:00
retrieval	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
rpc	Fix cuda init error in rpc (#957 )	2025-11-14 06:59:54 +02:00
save-load-state	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
server	Update grammar (#1023 )	2025-11-30 18:45:38 +01:00
simple	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
speculative	Support --device and --device-draft parameter (#866 )	2025-10-27 18:13:28 +02:00
sweep-bench	sweep-bench: be able to set TG tokens via -n (#897 )	2025-11-04 14:39:30 +02:00
sycl	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
tokenize	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
base-translate.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
chat-13B.bat	Create chat-13B.bat (#592 )	2023-03-29 20:21:09 +03:00
chat-13B.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
chat-persistent.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
chat-vicuna.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
chat.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
CMakeLists.txt	Port mdmd from mainline + Qwen2/2.5-VL support (#798 )	2025-09-27 08:45:29 +02:00
convert_legacy_llama.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
json_schema_pydantic_example.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
json_schema_to_grammar.py	Update grammar (#1023 )	2025-11-30 18:45:38 +01:00
llama.vim	llama.vim : added api key support (#5090 )	2024-01-23 08:51:27 +02:00
llm.vim	llm.vim : stop generation at multiple linebreaks, bind to <F2> (#2879 )	2023-08-30 09:50:55 +03:00
Miku.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
pydantic_models_to_grammar_examples.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
pydantic_models_to_grammar.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
reason-act.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
regex_to_grammar.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
server_embd.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
server-llama2-13B.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
ts-type-to-grammar.sh	JSON schema conversion: ⚡️ faster repetitions, min/maxLength for strings, cap number length (#6555 )	2024-04-12 19:43:38 +01:00