ik_llama.cpp

History

Anton Sokolchenko 3701fb1686 Function calling support for Kimi-K2 (#628 ) * Implement function calling / tools for ik_llama.cpp for Kimi K2 * Implement basic tool choice * Backport llama.cpp tool calls support * Enhance function calls with improved chat parser and string utilities - Add new chat.h/chat.cpp and chat-parser.h/chat-parser.cpp for better chat handling - Improve function calls parsing with fallback to llama.cpp builder pattern - Add string utility functions (starts_with, ends_with, find_partial_stop) - Update README with function calls testing instructions - Enhance Kimi K2 parser and function calls documentation - Add comprehensive test suite for function calls - Update CMakeLists.txt and Makefile for new components * Enhance function calling with unified streaming and parser improvements - Fix streaming content cleanup to prevent function syntax in output - Unify content extraction patterns with llama.cpp approach - Improve Kimi K2 parser robustness and partial content handling - Add comprehensive test coverage for function call scenarios - Optimize chat message parsing and diff computation * Replace hardcoded values in kimi_k2_parser.hpp with named constants - Add compile-time constants for all token format markers - Add compile-time constants for XML format markers - Add compile-time constants for simple format patterns - Replace all hardcoded string literals with named constants - Use compile-time length calculation to avoid manual counting - Improve maintainability and reduce magic numbers throughout parser * Fix duplicate common_chat_parse definition - Remove duplicate implementation from chat-parser.cpp - Keep single implementation in chat.cpp following llama.cpp patterns - Resolves linker error: multiple definition of common_chat_parse * Fix JSON assertion failure in function call parsing - Add proper validation that 'function' field is an object before accessing nested keys - Handle missing 'arguments' field gracefully with default "{}" - Prevents crash when parsing malformed tool call JSON structures * Add comprehensive Qwen3 XML tool calling support with unit tests - Implement Qwen3 XML parser with <tool_call>{"name": "func", "arguments": {...}}</tool_call> format - Add model detection and routing for Qwen3 vs Kimi-K2 formats - Create 8 comprehensive unit tests covering parsing, streaming, error handling - Fix token format cleaning bug in kimi_k2_parser.hpp processing order - Remove progressive parsing code and related utilities - Add tool injection support for Qwen3 format in server utils * Add DeepSeek R1 function calling support with comprehensive unit tests - Implement complete DeepSeek R1 tool call parsing in common_chat_parser.cpp - Add DeepSeek R1 model detection and tool injection in deepseek_r1_tools.hpp - Update function_calls.hpp with DeepSeek R1 integration and content extraction - Update documentation to reflect support for Kimi-K2, Qwen3, and DeepSeek R1 models - Add comprehensive unit tests for DeepSeek R1 reasoning, tool calls, and integration - Port exact implementation patterns from original llama.cpp for compatibility Key features: - Native DeepSeek R1 format: <｜tool▁calls▁begin｜>function<｜tool▁sep｜>name```json{}```<｜tool▁call▁end｜><｜tool▁calls▁end｜> - Reasoning content extraction from <think>...</think> tags - Multiple tool calls support with separate call blocks - Model detection for deepseek-r1, deepseek_r1 naming patterns - Integration with incremental parsing and streaming support * Add partial parsing support for JSON and regex - json-partial.h/cpp: JSON partial parsing functionality - regex-partial.h/cpp: Regex partial parsing functionality * Add format_chat integration tests for Qwen3 tool injection - Add test_qwen3_format_chat_integration() to validate tool injection pipeline - Test tool injection conditions and system message enhancement - Verify JSON formatting and anti-preamble instructions - Add comprehensive test documentation Tests confirm tool injection works correctly - conversational preamble issue is not in ik_llama.cpp but likely in UI configuration. * Fix Qwen3 tool call parsing - pass model name to parser Server was not passing model name to parse_chat_message_incremental(), causing Qwen3 to fall back to Kimi-K2 parser and return tool calls as content instead of proper tool_calls array. * Fix non-streaming path to use model-specific parsing Non-streaming responses were hardcoded to use Kimi-K2 format, causing Qwen3 XML tool calls to be returned as content instead of proper tool_calls array. Now uses same model detection as streaming path for consistency.		2025-07-23 18:11:42 +02:00
..
baby-llama	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
batched	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
batched-bench	MoE fix for R4 quants (#170 )	2025-01-12 13:19:14 +02:00
batched.swift	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
benchmark	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
convert-llama2c-to-ggml	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
cvector-generator	Merge vulkan code from mainline up to commit of 6/28/2025 (#563 )	2025-07-02 08:49:42 +02:00
deprecation-warning	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
embedding	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
eval-callback	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
export-lora	Merge vulkan code from mainline up to commit of 6/28/2025 (#563 )	2025-07-02 08:49:42 +02:00
gbnf-validator	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
gguf	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
gguf-hash	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
gguf-split	gguf-split : update (#444 )	2025-05-23 08:07:42 +03:00
gritlm	llama : allow pooled embeddings on any model (#7477 )	2024-06-21 08:38:22 +03:00
imatrix	Fix imatrix calculation for MLA models (#411 )	2025-05-13 17:53:38 +03:00
infill	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
jeopardy	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
llama-bench	Add copyright notices (#317 )	2025-04-07 10:43:26 +02:00
llama.android	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
llama.swiftui	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
llava	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
lookahead	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
lookup	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
main	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
main-cmake-pkg	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
parallel	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
passkey	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
perplexity	Fix KLD precision (#325 )	2025-04-12 16:17:50 +02:00
quantize	Adding IQ1_KT - 1.75 bpw SOTA quants (#616 )	2025-07-20 10:05:23 +02:00
quantize-stats	Adding IQ2_KL (#602 )	2025-07-14 18:55:08 +02:00
retrieval	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
rpc	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
save-load-state	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
server	Function calling support for Kimi-K2 (#628 )	2025-07-23 18:11:42 +02:00
simple	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
speculative	add dry sampler (#513 )	2025-06-19 10:24:53 +03:00
sweep-bench	Add batch warmup to sweep-bench (#375 )	2025-05-12 07:50:26 +03:00
sycl	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
tokenize	Merge mainline - Aug 12 2024 (#17 )	2024-08-12 15:14:32 +02:00
base-translate.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
chat-13B.bat	Create chat-13B.bat (#592 )	2023-03-29 20:21:09 +03:00
chat-13B.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
chat-persistent.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
chat-vicuna.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
chat.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
CMakeLists.txt	Add new sweep-bench benchmark (#225 )	2025-02-23 00:16:27 -06:00
convert_legacy_llama.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
json_schema_pydantic_example.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
json_schema_to_grammar.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
llama.vim	llama.vim : added api key support (#5090 )	2024-01-23 08:51:27 +02:00
llm.vim	llm.vim : stop generation at multiple linebreaks, bind to <F2> (#2879 )	2023-08-30 09:50:55 +03:00
Miku.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
pydantic_models_to_grammar_examples.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
pydantic_models_to_grammar.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
reason-act.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
regex_to_grammar.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
server_embd.py	Merge mainline llama.cpp (#3 )	2024-07-27 07:55:01 +02:00
server-llama2-13B.sh	`build`: rename main → llama-cli, server → llama-server, llava-cli → llama-llava-cli, etc... (#7809 )	2024-06-13 00:41:52 +01:00
ts-type-to-grammar.sh	JSON schema conversion: ⚡️ faster repetitions, min/maxLength for strings, cap number length (#6555 )	2024-04-12 19:43:38 +01:00