llama.cpp

mirror of https://github.com/ggerganov/llama.cpp synced 2026-03-23 07:30:53 +01:00

History

leo-pony 6b8447352d [CANN] Adapt to dynamically loadable backends mechanism (#9970 ) * [CANN] Adapt to dynamically loadable backends mechanism * Fix the Bug: inference running result is garbled in debug running model for LM models who's type is Q4_0 class * Handle the review comments of this pull request		2024-10-22 16:16:01 +08:00
..
CMakeLists.txt	llama : move vocab, grammar and sampling into separate files (#8508 )	2024-07-23 13:10:17 +03:00
llama-grammar.cpp	llama : refactor sampling v2 (#9294 )	2024-09-07 15:16:19 +03:00
llama-grammar.h	llama : refactor sampling v2 (#9294 )	2024-09-07 15:16:19 +03:00
llama-impl.h	log : add CONT level for continuing previous log entry (#9610 )	2024-09-24 10:15:35 +03:00
llama-sampling.cpp	llama : default sampling changes + greedy update (#9897 )	2024-10-21 09:46:40 +03:00
llama-sampling.h	llama : add infill sampler (#9896 )	2024-10-15 16:35:33 +03:00
llama-vocab.cpp	llama : infill sampling handle very long tokens (#9924 )	2024-10-17 22:32:47 +03:00
llama-vocab.h	llama : add infill sampler (#9896 )	2024-10-15 16:35:33 +03:00
llama.cpp	[CANN] Adapt to dynamically loadable backends mechanism (#9970 )	2024-10-22 16:16:01 +08:00
unicode-data.cpp	server : better security control for public deployments (#9776 )	2024-10-08 13:27:04 +02:00
unicode-data.h	llama : reduce compile time and binary size (#9712 )	2024-10-02 15:49:55 +02:00
unicode.cpp	llama : reduce compile time and binary size (#9712 )	2024-10-02 15:49:55 +02:00
unicode.h	llama : move vocab, grammar and sampling into separate files (#8508 )	2024-07-23 13:10:17 +03:00