llama.cpp

mirror of https://github.com/ggerganov/llama.cpp.git synced 2025-01-27 12:33:06 +01:00

History

Georgi Gerganov 55e47786e3 llama : default sampling changes + greedy update (#9897 ) * llama : deprecate softmax sampler + fix dist sampler ggml-ci * tests : replace macros with functions ggml-ci * sampling : change temperature sampler logic For t <= 0.0f, keep the max logit intact and set the rest to -inf * cont : no need for special "greedy" logic top-k == 1 is the same * tests : init prob correctly * llama : handle temp <= 0.0 in the temp_ext sampler too ggml-ci * cont : avoid extra loop in temperature sampler for sub-zero temp ggml-ci		2024-10-21 09:46:40 +03:00
..
CMakeLists.txt	llama : move vocab, grammar and sampling into separate files (#8508 )	2024-07-23 13:10:17 +03:00
llama-grammar.cpp	llama : refactor sampling v2 (#9294 )	2024-09-07 15:16:19 +03:00
llama-grammar.h	llama : refactor sampling v2 (#9294 )	2024-09-07 15:16:19 +03:00
llama-impl.h	log : add CONT level for continuing previous log entry (#9610 )	2024-09-24 10:15:35 +03:00
llama-sampling.cpp	llama : default sampling changes + greedy update (#9897 )	2024-10-21 09:46:40 +03:00
llama-sampling.h	llama : add infill sampler (#9896 )	2024-10-15 16:35:33 +03:00
llama-vocab.cpp	llama : infill sampling handle very long tokens (#9924 )	2024-10-17 22:32:47 +03:00
llama-vocab.h	llama : add infill sampler (#9896 )	2024-10-15 16:35:33 +03:00
llama.cpp	llama : remove all_pos_0, all_pos_1, all_seq_id from llama_batch (#9745 )	2024-10-18 23:18:01 +02:00
unicode-data.cpp	server : better security control for public deployments (#9776 )	2024-10-08 13:27:04 +02:00
unicode-data.h	llama : reduce compile time and binary size (#9712 )	2024-10-02 15:49:55 +02:00
unicode.cpp	llama : reduce compile time and binary size (#9712 )	2024-10-02 15:49:55 +02:00
unicode.h	llama : move vocab, grammar and sampling into separate files (#8508 )	2024-07-23 13:10:17 +03:00