llama.cpp

mirror of https://github.com/ggerganov/llama.cpp.git synced 2024-12-29 07:34:18 +01:00

Author	SHA1	Message	Date
Georgi Gerganov	af99c6fbfc	llama : remove memory_f16 and kv_f16 flags	2023-12-05 18:18:16 +02:00
Georgi Gerganov	4adb1d69d9	cuda : add comment	2023-12-05 18:15:51 +02:00
Georgi Gerganov	dd86df82e6	metal : use mm kernel only for quantum KV cache	2023-12-05 18:14:04 +02:00
slaren	903167a777	llama-bench : support type_k/type_v	2023-12-05 16:32:53 +01:00
Georgi Gerganov	b2acedeb1a	cuda : add F32 -> Q4_0 and F32 -> Q4_1 copy kernels	2023-12-05 16:47:34 +02:00
Georgi Gerganov	e8457c90a0	cuda : wip	2023-12-05 16:29:52 +02:00
Georgi Gerganov	6b58ae9892	metal : add F32 -> Q4_1 copy kernel	2023-12-05 16:09:16 +02:00
Georgi Gerganov	9d69ecc0c9	metal : add F32 -> Q4_0 copy kernel	2023-12-05 16:01:50 +02:00
Georgi Gerganov	7864a2cd9b	llama : fix build ggml-ci	2023-12-05 15:43:25 +02:00
Georgi Gerganov	3ce30e07c9	llama : pass KV cache type through API	2023-12-05 15:40:23 +02:00
Georgi Gerganov	b881f630ca	cuda : use mmv kernel for quantum cache ops	2023-12-04 15:41:20 +02:00
Georgi Gerganov	a1bf6c09f8	cuda : add F32 -> Q8_0 copy kernel ggml-ci	2023-12-04 15:09:43 +02:00
Georgi Gerganov	bcfebf241d	metal : add F32 -> Q8_0 copy kernel	2023-12-04 10:42:10 +02:00
Georgi Gerganov	d04ee928a2	llama : support quantum K cache (wip)	2023-12-03 21:34:50 +02:00
Georgi Gerganov	66aaac9867	llama : update session save/load	2023-12-03 21:10:16 +02:00
Georgi Gerganov	e262947d43	common : add command-line arg to disable KV cache offloading	2023-12-03 20:31:01 +02:00
Georgi Gerganov	c80b8a2bff	llama : remove mirrors, perform Device -> Host when partial offload	2023-12-03 19:46:06 +02:00
Georgi Gerganov	c44bc1ee00	llama : keep the KV related layers on the device	2023-12-03 19:22:47 +02:00
Georgi Gerganov	1fa91a4833	llama : enable offload debug temporarily	2023-12-03 18:36:02 +02:00
Georgi Gerganov	3d3e6bd0e4	llama : offload for rest of the model arches	2023-12-03 17:52:23 +02:00
Georgi Gerganov	f3dbfb9f60	llama : offload K shift tensors	2023-12-03 17:44:18 +02:00
Georgi Gerganov	986b3da76a	llama : offload KV cache per-layer	2023-12-03 17:34:39 +02:00
Georgi Gerganov	c294c78eb7	Merge branch 'master' into per-layer-kv	2023-12-03 16:35:53 +02:00
Georgi Gerganov	fbbc42827b	ggml : reuse ggml_get_n_tasks() in ggml_graph_plan() (#4308 ) * ggml : fix soft max out-of-bounds access ggml-ci * ggml : reuse ggml_get_n_tasks() in ggml_graph_plan() ggml-ci	2023-12-03 15:56:35 +02:00
Georgi Gerganov	adf3de4f69	ggml : fix soft max out-of-bounds access (#4307 ) ggml-ci	2023-12-03 15:56:22 +02:00
Ed Lee	33e171d1e9	server : fix OpenAI API `stop` field to be optional (#4299 ) (cherry picked from commit Mozilla-Ocho/llamafile@e8c92bcb84)	2023-12-03 11:10:43 +02:00
Rickard Edén	6949b50df5	py : add grammar to oai like api (#4294 )	2023-12-03 11:03:25 +02:00
Georgi Gerganov	d7b800b8bc	llama : pad KV cache size (#4280 ) * llama : pad KV cache size to 32 * metal : try to improve batched decoding	2023-12-03 10:58:16 +02:00
Georgi Gerganov	5a7d3125e7	llama : avoid using "optional" keyword (#4283 )	2023-12-01 20:39:12 +02:00
Georgi Gerganov	d5a1cbde60	llama : support optional tensors (#4283 )	2023-12-01 20:35:47 +02:00
Miwa / Ensan	b220222a64	swift : fix token_to_piece implementation (#4278 ) * Fix token_to_piece implementation in Swift * Fix errors	2023-12-01 20:19:45 +02:00
Jared Van Bortel	511f52c334	build : enable libstdc++ assertions for debug builds (#4275 )	2023-12-01 20:18:35 +02:00
CausalLM	03562f3a86	llama : support attention bias on LLaMA architecture (#4283 ) * Support attention_bias on LLaMA architecture QKVO bias, should fix InternLM (https://github.com/ggerganov/llama.cpp/issues/3133) and works for LLaMAfied Qwen models (https://github.com/ggerganov/llama.cpp/pull/3743#issuecomment-1825923608). * check existence of qkvo bias while loading llama models Tested on LLaMA2, CUDA and CPU. * Update llama.cpp	2023-12-01 20:17:06 +02:00
Shijie	37c746d687	llama : add Qwen support (#4281 ) * enable qwen to llama.cpp * llama : do not GPU split bias tensors --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-12-01 20:16:31 +02:00
Georgi Gerganov	880f57973b	llama : fix integer overflow during quantization (#4284 ) happens with multi-threaded quantization of Qwen-72B ggml-ci	2023-12-01 18:42:11 +02:00
Daniel Bevenius	8d6d9f033b	py : add requirements file for convert-hf-to-gguf.py (#4277 ) This commit adds a requirements file for the convert-hf-to-gguf.py script, and also add the torch and transformers packages to it. The motivation for this is that currently running convert-hf-to-gguf.py will produce the following error: ```console $ python3 -m venv venv $ source venv/bin/activate (venv) $ pip install -r requirements.txt Collecting numpy==1.24.4 Collecting sentencepiece==0.1.98 Collecting gguf>=0.1.0 Installing collected packages: sentencepiece, numpy, gguf Successfully installed gguf-0.5.1 numpy-1.24.4 sentencepiece-0.1.98 (venv) $ python convert-hf-to-gguf.py --help Traceback (most recent call last): File "llama.cpp/convert-hf-to-gguf.py", line 16, in <module> import torch ModuleNotFoundError: No module named 'torch' ``` With this commit, and using requirements-hf-to-gguf.txt instead of requirements.txt, the script can be run and shows the help output. Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com>	2023-12-01 11:41:56 +02:00
Georgi Gerganov	ef47ec18da	ggml : add ggml_soft_max_ext (#4256 ) * metal : implement soft_max_ext * cuda : implement soft_max_ext * ggml : implement soft_max_ext (CPU) * batched-bench : print threads ggml-ci * metal : simplify soft_max encoding ggml-ci * cuda : use 512 threads for soft_max instead of 32 * ggml : update soft max cpu * cuda : do warp-based block reduce * cuda : increase max block size to 1024 * cuda : fix warp reduction initialization of shared mem * metal : warp-based reduction for soft max kernel * metal : warp-based reduce for rms_norm * metal : simplify soft max kernel ggml-ci * alloc : fix build with debug	2023-12-01 10:51:24 +02:00
Ziad Ben Hadj-Alouane	1d144112c0	server : add --log-disable to disable logging to file (#4260 ) * * add --log-disable to disable logging to file in the server example * * typo fix	2023-12-01 00:25:49 +02:00
Ziad Ben Hadj-Alouane	f43f09366d	server : add single-client multi-prompt support (#4232 ) * * add multiprompt support * * cleanup * * more cleanup * * remove atomicity of id_gen, and change lock_guard to unique_lock on completion requests * * remove all references to mutex_multitasks * Update examples/server/server.cpp Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com> * Update examples/server/server.cpp Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com> * Update examples/server/server.cpp Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com> * Update examples/server/server.cpp Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com> * * change to set --------- Co-authored-by: Jared Van Bortel <cebtenzzre@gmail.com>	2023-12-01 00:25:04 +02:00
WillCorticesAI	d2809a3ba2	make : fix Apple clang determination bug (#4272 ) Co-authored-by: Will Findley <findley@gmail.com>	2023-12-01 00:23:44 +02:00
Jared Van Bortel	15f5d96037	build : fix build info generation and cleanup Makefile (#3920 ) * cmake : fix joining of REAL_GIT_DIR * fix includes with help from include-what-you-use * make : remove unneeded deps and add test-rope target * fix C includes in C++ source files * Revert "fix includes with help from include-what-you-use" This reverts commit `635e9fadfd`.	2023-12-01 00:23:08 +02:00
John	33c9892af5	llava : ShareGPT4V compatibility (vision encoder only loading) (#4172 ) * ShareGPT4 compatibility (vision encoder only loading) Load only a CLIP vision encoder (as supplied by ShareGPT finetunes) Corrects the argument parsing for --img_mean and --img_std (which were previously not parsed but attempted to access) Defines defaults for img_mean and img_std which are equal to the llava 1.5 CLIP encoder, so you do not have to provide them * Update convert-image-encoder-to-gguf.py	2023-11-30 23:11:14 +01:00
Andrew Godfrey	8efa0f6ebe	main : pass LOG_TEE callback to llama.cpp log (#4033 ) * main : Call llama_log_set to use LOG_TEE * tabs to spaces	2023-11-30 23:56:19 +02:00
vodkaslime	524907aa76	readme : fix (#4135 ) * fix: readme * chore: resolve comments * chore: resolve comments	2023-11-30 23:49:21 +02:00
Juraj Bednar	3bd2c7ce1b	docker : add finetune option (#4211 )	2023-11-30 23:46:01 +02:00
Miwa / Ensan	bde629bb53	batched.swift : update README.md (#4214 ) docs: update how to run	2023-11-30 23:45:17 +02:00
Li Tan	f7f9e06212	cmake : fix the metal file foder path (#4217 )	2023-11-30 23:44:11 +02:00
Dawid Wysocki	74daabae69	readme : fix typo (#4253 ) llama.cpp uses GitHub Actions, not Gitlab Actions.	2023-11-30 23:43:32 +02:00
Daniel Bevenius	b18c66ca6e	llama : fix alignment of general.name in print meta (#4254 ) * llama: fix alignment of general.name in print meta This commit fixes the alignment of the general.name field in the llm_load_print_meta function. Currently the output looks like this: ```console llm_load_print_meta: model ftype = mostly Q4_0 llm_load_print_meta: model params = 13.02 B llm_load_print_meta: model size = 6.86 GiB (4.53 BPW) llm_load_print_meta: general.name = LLaMA v2 ``` And with this commit it looks like this: ```console llm_load_print_meta: model ftype = mostly Q4_0 llm_load_print_meta: model params = 13.02 B llm_load_print_meta: model size = 6.86 GiB (4.53 BPW) llm_load_print_meta: general.name = LLaMA v2 ``` Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com> * llama: fix alignment of special tokens Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com> --------- Signed-off-by: Daniel Bevenius <daniel.bevenius@gmail.com>	2023-11-30 23:43:08 +02:00
slaren	f4d973cecb	convert.py : fix llama/llama2 conversion due to vocab_size=-1 (#4258 )	2023-11-30 23:42:23 +02:00

1 2 3 4 5 ...

1632 Commits