llama.cpp

mirror of https://github.com/ggerganov/llama.cpp.git synced 2024-12-27 06:39:25 +01:00

Author	SHA1	Message	Date
Matvey Soloviev	89d5d90f3b	Fix color codes emitting mid-UTF8 code. (#312 )	2023-03-21 19:11:01 +02:00
comex	16ffc013c6	Importer for GPTQ quantized LLaMA models (#301 ) * [WIP, broken] Importer for GPTQ quantized LLaMA models Based on: https://github.com/qwopqwop200/GPTQ-for-LLaMa Current status: Something is busted. The output starts out decent, but quickly degrades into gibberish. This doesn't happen with either the original GPTQ-for-LLaMa using the same weights, or llama.cpp when using weights quantized by its own quantizer. Is there a bug in the conversion script that somehow only comes into play with a large context size? I did notice one potential issue. It's clearly not the main cause of the gibberish, since it doesn't happen when using q4_1 weights quantized by llama.cpp itself, but it seems concerning. When doing a matrix multiplication of f16 * f32 => f32 or q4_1 * f32 => f32, at least when the multiplication is not done with BLAS, the intermediate results are stored in the smaller format rather than f32. This seems like an unnecessary waste of precision, especially in the q4_1 case. I was originally hoping to validate the results by matching the Python implementation's output exactly, but precision and non-associativity issues make this very difficult, including when performing matrix multiplications and, especially, computing norms. Anyway, design details: The models being imported store per-layer weights in essentially q4_1 format, although the addend and scale are shared across an entire row rather than every group of 32 weights. This script duplicates the addend and scale to match ggml's expectations, at the cost of wasting some memory. However, there are two differences which I accommodated changing the output format (and adding corresponding support to main.cpp) rather than having the script match the existing one: - The tok_embeddings and output weights (i.e. the weights that aren't per-layer) are f16 instead of q4_1. They could be converted to q4_1, and the impact of the loss of precision would probably be low, but this would rule out exactly matching the Python implementation's output for validation. - There is no sharding, since the input doesn't have it, and for a CPU-only implementation it seems more useful to avoid having to deal with multiple files. The new format is differentiated from existing q4_1 format by changing the 'f16' header flag to a new value, 4. That said, I think a cleaner approach would be to change main.cpp to support loading each tensor with an arbitrary sharding configuration and type rather than hardcoding specific combinations of types. So far I've wasted too much time debugging to try implementing this... * Add missing permutation. Now it works. --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-21 18:42:25 +02:00
Gary Linscott	486ae645fd	Compute perplexity over prompt (#270 ) * Compute perplexity over prompt * More accurate perplexity calculation - over all logits in the context window (so 512x more tokens!) * Output all perplexitiies * Add timing/ETA	2023-03-21 18:27:42 +02:00
Jean-Christophe Hoelt	3ab3e6582f	Add chatLLaMa script (#198 ) * Add chatLLaMa script * Fix shellcheck errors and do some cleanup * Move chatLLaMa script to `examples` directory * Reduce chatLLaMa context size to 2048 Ref `d7def1a752` * Include n_predict to 2048 in examples/chatLLaMa	2023-03-21 18:23:15 +02:00
Alex von Gluck IV	f157088cb7	makefile: Fix CPU feature detection on Haiku (#218 )	2023-03-21 18:21:06 +02:00
anzz1	c86ba036e6	Enable ANSI colors on Windows 10+ (#311 ) * Enable ANSI colors on Windows 10+ On older versions function will silently fail without any ill effects * Do not call SetConsoleMode if the mode is already set * Update main.cpp --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-21 18:14:46 +02:00
Georgi Gerganov	1daf4dd712	Minor style changes	2023-03-21 18:10:32 +02:00
Georgi Gerganov	dc6a845b85	Add chat.sh script	2023-03-21 18:09:46 +02:00
tjohnman	6a612959e1	Check for reverse prompt by characters instead of tokens (#292 ) (#330 ) * Check for reverse prompt by characters instead of tokens (#292) * Update main.cpp Wording. * Cleanup. * Remove unnecessary use of std::stringstream. --------- Co-authored-by: Johnman <tjohnman@github> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-21 18:05:06 +02:00
tjohnman	d5f56a5e5a	Check for reverse prompt by characters instead of tokens (#292 ) (#330 ) * Check for reverse prompt by characters instead of tokens (#292) * Update main.cpp Wording. * Cleanup. * Remove unnecessary use of std::stringstream. --------- Co-authored-by: Johnman <tjohnman@github> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-21 18:04:43 +02:00
Georgi Gerganov	3bfa3b43b7	Fix convert script, warnings alpaca instructions, default params	2023-03-21 17:59:16 +02:00
Kevin Lo	715d292ee0	Add OpenBSD support (#314 )	2023-03-21 17:50:09 +02:00
Mack Straight	c98ae02668	fix typo in comment (#318 )	2023-03-21 17:49:43 +02:00
Qingyou Meng	c3b2306b18	Makefile: slightly cleanup for Mac Intel; echo instead of run ./main -h (#335 )	2023-03-21 17:44:11 +02:00
anzz1	975d2cebf9	cmdline option for custom amount of model parts (--n_parts N) (#348 ) * cmdline option for custom amount of model parts (--n_parts N) * Update main.cpp --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-21 17:42:43 +02:00
Kevin Kwok	e0ffc861fa	Update IPFS links to quantized alpaca with new tokenizer format (#352 )	2023-03-21 17:34:49 +02:00
Georgi Gerganov	8f644a0a85	Change default repeat_penalty to 1.0 I feel this penalty is not really helping. Especially for the example from the README it makes results pretty bad	2023-03-21 17:32:14 +02:00
Georgi Gerganov	eb34620aec	Add tokenizer test + revert to C++11 (#355 ) * Add test-tokenizer-0 to do a few tokenizations - feel free to expand * Added option to convert-pth-to-ggml.py script to dump just the vocabulary * Added ./models/ggml-vocab.bin containing just LLaMA vocab data (used for tests) * Added utility to load vocabulary file from previous point (temporary implementation) * Avoid using std::string_view and drop back to C++11 (hope I didn't break something) * Rename gpt_vocab -> llama_vocab * All CMake binaries go into ./bin/ now	2023-03-21 17:29:41 +02:00
Casey Primozic	2e664f1ff4	Add initial AVX512 support for dot product on Linux (#320 ) * Update Makefile to detect AVX512 support and add compiler flags if it's available * Based on existing AVX2 implementation, dot product on one 32-value block of 4-bit quantized ints at a time * Perform 8 bit -> 16 bit sign extension and multiply+add on 32 values at time instead of 16 * Use built-in AVX512 horizontal reduce add to get sum at the end * Manual unrolling on inner dot product loop to reduce loop counter overhead	2023-03-21 15:35:42 +01:00
nusu-github	8cf9f34edd	Adding missing features of CMakeLists.txt & Refactoring (#131 ) * Functionality addition CMakeLists.txt Refactoring: 1. Simplify more options that are negation of negation. LLAMA_NO_ACCELERATE -> LLAMA_ACCELERATE 2. Changed to an optional expression instead of forcing to enable AVX2 in MSVC. 3. Make CMAKE_CXX_STANDARD, which is different from Makefile, the same. 4. Use add_compile_options instead of adding options to CMAKE_C_FLAGS. 5. Make utils use target_link_libraries instead of directly referencing code. Added features: 1. Added some options. LLAMA_STATIC_LINK,LLAMA_NATIVE,LLAMA_LTO,LLAMA_GPROF,LLAMA_OPENBLAS * Fix Accelerate link in CMake * Windows build Fix * C++11 to C++17 * Reflects C/C++ standard individually * Change the version to 3.12 --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-21 01:37:16 +01:00
Ben Siraphob	bd4b46d6ba	Nix flake: set meta.mainProgram to llama	2023-03-20 22:50:22 +01:00
Qingyou Meng	6b6d5b5024	Fixed tokenizer.model not found error when model dir is symlink (#325 )	2023-03-20 19:33:10 +00:00
Mack Straight	a791a68b61	move file magic/version to header, print expected version (#319 )	2023-03-20 19:26:01 +00:00
Bernat Vadell	0f1b21cb90	Docker - Fix publish docker image in GitHub Registry (#235 ) * fix publish permission * try to fix docker pipeline using as password github_token & username repository_owner	2023-03-20 18:05:20 +01:00
Mack Straight	074bea2eb1	sentencepiece bpe compatible tokenizer (#252 ) * potential out of bounds read * fix quantize * style * Update convert-pth-to-ggml.py * mild cleanup * don't need the space-prefixing here rn since main.cpp already does it * new file magic + version header field * readme notice * missing newlines Co-authored-by: slaren <2141330+slaren@users.noreply.github.com>	2023-03-20 03:17:23 -07:00
Stephan Walter	5cb63e2493	Add tqdm to Python requirements (#293 ) * Add tqdm to Python requirements * Remove torchvision torchaudio, add requests	2023-03-20 09:24:11 +01:00
cocktailpeanut	da5303c1ea	bugfix: default should not be interactive (#304 )	2023-03-19 23:44:20 +02:00
Georgi Gerganov	4545539d71	Rename script	2023-03-19 21:58:51 +02:00
Georgi Gerganov	edeba28366	Add temporary helper script for Alpaca chat	2023-03-19 21:57:48 +02:00
Rickey Bowers Jr	5c19c70ba6	fix coloring of last `n_batch` of prompt, and refactor line input (#221 ) * fix coloring of last `n_batch` of prompt, and refactor line input * forgot the newline that needs to be sent to the model * (per #283) try to force flush of color reset in SIGINT handler	2023-03-19 19:44:30 +00:00
tjohnman	24568371ae	Support for multiple reverse prompts. (#299 ) Co-authored-by: Johnman <> Co-authored-by: Johnman <tjohnman@github>	2023-03-19 21:33:06 +02:00
Suaj Carrot	7392f1cd2c	Improved quantize script (#222 ) * Improved quantize script I improved the quantize script by adding error handling and allowing to select many models for quantization at once in the command line. I also converted it to Python for generalization as well as extensibility. * Fixes and improvements based on Matt's observations Fixed and improved many things in the script based on the reviews made by @mattsta. The parallelization suggestion is still to be revised, but code for it was still added (commented). * Small fixes to the previous commit * Corrected to use the original glob pattern The original Bash script uses a glob pattern to match files that have endings such as ...bin.0, ...bin.1, etc. That has been translated correctly to Python now. * Added support for Windows and updated README to use this script New code to set the name of the quantize script binary depending on the platform has been added (quantize.exe if working on Windows) and the README.md file has been updated to use this script instead of the Bash one. * Fixed a typo and removed shell=True in the subprocess.run call Fixed a typo regarding the new filenames of the quantized models and removed the shell=True parameter in the subprocess.run call as it was conflicting with the list of parameters. * Corrected previous commit * Small tweak: changed the name of the program in argparse This was making the automatic help message to be suggesting the program's usage as being literally "$ Quantization Script [arguments]". It should now be something like "$ python3 quantize.py [arguments]".	2023-03-19 20:38:44 +02:00
tjohnman	ad5fd5b60c	Make prompt randomization optional. (#300 ) Co-authored-by: Johnman <>	2023-03-19 20:36:19 +02:00
tjohnman	368d0c8a9e	Respect the maximum number of tokens in interactive. (#298 ) Co-authored-by: Johnman <johnman@github> Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-19 20:31:17 +02:00
slaren	50fae10d03	Add --ignore-eos parameter (#181 ) Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-19 20:22:48 +02:00
Qingyou Meng	084e2f0ec0	interactive mode: print '\n' in sigint_handler, this flush stdout thus ensure color reset. (#283 )	2023-03-19 20:10:00 +02:00
Erik Scholz	0b366e7357	Command line switch to use F16 for memory_k and memory_v (refactor of #154 ) (#294 ) * Use F16 for memory_k and memory_v * add command line switch to use f16 instead of f32 for memory k+v --------- Co-authored-by: Ty Everett <ty@tyweb.us>	2023-03-19 19:57:00 +02:00
Georgi Gerganov	160bfb217d	Update hot topics to mention Alpaca support	2023-03-19 19:51:55 +02:00
Georgi Gerganov	c494ed5b94	Fix off-by-one bug (#115 )	2023-03-19 19:46:32 +02:00
Georgi Gerganov	c1c7026b47	Fix python stuff (#109 )	2023-03-19 19:33:18 +02:00
qunash	467b149761	Refactoring `convert-pth-to-ggml.py`: more concise and readable (#109 ) * Refactor get_n_parts function to simplify code and improve readability * Use f-strings instead of concatenation * Refactoring: more concise and readable * modularize --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>	2023-03-19 19:17:39 +02:00
Georgi Gerganov	70f01cb863	Drop trailing new line from file prompts (#80 )	2023-03-19 19:05:04 +02:00
Georgi Gerganov	a4e63b73df	Add instruction for using Alpaca (#240 )	2023-03-19 18:49:50 +02:00
Georgi Gerganov	9e1707218a	Add "--instruct" argument for usage with Alpaca (#240 ) Also start adding prompts in "./prompts"	2023-03-19 18:37:02 +02:00
Georgi Gerganov	22213a17b5	Change RMSNorm eps to 1e-6 (#173 ) I think this is what is used in the Python code	2023-03-19 17:30:00 +02:00
Ronsor	d7def1a752	Warn user if a context size greater than 2048 tokens is specified (#274 ) LLaMA doesn't support more than 2048 token context sizes, and going above that produces terrible results.	2023-03-18 20:10:47 -04:00
Pavol Rusnak	6f61c18ec9	Fix typo in readme	2023-03-18 23:18:04 +01:00
Pavol Rusnak	1e5a6d088d	Add note about Python 3.11 to readme	2023-03-18 22:25:35 +01:00
Pavol Rusnak	554b541521	Add memory/disk requirements to readme	2023-03-18 22:25:35 +01:00
Alex Nguyen	d3f202d57b	Remove unused code since n_vocab is model.hparams.n_vocab (#262 )	2023-03-18 13:51:49 +00:00

... 12 13 14 15 16

786 Commits