llama.cpp/examples/server/tests/features/server.feature

@llama.cpp
Feature: llama.cpp server

  Background: Server startup
    Given a server listening on localhost:8080
    And   a model file stories260K.gguf
    And   a model alias tinyllama-2
    And   42 as server seed
      # KV Cache corresponds to the total amount of tokens
      # that can be stored across all independent sequences: #4130
      # see --ctx-size and #5568
    And   32 KV cache size
    And   1 slots
    And   embeddings extraction
    And   32 server max tokens to predict
    And   prometheus compatible metrics exposed
    Then  the server is starting
    Then  the server is healthy

  Scenario: Health
    Then the server is ready
    And  all slots are idle

  Scenario Outline: Completion
    Given a prompt <prompt>
    And   <n_predict> max tokens to predict
    And   a completion request with no api error
    Then  <n_predicted> tokens are predicted matching <re_content>
    And   prometheus metrics are exposed

    Examples: Prompts
      | prompt                           | n_predict | re_content                             | n_predicted |
      | I believe the meaning of life is | 8         | (read<or>going)+                       | 8           |
      | Write a joke about AI            | 64        | (park<or>friends<or>scared<or>always)+ | 32          |

  Scenario Outline: OAI Compatibility
    Given a model <model>
    And   a system prompt <system_prompt>
    And   a user prompt <user_prompt>
    And   <max_tokens> max tokens to predict
    And   streaming is <enable_streaming>
    Given an OAI compatible chat completions request with no api error
    Then  <n_predicted> tokens are predicted matching <re_content>

    Examples: Prompts
      | model        | system_prompt               | user_prompt                          | max_tokens | re_content                 | n_predicted | enable_streaming |
      | llama-2      | Book                        | What is the best book                | 8          | (Mom<or>what)+             | 8           | disabled         |
      | codellama70b | You are a coding assistant. | Write the fibonacci function in c++. | 64         | (thanks<or>happy<or>bird)+ | 32          | enabled          |

  Scenario: Embedding
    When embeddings are computed for:
    """
    What is the capital of Bulgaria ?
    """
    Then embeddings are generated

  Scenario: OAI Embeddings compatibility
    Given a model tinyllama-2
    When an OAI compatible embeddings computation request for:
    """
    What is the capital of Spain ?
    """
    Then embeddings are generated

  Scenario: OAI Embeddings compatibility with multiple inputs
    Given a model tinyllama-2
    Given a prompt:
      """
      In which country Paris is located ?
      """
    And a prompt:
      """
      Is Madrid the capital of Spain ?
      """
    When an OAI compatible embeddings computation request for multiple inputs
    Then embeddings are generated


  Scenario: Tokenize / Detokenize
    When tokenizing:
    """
    What is the capital of France ?
    """
    Then tokens can be detokenize
server: init functional tests (#5566) * server: tests: init scenarios - health and slots endpoints - completion endpoint - OAI compatible chat completion requests w/ and without streaming - completion multi users scenario - multi users scenario on OAI compatible endpoint with streaming - multi users with total number of tokens to predict exceeds the KV Cache size - server wrong usage scenario, like in Infinite loop of "context shift" #3969 - slots shifting - continuous batching - embeddings endpoint - multi users embedding endpoint: Segmentation fault #5655 - OpenAI-compatible embeddings API - tokenize endpoint - CORS and api key scenario * server: CI GitHub workflow --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> 2024-02-24 12:28:55 +01:00			`@llama.cpp`
			`Feature: llama.cpp server`

			`Background: Server startup`
			`Given a server listening on localhost:8080`
			`And a model file stories260K.gguf`
			`And a model alias tinyllama-2`
			`And 42 as server seed`
			`# KV Cache corresponds to the total amount of tokens`
			`# that can be stored across all independent sequences: #4130`
			`# see --ctx-size and #5568`
			`And 32 KV cache size`
			`And 1 slots`
			`And embeddings extraction`
			`And 32 server max tokens to predict`
server: concurrency fix + monitoring - add /metrics prometheus compatible endpoint (#5708) * server: monitoring - add /metrics prometheus compatible endpoint * server: concurrency issue, when 2 task are waiting for results, only one call thread is notified * server: metrics - move to a dedicated struct 2024-02-25 13:49:43 +01:00			`And prometheus compatible metrics exposed`
server: init functional tests (#5566) * server: tests: init scenarios - health and slots endpoints - completion endpoint - OAI compatible chat completion requests w/ and without streaming - completion multi users scenario - multi users scenario on OAI compatible endpoint with streaming - multi users with total number of tokens to predict exceeds the KV Cache size - server wrong usage scenario, like in Infinite loop of "context shift" #3969 - slots shifting - continuous batching - embeddings endpoint - multi users embedding endpoint: Segmentation fault #5655 - OpenAI-compatible embeddings API - tokenize endpoint - CORS and api key scenario * server: CI GitHub workflow --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> 2024-02-24 12:28:55 +01:00			`Then the server is starting`
			`Then the server is healthy`

			`Scenario: Health`
			`Then the server is ready`
			`And all slots are idle`

			`Scenario Outline: Completion`
			`Given a prompt <prompt>`
			`And <n_predict> max tokens to predict`
			`And a completion request with no api error`
			`Then <n_predicted> tokens are predicted matching <re_content>`
server: concurrency fix + monitoring - add /metrics prometheus compatible endpoint (#5708) * server: monitoring - add /metrics prometheus compatible endpoint * server: concurrency issue, when 2 task are waiting for results, only one call thread is notified * server: metrics - move to a dedicated struct 2024-02-25 13:49:43 +01:00			`And prometheus metrics are exposed`
server: init functional tests (#5566) * server: tests: init scenarios - health and slots endpoints - completion endpoint - OAI compatible chat completion requests w/ and without streaming - completion multi users scenario - multi users scenario on OAI compatible endpoint with streaming - multi users with total number of tokens to predict exceeds the KV Cache size - server wrong usage scenario, like in Infinite loop of "context shift" #3969 - slots shifting - continuous batching - embeddings endpoint - multi users embedding endpoint: Segmentation fault #5655 - OpenAI-compatible embeddings API - tokenize endpoint - CORS and api key scenario * server: CI GitHub workflow --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> 2024-02-24 12:28:55 +01:00
			`Examples: Prompts`
server: logs - unified format and --log-format option (#5700) * server: logs - always use JSON logger, add add thread_id in message, log task_id and slot_id * server : skip GH copilot requests from logging * server : change message format of server_log() * server : no need to repeat log in comment * server : log style consistency * server : fix compile warning * server : fix tests regex patterns on M2 Ultra * server: logs: PR feedback on log level * server: logs: allow to choose log format in json or plain text * server: tests: output server logs in text * server: logs switch init logs to server logs macro * server: logs ensure value json value does not raised error * server: logs reduce level VERBOSE to VERB to max 4 chars * server: logs lower case as other log messages * server: logs avoid static in general Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> * server: logs PR feedback: change text log format to: LEVEL [function_name] message \| additional=data --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> 2024-02-25 13:50:32 +01:00			`\| prompt \| n_predict \| re_content \| n_predicted \|`
			`\| I believe the meaning of life is \| 8 \| (read<or>going)+ \| 8 \|`
			`\| Write a joke about AI \| 64 \| (park<or>friends<or>scared<or>always)+ \| 32 \|`
server: init functional tests (#5566) * server: tests: init scenarios - health and slots endpoints - completion endpoint - OAI compatible chat completion requests w/ and without streaming - completion multi users scenario - multi users scenario on OAI compatible endpoint with streaming - multi users with total number of tokens to predict exceeds the KV Cache size - server wrong usage scenario, like in Infinite loop of "context shift" #3969 - slots shifting - continuous batching - embeddings endpoint - multi users embedding endpoint: Segmentation fault #5655 - OpenAI-compatible embeddings API - tokenize endpoint - CORS and api key scenario * server: CI GitHub workflow --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> 2024-02-24 12:28:55 +01:00
			`Scenario Outline: OAI Compatibility`
			`Given a model <model>`
			`And a system prompt <system_prompt>`
			`And a user prompt <user_prompt>`
			`And <max_tokens> max tokens to predict`
			`And streaming is <enable_streaming>`
			`Given an OAI compatible chat completions request with no api error`
			`Then <n_predicted> tokens are predicted matching <re_content>`

			`Examples: Prompts`
			`\| model \| system_prompt \| user_prompt \| max_tokens \| re_content \| n_predicted \| enable_streaming \|`
			`\| llama-2 \| Book \| What is the best book \| 8 \| (Mom<or>what)+ \| 8 \| disabled \|`
			`\| codellama70b \| You are a coding assistant. \| Write the fibonacci function in c++. \| 64 \| (thanks<or>happy<or>bird)+ \| 32 \| enabled \|`

			`Scenario: Embedding`
			`When embeddings are computed for:`
			`"""`
			`What is the capital of Bulgaria ?`
			`"""`
			`Then embeddings are generated`

			`Scenario: OAI Embeddings compatibility`
			`Given a model tinyllama-2`
			`When an OAI compatible embeddings computation request for:`
			`"""`
			`What is the capital of Spain ?`
			`"""`
			`Then embeddings are generated`

server: continue to update other slots on embedding concurrent request (#5699) * server: #5655 - continue to update other slots on embedding concurrent request. * server: tests: add multi users embeddings as fixed * server: tests: adding OAI compatible embedding concurrent endpoint * server: tests: adding OAI compatible embedding with multiple inputs 2024-02-24 19:16:04 +01:00			`Scenario: OAI Embeddings compatibility with multiple inputs`
			`Given a model tinyllama-2`
			`Given a prompt:`
			`"""`
			`In which country Paris is located ?`
			`"""`
			`And a prompt:`
			`"""`
			`Is Madrid the capital of Spain ?`
			`"""`
			`When an OAI compatible embeddings computation request for multiple inputs`
			`Then embeddings are generated`

server: init functional tests (#5566) * server: tests: init scenarios - health and slots endpoints - completion endpoint - OAI compatible chat completion requests w/ and without streaming - completion multi users scenario - multi users scenario on OAI compatible endpoint with streaming - multi users with total number of tokens to predict exceeds the KV Cache size - server wrong usage scenario, like in Infinite loop of "context shift" #3969 - slots shifting - continuous batching - embeddings endpoint - multi users embedding endpoint: Segmentation fault #5655 - OpenAI-compatible embeddings API - tokenize endpoint - CORS and api key scenario * server: CI GitHub workflow --------- Co-authored-by: Georgi Gerganov <ggerganov@gmail.com> 2024-02-24 12:28:55 +01:00
			`Scenario: Tokenize / Detokenize`
			`When tokenizing:`
			`"""`
			`What is the capital of France ?`
			`"""`
			`Then tokens can be detokenize`