llama.cpp/fattn-wmma-f16-instance-kqhalf-cpb8.cu at 2decf57bc6e4a6b45176c3727d964a01161beecc - llama.cpp - Gitea: Git with a cup of tea

Mirrors/llama.cpp

mirror of https://github.com/ggerganov/llama.cpp.git synced 2024-11-01 07:30:17 +01:00

Johannes Gäßler 7d1a378b8f

CUDA: refactor mmq, dmmv, mmvq (#7716 )

* CUDA: refactor mmq, dmmv, mmvq

* fix out-of-bounds write

* struct for qk, qr, qi

* fix cmake build

* mmq_type_traits

2024-06-05 16:53:00 +02:00

9 lines

276 B

Plaintext

Raw Blame History

 // This file has been autogenerated by generate_cu_files.py, do not edit manually.
 #include "../fattn-wmma-f16.cuh"
 DECL_FATTN_WMMA_F16_CASE(64, 8, half);
 DECL_FATTN_WMMA_F16_CASE(96, 8, half);
 DECL_FATTN_WMMA_F16_CASE(128, 8, half);
 DECL_FATTN_WMMA_F16_CASE(256, 8, half);