text-generation-inference/server
Tiezhen WANG b07a2518d9 Update the link for qwen2 (#2068)
* Update the link for qwen2

* Fix Qwen2 model URL in model table

* Fix too eager staging

---------

Co-authored-by: Daniël de Kok <me@danieldk.eu>
2024-09-24 03:43:30 +00:00
..
custom_kernels chore: add pre-commit (#1569) 2024-04-24 15:32:02 +03:00
exllama_kernels MI300 compatibility (#1764) 2024-07-17 05:36:58 +00:00
exllamav2_kernels chore: add pre-commit (#1569) 2024-04-24 15:32:02 +03:00
marlin Add support for GPTQ Marlin (#2052) 2024-09-24 03:43:30 +00:00
tests server: use chunked inputs 2024-09-24 03:42:29 +00:00
text_generation_server Update the link for qwen2 (#2068) 2024-09-24 03:43:30 +00:00
.gitignore Impl simple mamba model (#1480) 2024-04-23 11:45:11 +03:00
Makefile Add support for GPTQ Marlin (#2052) 2024-09-24 03:43:30 +00:00
Makefile-awq chore: add pre-commit (#1569) 2024-04-24 15:32:02 +03:00
Makefile-eetq Upgrade EETQ (Fixes the cuda graphs). (#1729) 2024-04-25 17:58:27 +03:00
Makefile-flash-att Hotfixing make install. (#2008) 2024-09-24 03:29:29 +00:00
Makefile-flash-att-v2 Hotfixing make install. (#2008) 2024-09-24 03:29:29 +00:00
Makefile-selective-scan chore: add pre-commit (#1569) 2024-04-24 15:32:02 +03:00
Makefile-vllm Update LLMM1 bound (#2050) 2024-09-24 03:42:29 +00:00
poetry.lock Making make install work better by default. (#2004) 2024-09-24 03:29:29 +00:00
pyproject.toml Making make install work better by default. (#2004) 2024-09-24 03:29:29 +00:00
README.md chore: add pre-commit (#1569) 2024-04-24 15:32:02 +03:00
requirements_cuda.txt Modifing the version number. 2024-07-17 05:36:58 +00:00
requirements_intel.txt reable xpu, broken by gptq and setuptool upgrade (#1988) 2024-09-24 03:26:17 +00:00
requirements_rocm.txt Modifing the version number. 2024-07-17 05:36:58 +00:00

Text Generation Inference Python gRPC Server

A Python gRPC server for Text Generation Inference

Install

make install

Run

make run-dev