text-generation-inference/Dockerfile

# Rust builder
FROM lukemathwalker/cargo-chef:latest-rust-1.80 AS chef
WORKDIR /usr/src

ARG CARGO_REGISTRIES_CRATES_IO_PROTOCOL=sparse

FROM chef AS planner
COPY Cargo.lock Cargo.lock
COPY Cargo.toml Cargo.toml
COPY rust-toolchain.toml rust-toolchain.toml
COPY proto proto
COPY benchmark benchmark
COPY router router
COPY backends backends
COPY launcher launcher

RUN cargo chef prepare --recipe-path recipe.json

FROM chef AS builder

RUN apt-get update && DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends \
    python3.11-dev
RUN PROTOC_ZIP=protoc-21.12-linux-x86_64.zip && \
    curl -OL https://github.com/protocolbuffers/protobuf/releases/download/v21.12/$PROTOC_ZIP && \
    unzip -o $PROTOC_ZIP -d /usr/local bin/protoc && \
    unzip -o $PROTOC_ZIP -d /usr/local 'include/*' && \
    rm -f $PROTOC_ZIP

COPY --from=planner /usr/src/recipe.json recipe.json
RUN cargo chef cook --profile release-opt --recipe-path recipe.json

ARG GIT_SHA
ARG DOCKER_LABEL

COPY Cargo.toml Cargo.toml
COPY rust-toolchain.toml rust-toolchain.toml
COPY proto proto
COPY benchmark benchmark
COPY router router
COPY backends backends
COPY launcher launcher
RUN cargo build --profile release-opt

# Text Generation Inference base image
FROM vault.habana.ai/gaudi-docker/1.17.0/ubuntu22.04/habanalabs/pytorch-installer-2.3.1:latest as base

# Text Generation Inference base env
ENV HF_HOME=/data \
    HF_HUB_ENABLE_HF_TRANSFER=1 \
    PORT=80

# libssl.so.1.1 is not installed on Ubuntu 22.04 by default, install it
RUN wget http://nz2.archive.ubuntu.com/ubuntu/pool/main/o/openssl/libssl1.1_1.1.1f-1ubuntu2_amd64.deb && \
    dpkg -i ./libssl1.1_1.1.1f-1ubuntu2_amd64.deb

WORKDIR /usr/src

RUN apt-get update && DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends \
        libssl-dev \
        ca-certificates \
        make \
        curl \
        git \
        python3.11-dev \
        && rm -rf /var/lib/apt/lists/*

# Install server
COPY proto proto
COPY server server
COPY server/Makefile server/Makefile
RUN cd server && \
    make gen-server && \
    pip install -r requirements.txt && \
    bash ./dill-0.3.8-patch.sh && \
    pip install git+https://github.com/HabanaAI/DeepSpeed.git@1.17.0 && \
    BUILD_CUDA_EXT=0 pip install git+https://github.com/AutoGPTQ/AutoGPTQ.git@097dd04e --no-build-isolation && \
    pip install . --no-cache-dir

# Install benchmarker
COPY --from=builder /usr/src/target/release-opt/text-generation-benchmark /usr/local/bin/text-generation-benchmark
# Install router
COPY --from=builder /usr/src/target/release-opt/text-generation-router /usr/local/bin/text-generation-router
# Install launcher
COPY --from=builder /usr/src/target/release-opt/text-generation-launcher /usr/local/bin/text-generation-launcher


# AWS Sagemaker compatible image
FROM base AS sagemaker

COPY sagemaker-entrypoint.sh entrypoint.sh
RUN chmod +x entrypoint.sh

ENTRYPOINT ["./entrypoint.sh"]

# Final image
FROM base

COPY ./tgi-entrypoint.sh /tgi-entrypoint.sh
RUN chmod +x /tgi-entrypoint.sh

ENTRYPOINT ["/tgi-entrypoint.sh"]
CMD ["--json-output"]
fix(docker): fix docker image dependencies (#187) 2023-04-16 22:26:47 +00:00			`# Rust builder`
Add some missing modification of 2.3.0 because of conflict Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-09-25 07:49:49 +00:00			`FROM lukemathwalker/cargo-chef:latest-rust-1.80 AS chef`
feat(ci): improve CI speed (#94) 2023-03-03 14:07:27 +00:00			`WORKDIR /usr/src`

Add some missing modification of 2.3.0 because of conflict Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-09-25 07:49:49 +00:00			`ARG CARGO_REGISTRIES_CRATES_IO_PROTOCOL=sparse`

			`FROM chef AS planner`
Fix cargo-chef prepare (#2101) * Fix cargo-chef prepare In prepare stage, cargo-chef reads Cargo.lock and transforms it accordingly. If Cargo.lock is not present, cargo-chef will generate a new one first, which might vary a lot and invalidate docker build caches. * Fix Dockerfile_amd and Dockerfile_intel 2024-06-24 16:16:36 +00:00			`COPY Cargo.lock Cargo.lock`
feat(ci): improve CI speed (#94) 2023-03-03 14:07:27 +00:00			`COPY Cargo.toml Cargo.toml`
			`COPY rust-toolchain.toml rust-toolchain.toml`
			`COPY proto proto`
fix(docker): fix docker build (#299) 2023-05-09 12:39:59 +00:00			`COPY benchmark benchmark`
feat(ci): improve CI speed (#94) 2023-03-03 14:07:27 +00:00			`COPY router router`
Rebase TRT-llm (#2331) * wip wip refacto refacto Initial setup for CXX binding to TRTLLM Working FFI call for TGI and TRTLLM backend Remove unused parameters annd force tokenizer name to be set Overall build TRTLLM and deps through CMake build system Enable end to end CMake build First version loading engines and making it ready for inference Remembering to check how we can detect support for chunked context Move to latest TensorRT-LLM version Specify which default log level to use depending on CMake build type make leader executor mode working unconditionally call InitializeBackend on the FFI layer bind to CUDA::nvml to retrieve compute capabilities at runtime updated logic and comment to detect cuda compute capabilities implement the Stream method to send new tokens through a callback use spdlog release 1.14.1 moving forward update trtllm to latest version a96cccafcf6365c128f004f779160951f8c0801c correctly tell cmake to build dependent tensorrt-llm required libraries create cmake install target to put everything relevant in installation folder add auth_token CLI argument to provide hf hub authentification token allow converting huggingface::tokenizers error to TensorRtLlmBackendError use correct include for spdlog include guard to build example in cmakelists working setup of the ffi layer remove fmt import use external fmt lib end to end ffi flow working make sure to track include/ffi.h to trigger rebuild from cargo impl the rust backend which currently cannot move the actual computation in background thread expose shutdown function at ffi layer impl RwLock scenario for TensorRtLllmBackend oops missing c++ backend definitions compute the number of maximum new tokens for each request independently make sure the context is not dropped in the middle of the async decoding. remove unnecessary log add all the necessary plumbery to return the generated content update invalid doc in cpp file correctly forward back the log probabilities remove unneeded scope variable for now refactor Stream impl for Generation to factorise code expose the internal missing start/queue timestamp forward tgi parameters rep/freq penalty add some more validation about grammar not supported define a shared struct to hold the result of a decoding step expose information about potential error happening while decoding remove logging add logging in case of decoding error make sure executor_worker is provided add initial Dockerfile for TRTLLM backend add some more information in CMakeLists.txt to correctly install executorWorker add some more information in CMakeLists.txt to correctly find and install nvrtc wrapper simplify prebuilt trtllm libraries name definition do the same name definition stuff for tensorrt_llm_executor_static leverage pkg-config to probe libraries paths and reuse new install structure from cmake fix bad copy/past missing nvinfer linkage direction align all the linker search dependency add missing pkgconfig folder for MPI in Dockerfile correctly setup linking search path for runtime layer fix missing / before tgi lib path adding missing ld_library_path for cuda stubs in Dockerfile update tgi entrypoint commenting out Python part for TensorRT installation refactored docker image move to TensorRT-LLM v0.11.0 make docker linter happy with same capitalization rule fix typo refactor the compute capabilities detection along with num gpus update TensorRT-LLM to latest version update TensorRT install script to latest update build.rs to link to cuda 12.5 add missing dependant libraries for linking clean up a bit install to decoder_attention target add some custom stuff for nccl linkage fix envvar CARGO_CFG_TARGET_ARCH set at runtime vs compile time use std::env::const::ARCH make sure variable live long enough... look for cuda 12.5 add some more basic info in README.md * Rebase. * Fix autodocs. * Let's try to enable trtllm backend. * Ignore backends/v3 by default. * Fixing client. * Fix makefile + autodocs. * Updating the schema thing + redocly. * Fix trtllm lint. * Adding pb files ? * Remove cargo fmt temporarily. * ? * Tmp. * Remove both check + clippy ? * Backporting telemetry. * Backporting 457fb0a1 * Remove PB from git. * Fixing PB with default member backends/client * update TensorRT-LLM to latest version * provided None for api_key * link against libtensorrt_llm and not libtensorrt-llm --------- Co-authored-by: OlivierDehaene <23298448+OlivierDehaene@users.noreply.github.com> Co-authored-by: Morgan Funtowicz <morgan@huggingface.co> 2024-07-31 08:33:10 +00:00			`COPY backends backends`
feat(ci): improve CI speed (#94) 2023-03-03 14:07:27 +00:00			`COPY launcher launcher`
Fix tokenization yi (#2507) * Fixing odd tokenization self modifications on the Rust side (load and resave in Python). * Fixing the builds ? * Fix the gh action? * Fixing the location ? * Validation is odd. * Try a faster runner * Upgrade python version. * Remove sccache * No sccache. * Getting libpython maybe ? * List stuff. * Monkey it up. * have no idea at this point * Tmp. * Shot in the dark. * Tmate the hell out of this. * Desperation. * WTF. * -y. * Apparently 3.10 is not available anymore. * Updating the dockerfile to make libpython discoverable at runtime too. * Put back rust tests. * Why do we want mkl on AMD ? * Forcing 3.11 ? 2024-09-11 20:41:56 +00:00
feat(ci): improve CI speed (#94) 2023-03-03 14:07:27 +00:00			`RUN cargo chef prepare --recipe-path recipe.json`

			`FROM chef AS builder`
feat: add distributed tracing (#62) 2023-02-13 12:02:45 +00:00
Fix tokenization yi (#2507) * Fixing odd tokenization self modifications on the Rust side (load and resave in Python). * Fixing the builds ? * Fix the gh action? * Fixing the location ? * Validation is odd. * Try a faster runner * Upgrade python version. * Remove sccache * No sccache. * Getting libpython maybe ? * List stuff. * Monkey it up. * have no idea at this point * Tmp. * Shot in the dark. * Tmate the hell out of this. * Desperation. * WTF. * -y. * Apparently 3.10 is not available anymore. * Updating the dockerfile to make libpython discoverable at runtime too. * Put back rust tests. * Why do we want mkl on AMD ? * Forcing 3.11 ? 2024-09-11 20:41:56 +00:00			`RUN apt-get update && DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends \`
			`python3.11-dev`
feat: add distributed tracing (#62) 2023-02-13 12:02:45 +00:00			`RUN PROTOC_ZIP=protoc-21.12-linux-x86_64.zip && \`
			`curl -OL https://github.com/protocolbuffers/protobuf/releases/download/v21.12/$PROTOC_ZIP && \`
			`unzip -o $PROTOC_ZIP -d /usr/local bin/protoc && \`
			`unzip -o $PROTOC_ZIP -d /usr/local 'include/*' && \`
			`rm -f $PROTOC_ZIP`
feat: Docker image 2022-10-14 13:56:21 +00:00
feat(ci): improve CI speed (#94) 2023-03-03 14:07:27 +00:00			`COPY --from=planner /usr/src/recipe.json recipe.json`
Add some missing modification of 2.3.0 because of conflict Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-09-25 07:49:49 +00:00			`RUN cargo chef cook --profile release-opt --recipe-path recipe.json`

			`ARG GIT_SHA`
			`ARG DOCKER_LABEL`
feat: Docker image 2022-10-14 13:56:21 +00:00
feat(ci): improve CI speed (#94) 2023-03-03 14:07:27 +00:00			`COPY Cargo.toml Cargo.toml`
fix(server): Fix Transformers fork version 2022-11-08 16:42:38 +00:00			`COPY rust-toolchain.toml rust-toolchain.toml`
feat: Docker image 2022-10-14 13:56:21 +00:00			`COPY proto proto`
fix(docker): fix docker build (#299) 2023-05-09 12:39:59 +00:00			`COPY benchmark benchmark`
feat: Docker image 2022-10-14 13:56:21 +00:00			`COPY router router`
Add some missing modification of 2.3.0 because of conflict Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-09-25 07:49:49 +00:00			`COPY backends backends`
v0.1.0 2022-10-18 13:19:03 +00:00			`COPY launcher launcher`
Add some missing modification of 2.3.0 because of conflict Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-09-25 07:49:49 +00:00			`RUN cargo build --profile release-opt`
v0.1.0 2022-10-18 13:19:03 +00:00
fix(docker): fix docker image dependencies (#187) 2023-04-16 22:26:47 +00:00			`# Text Generation Inference base image`
Upgrade SynapseAI version to 1.17.0 (#208) Signed-off-by: yuanwu <yuan.wu@intel.com> Co-authored-by: Thanaji Rao Thakkalapelli <tthakkalapelli@habana.ai> Co-authored-by: regisss <15324346+regisss@users.noreply.github.com> 2024-08-26 08:49:29 +00:00			`FROM vault.habana.ai/gaudi-docker/1.17.0/ubuntu22.04/habanalabs/pytorch-installer-2.3.1:latest as base`
fix(docker): fix docker image dependencies (#187) 2023-04-16 22:26:47 +00:00
			`# Text Generation Inference base env`
Using HF_HOME instead of CACHE to get token read in addition to models. (#2288) 2024-08-09 12:25:44 +00:00			`ENV HF_HOME=/data \`
feat(server): enable hf-transfer (#76) 2023-02-18 13:04:11 +00:00			`HF_HUB_ENABLE_HF_TRANSFER=1 \`
fix(docker): fix docker image dependencies (#187) 2023-04-16 22:26:47 +00:00			`PORT=80`
feat: Docker image 2022-10-14 13:56:21 +00:00
Add changes from Optimum Habana's TGI folder 2023-12-05 10:12:16 +00:00			`# libssl.so.1.1 is not installed on Ubuntu 22.04 by default, install it`
			`RUN wget http://nz2.archive.ubuntu.com/ubuntu/pool/main/o/openssl/libssl1.1_1.1.1f-1ubuntu2_amd64.deb && \`
			`dpkg -i ./libssl1.1_1.1.1f-1ubuntu2_amd64.deb`

fix(docker): revert dockerfile changes (#186) 2023-04-14 17:30:30 +00:00			`WORKDIR /usr/src`
feat(server): flash neoX (#133) 2023-03-24 13:02:14 +00:00
fix(docker): fix docker image dependencies (#187) 2023-04-16 22:26:47 +00:00			`RUN apt-get update && DEBIAN_FRONTEND=noninteractive apt-get install -y --no-install-recommends \`
			`libssl-dev \`
			`ca-certificates \`
			`make \`
Install curl to be able to perform more advanced healthchecks (#1033) # What does this PR do? Install curl within base image, negligible regarding the image volume and will allow to easily perform a better health check. Not sure about the failing github actions though. Should I fix something ? Signed-off-by: Raphael <oOraph@users.noreply.github.com> Co-authored-by: Raphael <oOraph@users.noreply.github.com> 2023-09-26 13:23:47 +00:00			`curl \`
Pali gemma modeling (#1895) This PR adds paligemma modeling code Blog post: https://huggingface.co/blog/paligemma Transformers PR: https://github.com/huggingface/transformers/pull/30814 install the latest changes and run with ```bash # get the weights # text-generation-server download-weights gv-hf/PaliGemma-base-224px-hf # run TGI text-generation-launcher --model-id gv-hf/PaliGemma-base-224px-hf ``` basic example sending various requests ```python from huggingface_hub import InferenceClient client = InferenceClient("http://127.0.0.1:3000") images = [ "https://huggingface.co/datasets/hf-internal-testing/fixtures-captioning/resolve/main/cow_beach_1.png", "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/transformers/rabbit.png", ] prompts = [ "What animal is in this image?", "Name three colors in this image.", "What are 10 colors in this image?", "Where is the cow standing?", "answer en Where is the cow standing?", "Is there a bird in the image?", "Is ther a cow in the image?", "Is there a rabbit in the image?", "how many birds are in the image?", "how many rabbits are in the image?", ] for img in images: print(f"\nImage: {img.split('/')[-1]}") for prompt in prompts: inputs = f"![]({img}){prompt}\n" json_data = { "inputs": inputs, "parameters": { "max_new_tokens": 30, "do_sample": False, }, } generated_output = client.text_generation(prompt, max_new_tokens=30, stream=False) print([f"{prompt}\n{generated_output}"]) ``` --------- Co-authored-by: Nicolas Patry <patry.nicolas@protonmail.com> 2024-05-16 04:58:47 +00:00			`git \`
Make Gaudi adapt to the tgi 2.3.0 Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-09-26 01:53:52 +00:00			`python3.11-dev \`
fix(docker): fix docker image dependencies (#187) 2023-04-16 22:26:47 +00:00			`&& rm -rf /var/lib/apt/lists/*`
feat: Docker image 2022-10-14 13:56:21 +00:00
			`# Install server`
feat(server): Use safetensors Co-authored-by: OlivierDehaene <23298448+OlivierDehaene@users.noreply.github.com> 2022-10-22 18:00:15 +00:00			`COPY proto proto`
feat: Docker image 2022-10-14 13:56:21 +00:00			`COPY server server`
fix(docker): fix docker image dependencies (#187) 2023-04-16 22:26:47 +00:00			`COPY server/Makefile server/Makefile`
feat: Docker image 2022-10-14 13:56:21 +00:00			`RUN cd server && \`
feat(server): Use safetensors Co-authored-by: OlivierDehaene <23298448+OlivierDehaene@users.noreply.github.com> 2022-10-22 18:00:15 +00:00			`make gen-server && \`
fix(docker): fix docker image dependencies (#187) 2023-04-16 22:26:47 +00:00			`pip install -r requirements.txt && \`
A patch to address HPU Graphs issue with DILL A temp solution to address overriding issue installing dill with habana torch from gaudi-docker/1.15.0 - Having `import __main__ as _main_module` in the global space of the dill module causes some overriding issue on hpu graph destructor 2024-04-23 19:57:39 +00:00			`bash ./dill-0.3.8-patch.sh && \`
Upgrade SynapseAI version to 1.17.0 (#208) Signed-off-by: yuanwu <yuan.wu@intel.com> Co-authored-by: Thanaji Rao Thakkalapelli <tthakkalapelli@habana.ai> Co-authored-by: regisss <15324346+regisss@users.noreply.github.com> 2024-08-26 08:49:29 +00:00			`pip install git+https://github.com/HabanaAI/DeepSpeed.git@1.17.0 && \`
Enable the AutoGPTQ (#217) Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-09-25 16:55:02 +00:00			`BUILD_CUDA_EXT=0 pip install git+https://github.com/AutoGPTQ/AutoGPTQ.git@097dd04e --no-build-isolation && \`
Add changes from Optimum Habana's TGI folder 2023-12-05 10:12:16 +00:00			`pip install . --no-cache-dir`
feat: Docker image 2022-10-14 13:56:21 +00:00
feat(docker): add benchmarking tool to docker image (#298) 2023-05-09 11:19:31 +00:00			`# Install benchmarker`
Add some missing modification of 2.3.0 because of conflict Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-09-25 07:49:49 +00:00			`COPY --from=builder /usr/src/target/release-opt/text-generation-benchmark /usr/local/bin/text-generation-benchmark`
feat: Docker image 2022-10-14 13:56:21 +00:00			`# Install router`
Add some missing modification of 2.3.0 because of conflict Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-09-25 07:49:49 +00:00			`COPY --from=builder /usr/src/target/release-opt/text-generation-router /usr/local/bin/text-generation-router`
feat(server): Use safetensors Co-authored-by: OlivierDehaene <23298448+OlivierDehaene@users.noreply.github.com> 2022-10-22 18:00:15 +00:00			`# Install launcher`
Add some missing modification of 2.3.0 because of conflict Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-09-25 07:49:49 +00:00			`COPY --from=builder /usr/src/target/release-opt/text-generation-launcher /usr/local/bin/text-generation-launcher`


			`# AWS Sagemaker compatible image`
			`FROM base AS sagemaker`

			`COPY sagemaker-entrypoint.sh entrypoint.sh`
			`RUN chmod +x entrypoint.sh`

			`ENTRYPOINT ["./entrypoint.sh"]`
feat: Docker image 2022-10-14 13:56:21 +00:00
fea(dockerfile): better layer caching (#159) 2023-04-14 08:12:21 +00:00			`# Final image`
feat: aws sagemaker compatible image (#147) The only difference is that now it pushes to registry.internal.huggingface.tech/api-inference/community/text-generation-inference/sagemaker:... instead of registry.internal.huggingface.tech/api-inference/community/text-generation-inference:sagemaker-... --------- Co-authored-by: Philipp Schmid <32632186+philschmid@users.noreply.github.com> 2023-03-29 19:38:30 +00:00			`FROM base`

Add some missing modification of 2.3.0 because of conflict Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-09-25 07:49:49 +00:00			`COPY ./tgi-entrypoint.sh /tgi-entrypoint.sh`
			`RUN chmod +x /tgi-entrypoint.sh`

Pass the max_batch_total_tokens to causal_lm refine the warmup Signed-off-by: yuanwu <yuan.wu@intel.com> 2024-10-10 07:31:50 +00:00			`ENTRYPOINT ["/tgi-entrypoint.sh"]`
			`CMD ["--json-output"]`