mirror of https://github.com/huggingface/text-generation-inference.git synced 2025-07-14 20:00:17 +00:00

* Upgrade the version number.

* Remove modifications in Lock.

* Tmp branch to test transformers backend with 2.5.1 and TP>1

* Fixing the transformers backend.

inference_mode forces the use of `aten.matmul` instead of `aten.mm` the
former doesn't have sharding support crashing the transformers TP
support.

`lm_head.forward` also crashes because it skips the hook that
cast/decast the DTensor.

Torch 2.5.1 is required for sharding support.

* Put back the attention impl.

* Revert the flashinfer (this will fails).

* Building AOT.

* Using 2.5 kernels.

* Remove the archlist, it's defined in the docker anyway.

2025-01-23 18:07:30 +01:00

854 B

Raw Blame History

Serving Private & Gated Models

If the model you wish to serve is behind gated access or the model repository on Hugging Face Hub is private, and you have access to the model, you can provide your Hugging Face Hub access token. You can generate and copy a read token from Hugging Face Hub tokens page

If you're using the CLI, set the HF_TOKEN environment variable. For example:

export HF_TOKEN=<YOUR READ TOKEN>

If you would like to do it through Docker, you can provide your token by specifying HF_TOKEN as shown below.

model=meta-llama/Llama-2-7b-chat-hf
volume=$PWD/data
token=<your READ token>

docker run --gpus all \
    --shm-size 1g \
    -e HF_TOKEN=$token \
    -p 8080:80 \
    -v $volume:/data ghcr.io/huggingface/text-generation-inference:3.0.2 \
    --model-id $model

854 B Raw Blame History

Serving Private & Gated Models

854 B

Raw Blame History