text-generation-inference

mirror of https://github.com/huggingface/text-generation-inference.git synced 2025-10-09 06:55:24 +00:00

History

janne-alatalo 7eeefa3b57 Qwen2-VL runtime error fix when prompted with multiple images (#2840 ) * Fix runtime error when Qwen2-VL was prompted with multiple images Fix runtime error when Qwen2-VL model is prompted with prompt with more than one image. The runtime error was: File "text-generation-inference/server/text_generation_server/models/custom_modeling/qwen2_vl.py", line 459, in get_position_ids text_pos_ids = torch.arange(text_length, device=d) RuntimeError: upper bound and larger bound inconsistent with step sign The error was caused by text_length variable going to negative value when multiple images caused multiple loops in the get_position_ids function's main loop. The error is a simple logic mistake where next_image_pos is initialized as relative offset from current_pos, but was used like it was absolute position from zero. * Fix runtime error when Qwen2-VL was prompted with multiple images Fix runtime error when Qwen2-VL model is prompted with prompt with more than one image. The runtime error was: File "text-generation-inference/server/text_generation_server/models/custom_modeling/qwen2_vl.py", line 534, in forward inputs_embeds[input_ids == self.image_token_id] = image_embeds RuntimeError: shape mismatch: value tensor of shape [512, 3584] cannot be broadcast to indexing result of shape [1024, 3584] (The error message shape numbers can be different depending on the input image resolutions) The error was caused by adding the wrong number of <\|image_pad\|> tokens to the tokenized input in the image_text_replacement function. The error is a simple logical mistake where the number of image pad tokens is checked from pixel_value_shape tensor's first dimension length. However, the pixel_value_shape contains patches from all of the images. Therefore the code added the total number of required image pad tokens for the whole input to each of the images locations. This resulted to extra image pad tokens to be present in the tokenized input. The fix was to check the number of required tokens from the image_grid_thw tensor. The tensor includes grid_t, grid_h, and grid_w values for each image. grid_t * grid_h * grid_w results to the total number of patches for the image [1]. The number of required image pad tokens is number_of_patches // 4. [1] `31f9a289a6/src/transformers/models/qwen2_vl/image_processing_qwen2_vl.py (L311)` --------- Co-authored-by: Janne Alatalo <janne.alatalo@jamk.fi>		2024-12-16 22:55:11 -05:00
..
custom_modeling	Qwen2-VL runtime error fix when prompted with multiple images (#2840 )	2024-12-16 22:55:11 -05:00
__init__.py	Use FP8 KV cache when specified by compressed-tensors (#2761 )	2024-11-26 08:27:41 +01:00
bloom.py	Refactor dead code - Removing all `flash_xxx.py` files. (#2166 )	2024-07-05 10:29:56 +02:00
causal_lm.py	Sync (most) server dependencies with Nix (#2782 )	2024-12-03 04:04:06 +01:00
flash_causal_lm.py	Using both value from config as they might not be correct. (#2817 )	2024-12-10 19:37:09 +01:00
galactica.py	feat: add ruff and resolve issue (#2262 )	2024-07-26 10:29:09 -04:00
globals.py	Attempt for cleverer auto batch_prefill values (some simplifications). (#2808 )	2024-12-09 19:44:32 +01:00
idefics_causal_lm.py	feat: prefill chunking (#2600 )	2024-10-16 12:49:33 +02:00
mamba.py	Choosing input/total tokens automatically based on available VRAM? (#2673 )	2024-10-28 04:59:49 +01:00
metadata_kernels.py	feat: add payload limit (#2726 )	2024-11-21 18:20:15 +00:00
mllama_causal_lm.py	feat: add triton kernels to decrease latency of large batches (#2687 )	2024-10-25 21:10:00 +00:00
model.py	Removing experimental to prefill chunking.	2024-12-06 19:09:40 +01:00
pali_gemma.py	feat: add ruff and resolve issue (#2262 )	2024-07-26 10:29:09 -04:00
seq2seq_lm.py	feat: prefill chunking (#2600 )	2024-10-16 12:49:33 +02:00
types.py	feat: prefill chunking (#2600 )	2024-10-16 12:49:33 +02:00
vlm_causal_lm.py	Qwen2-VL runtime error fix when prompted with multiple images (#2840 )	2024-12-16 22:55:11 -05:00