[Gaudi] Fix the OOM issue of Llama-4-Scout-17B-16E-Instruct (#3245)

Signed-off-by: yuanwu <yuan.wu@intel.com>
2025-10-09 06:55:24 +00:00 · 2025-05-29 15:58:24 +08:00 · 2025-05-29 15:58:24 +08:00 · 70217ac345
commit 70217ac345
parent f14044009a
1 changed files with 8 additions and 6 deletions
--- a/backends/gaudi/server/text_generation_server/models/custom_modeling/flash_llama_modeling.py
+++ b/backends/gaudi/server/text_generation_server/models/custom_modeling/flash_llama_modeling.py
@ -143,6 +143,8 @@ class FlashLlamaAttention(torch.nn.Module):
        config.num_key_value_heads = getattr(
            config, "num_key_value_heads", config.num_attention_heads
        )
+
+        if config.model_type != "llama4_text":
            self.rotary_emb = PositionRotaryEmbedding.static(
                config=config,
                dim=self.head_size,