text-generation-inference/server/text_generation_server/layers/moe
Mohit Sharma 4ef2e045c9
Add fp8 support moe models (#2928)
* Add fp8 support moe models

* flatten condition
2025-01-29 13:56:32 +01:00
..
__init__.py Add fp8 support moe models (#2928) 2025-01-29 13:56:32 +01:00
fp8.py Add fp8 support moe models (#2928) 2025-01-29 13:56:32 +01:00
fused_moe_ipex.py fix moe in quantization path (#2935) 2025-01-22 14:36:15 +01:00
gptq_marlin.py Add support for fused MoE Marlin for AWQ (#2616) 2024-10-08 11:56:41 +02:00
unquantized.py Add fp8 support moe models (#2928) 2025-01-29 13:56:32 +01:00