IT & AI Technology
Native FP8 and FP4 Support, and "fp8" Quantization with vLLM
Hello, this is the Product Development Department at Qualiteg. When a new model is released, we often go hunting for the right recipe to speed up inference — quantizing with various techniques and switching between multiple inference engines. In the course of that work, we ran into the following issue. Since