Skip to main content
Quantization reduces the number of bits used to serve a model, improving performance and reducing cost by 30-50%. However, this can change model numerics which may introduce small changes to the output.
Read our blog post for a detailed treatment of how quantization affects model quality.

Checking available precisions

Models may support different numerical precisions like FP16, FP8, BF16, or INT8, which affect memory usage and inference speed. Check default precision:
Check supported precisions:
The Precisions field indicates what precisions the model has been prepared for.

Quantizing a model

A model can be quantized to 8-bit floating-point (FP8) precision.
This is an additive process that enables creating deployments with additional precisions. The original FP16 checkpoint is still available for use.
You can check on the status of preparation by running:
and checking if the state is still in PREPARING. A successfully prepared model will have the desired precision added to the Precisions list.

Creating an FP8 deployment

By default, creating a deployment uses the FP16 checkpoint. To use a quantized FP8 checkpoint, first ensure the model has been prepared for FP8 (see Checking available precisions above). Do not deploy without a deployment shape — a shape that includes FP8 precision is the only validated path. List the model’s shapes with firectl deployment-shape-version list --base-model <MODEL> and look for one with the precision you want. If no FP8 shape exists for your model, start from the closest shape and override the precision with --precision. Do not create the deployment without a shape entirely: it skips validation, is the most common cause of failed deployment creations, and the unshaped path may be deprecated in the future:
Quantized deployments can only be served using H100 GPUs.