> ## Documentation Index
> Fetch the complete documentation index at: https://docs.fireworks.ai/llms.txt
> Use this file to discover all available pages before exploring further.

> ## Agent Instructions
> For Fireworks Nexus, start at https://docs.fireworks.ai/nexus.
> Use https://docs.fireworks.ai/nexus/quickstart for coding harnesses, custom agents, APIs, SDKs, and LLM gateways.
> Use https://docs.fireworks.ai/nexus/firerouter for how model routers work, the supported model list, composition, closed-model credentials, and pricing.
> Prefer canonical short model IDs such as firerouter/opus. In LiteLLM litellm_params.model, use the full path fireworks_ai/accounts/fireworks/routers/firerouter/opus.
> Family names such as opus track the latest evaluated family version; do not describe them as fixed model versions.

# Inference performance

> Understanding model performance, quantization, and batching capabilities.

## Model quantization

**Q: What quantization format is used for the Llama 3.1 405B model?**

The **Llama 3.1 405B model** uses the **FP8 quantization format**, which:

* Closely matches **Meta's reference implementation**
* Provides further details in the model description at [fireworks.ai/models/fireworks/llama-v3p1-405b-instruct](https://fireworks.ai/models/llama-v3p1-405b-instruct)
* Has a general quantization methodology documented in our [Quantization blog](https://fireworks.ai/blog/fireworks-quantization)

*Note*: **BF16 precision** will be available soon for on-demand deployments.

***

## API capabilities

**Q: Does the API support batching and load balancing?**

Current capabilities include:

* **Load balancing**: Yes, supported out of the box
* **Continuous batching**: Yes, supported
* **Batch inference**: Not currently supported (on the roadmap)
  * Note: For batch use cases, we recommend sending multiple parallel HTTP requests to the deployment while maintaining some fixed level of concurrency.
* **Streaming**: Yes, supported

***

## Request handling

**Q: What factors affect the number of simultaneous requests that can be handled?**

Request handling capacity depends on several factors:

* **Model size and type**
* **Number of GPUs allocated** to the deployment
* **GPU type** (e.g., A100, H100)
* **Prompt size**
* **Generation token length**
* **Deployment type** (serverless vs. on-demand)

***

## Additional information

If you experience any issues during these processes, you can:

* Contact support through Discord at [discord.gg/fireworks-ai](https://discord.gg/fireworks-ai)
* Reach out to your account representative (Enterprise customers)
* Email [inquiries@fireworks.ai](mailto:inquiries@fireworks.ai)
