Skip to main content

Hardware selection

Q: Which accelerator/GPU should I use? It depends on your workload, budget, and which regions you deploy in. Fireworks offers NVIDIA and AMD accelerators across a global fleet. See the on-demand deployments guide and Create Deployment API reference for the authoritative list of supported accelerator types. At a high level:
  • A100 — Lower cost per GPU; a good fit for lower-volume workloads and smaller models.
  • H100 / H200 — Strong general-purpose inference throughput. H200 is often the best fit for very large models (for example, DeepSeek V3 and R1 at scale).
  • B200 / B300 — Latest-generation NVIDIA Blackwell GPUs with high throughput and memory bandwidth.
  • MI325X / MI350X — AMD accelerators with high memory capacity, useful for large models and longer contexts with fewer shards.
GPU availability varies by region. Check the regions table before pinning a deployment to a specific location.

Best Practices for Selection

  1. Analyze your workload requirements to determine which GPU fits your processing needs.
  2. Consider your throughput needs and the scale of your deployment.
  3. Calculate the cost-performance ratio for each hardware option.
  4. Factor in future scaling needs to ensure the selected GPU can support growth.

Additional resources