Hardware selection
Q: Which accelerator/GPU should I use? It depends on your workload, budget, and which regions you deploy in. Fireworks offers NVIDIA and AMD accelerators across a global fleet. See the on-demand deployments guide and Create Deployment API reference for the authoritative list of supported accelerator types. At a high level:- A100 — Lower cost per GPU; a good fit for lower-volume workloads and smaller models.
- H100 / H200 — Strong general-purpose inference throughput. H200 is often the best fit for very large models (for example, DeepSeek V3 and R1 at scale).
- B200 / B300 — Latest-generation NVIDIA Blackwell GPUs with high throughput and memory bandwidth.
- MI325X / MI350X — AMD accelerators with high memory capacity, useful for large models and longer contexts with fewer shards.
Best Practices for Selection
- Analyze your workload requirements to determine which GPU fits your processing needs.
- Consider your throughput needs and the scale of your deployment.
- Calculate the cost-performance ratio for each hardware option.
- Factor in future scaling needs to ensure the selected GPU can support growth.
Additional resources
- Discord Community: discord.gg/fireworks-ai
- Email Support: inquiries@fireworks.ai
- Contact our sales team for custom pricing options