Choose by Use Case
| Category | Use Case | Recommended Models |
|---|---|---|
| Code & Development | Code generation, reasoning & agentic tasks | DeepSeek V4 Pro, Kimi K2.7 Code, GLM 5.2, MiniMax M2.7 |
| AI Applications | AI agents with tool use | Kimi K2.6, DeepSeek V4 Pro, GLM 5.2, MiniMax M2.7 |
| General reasoning & planning | DeepSeek V4 Pro, Kimi K2.6, GLM 5.2, GPT-OSS 120B (medium) | |
| Long context & summarization | DeepSeek V4 Pro, Kimi K2.6, Qwen3.6 Plus, GLM 5.2, DeepSeek V4 Flash | |
| Fast extraction, classification & search | DeepSeek V4 Flash, MiniMax M2.5, Kimi K2.5, Step 3.7 Flash, GPT-OSS 20B (small) | |
| Vision & Multimodal | Vision & document understanding | Kimi K2.6, Qwen3.6 Plus, Step 3.7 Flash, Gemma 4 31B (small) |
| Audio & video understanding | Qwen3 Omni 30B A3B Instruct, NVIDIA Nemotron 3 Nano Omni 30B A3B | |
| Search & Retrieval | Embeddings & reranking | Qwen3 Embedding 8B, Qwen3 Reranker 8B |
Migrating from Closed Models?
If you’re currently using Claude, OpenAI / GPT, or Gemini models, here’s a guide to the best open source alternatives on Fireworks by use case and latency requirements.Claude Alternatives
OpenAI GPT Alternatives
Google Gemini Alternatives
Understanding Latency Budget:
- High latency budget: Quality is priority. Best for complex reasoning, multi-step workflows, and research tasks where accuracy matters more than speed.
- Low latency budget: Speed is priority. Best for user-facing applications like chatbots, real-time search, and high-throughput classification.
Last updated: July 2026