> ## Documentation Index
> Fetch the complete documentation index at: https://docs.fireworks.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Hugging Face

> Learn how developers can integrate and use Fireworks.ai inference capabilities via the Hugging Face ecosystem.

## Hugging Face integration

This documentation provides a concise guide for developers to integrate and use Fireworks.ai inference capabilities via the Hugging Face ecosystem.

## Authentication and billing

When using Fireworks.ai through Hugging Face, you have two options for authentication:

* **Direct Requests:** Use your Fireworks.ai API key in your Hugging Face user account settings. In this mode, inference requests are sent directly to Fireworks.ai, and billing is handled by your Fireworks.ai account.
* **Routed Requests:** If you don't configure a Fireworks.ai API key, your requests will be routed through Hugging Face. In this case, you can use a Hugging Face token for authentication. Billing for routed requests is applied to your Hugging Face account at standard provider API rates. You don’t need an account on Fireworks.ai to do this, just use your HF one!

To add a Fireworks.ai API key to your Hugging Face settings, follow these steps:

1. Go to your [Hugging Face user account settings](https://huggingface.co/settings/inference-providers).
2. Locate the "Inference Providers" section.
3. You can add your API keys for different providers, including Fireworks.ai
4. You can also set your preferred provider order, which will influence the display order in model widgets and code snippets.

<Tip>
  You can search for all [Fireworks.ai models](https://huggingface.co/models?inference_provider=fireworks-ai\&sort=trending) on the hub and directly try out the available models via the Model Page widget too.
</Tip>

## Usage examples

The examples below demonstrate how to interact with various models using Python and JavaScript.

First, ensure you have the [`huggingface_hub`](https://huggingface.co/docs/huggingface_hub/en/guides/inference) library installed (version v0.29.0 or later):

<CodeGroup>
  ```bash Python theme={null}
  pip install huggingface_hub>=0.29.0
  ```

  ```bash JavaScript (Node.js) theme={null}
  npm install @huggingface/inference
  ```
</CodeGroup>

1. Chat Completion (LLMs) with Hugging Face Hub library

<CodeGroup>
  ```python Python theme={null}
  from huggingface_hub import InferenceClient

  # Initialize the InferenceClient with Fireworks.ai as the provider
  client = InferenceClient(
      provider="fireworks-ai", 
      api_key="xxxxxxxxxxxxxxxxxxxxxxxx"  # Replace with your API key (HF or custom)
  )

  # Define the chat messages
  messages = [
      {
          "role": "user",
          "content": "What is the capital of France?"
      }
  ]

  # Generate a chat completion
  completion = client.chat.completions.create(
      model="deepseek-ai/DeepSeek-R1",  
      messages=messages, 
      max_tokens=500
  )

  # Print the response
  print(completion.choices[0].message)
  ```

  ```javascript JavaScript theme={null}
  import { HfInference } from "@huggingface/inference";

  const client = new HfInference("xxxxxxxxxxxxxxxxxxxxxxxx");

  // Define the chat messages
  const messages = [
      {
          role: "user",
          content: "What is the capital of France?"
      }
  ];

  const completion = await client.chatCompletion({
      model: "deepseek-ai/DeepSeek-R1",
      messages: messages,
      max_tokens: 500,
  });

  console.log(completion.choices[0].message);

  ```
</CodeGroup>

You can swap this for any compatible LLM from Fireworks.ai, here’s a handy URL to find the list: [here](https://huggingface.co/models?inference_provider=fireworks-ai\&sort=trending)

2. Vision Language Models (VLMs) with Hugging Face Hub Library

<CodeGroup>
  ```python Python theme={null}
  import os
  from huggingface_hub import InferenceClient

  client = InferenceClient(
      provider="fireworks-ai",
      api_key=os.environ["HF_TOKEN"],
  )

  completion = client.chat.completions.create(
      model="Qwen/Qwen2.5-VL-32B-Instruct",
      messages=[
          {
              "role": "user",
              "content": [
                  {
                      "type": "text",
                      "text": "Describe this image in one sentence."
                  },
                  {
                      "type": "image_url",
                      "image_url": {
                          "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
                      }
                  }
              ]
          }
      ],
  )

  print(completion.choices[0].message)
  ```

  ```javascript JavaScript theme={null}
  import { HfInference } from "@huggingface/inference";

  const client = new HfInference("xxxxxxxxxxxxxxxxxxxxxxxx");

  const messages = [
      {
          role: "user",
          content: [
              {
                  type: "text",
                  text: "Describe this image in one sentence."
              },
              {
                  type: "image_url",
                  image_url: {
                      url: "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg"
                  }
              }
          ]
      }
  ];

  const completion = await client.chatCompletion({
      model: "Qwen/Qwen2.5-VL-32B-Instruct",
      messages: messages,
      max_tokens: 500,
  });

  console.log(completion.choices[0].message);
  ```
</CodeGroup>

Similar to LLMs, you can use any compatible VLM model from the list [here](https://huggingface.co/models?inference_provider=fireworks-ai\&sort=trending)

You can also call inference providers via the [OpenAI python client](https://github.com/openai/openai-python). You will need to specify the `base_url` and `model` parameters in the client and call respectively.

The easiest way is to go to [a model’s page](https://huggingface.co/deepseek-ai/DeepSeek-R1?inference_api=true\&inference_provider=\<provider_name>\&language=python) on the hub and copy the snippet.

<CodeGroup>
  ```python Python theme={null}
  from openai import OpenAI

  client = OpenAI(
  	base_url="https://router.huggingface.co/fireworks-ai/inference/v1",
  	api_key="xxxxxxxxxxxxxxxxxxxxxxxx" #fireworks or Hugging Face API key 
  )

  messages = [
  	{
  		"role": "user",
  		"content": "What is the capital of France?"
  	}
  ]

  completion = client.chat.completions.create(
  	model="<provider_specific_model_name>", 
  	messages=messages, 
  	max_tokens=500,
  )

  print(completion.choices[0].message)
  ```

  ```javascript JavaScript theme={null}
  import { HfInference } from "@huggingface/inference";

  // Initialize the HfInference client with your API key
  const client = new HfInference("xxxxxxxxxxxxxxxxxxxxxxxx");

  // Generate a chat completion
  const chatCompletion = await client.chatCompletion({
      model: "deepseek-ai/DeepSeek-R1",  // Replace with your desired model
      messages: [
          {
              role: "user",
              content: "What is the capital of France?"
          }
      ],
      provider: "fireworks-ai", 
      max_tokens: 500
  });

  // Log the response
  console.log(chatCompletion.choices[0].message);
  ```
</CodeGroup>

You can search for all [Fireworks.ai models](https://huggingface.co/models?inference_provider=fireworks-ai\&sort=trending) on the hub and directly try out the available models via the Model Page widget too.

We’ll continue to increase the number of models and ways to try it out!
