Weaviate's integration with NVIDIA's APIs allows you to access their models' capabilities directly from Weaviate.

[Configure a Weaviate collection](#configure-the-reranker) to use an NVIDIA reranker model, and Weaviate will use the specified model and your NVIDIA NIM API key to rerank search results.

This two-step process involves Weaviate first performing a search and then reranking the results using the specified model.

![Reranker integration illustration](/assets/docs/weaviate/model-providers/_includes/integration_nvidia_reranker.png)

## Requirements

### Weaviate configuration

Your Weaviate instance must be configured with the NVIDIA reranker integration (`reranker-nvidia`) module.

:::accordion{title="For Weaviate Cloud (WCD) users"}
This integration is enabled by default on Weaviate Cloud (WCD) instances.
:::

:::accordion{title="For self-hosted users"}
- Check the [cluster metadata](../monitoring-and-logging/status.md#cluster-metadata) to verify if the module is enabled.
- Follow the [how-to configure modules](../how-to-configure-weaviate/modules.md) guide to enable the module in Weaviate.
:::

### API credentials

You must provide a valid NVIDIA NIM API key to Weaviate for this integration. Go to [NVIDIA](https://build.nvidia.com/) to sign up and obtain an API key.

Provide the API key to Weaviate using one of the following methods:

- Set the `NVIDIA_APIKEY` environment variable that is available to Weaviate.
- Provide the API key at runtime, as shown in the examples below.

:::code-group{sync="languages"}
```python title="Python"
# Recommended: save sensitive data as environment variables
nvidia_key = os.getenv("NVIDIA_API_KEY")
```

```typescript title="JavaScript/TypeScript"
const nvidiaApiKey = process.env.NVIDIA_API_KEY || '';  // Replace with your inference API key
```
:::

## Configure the reranker

:::callout{intent="info" title="Reranker model integration mutable from `v1.25.23`, `v1.26.8` and `v1.27.1`"}
A collection's `reranker` model integration configuration is mutable from `v1.25.23`, `v1.26.8` and `v1.27.1`. See [this section](../how-to-manage-collections/generative-reranker-models.md#update-the-reranker-model-integration) for details on how to update the collection configuration.
:::

Configure a Weaviate collection to use an NVIDIA reranker model as follows:

:::code-group{sync="languages"}
```python title="Python" {5}
from weaviate.classes.config import Configure

client.collections.create(
    "DemoCollection",
    reranker_config=Configure.Reranker.nvidia()
    # Additional parameters not shown
)
```

```typescript title="JavaScript/TypeScript" {3}
await client.collections.create({
  name: 'DemoCollection',
  reranker: weaviate.configure.reranker.nvidia(),
});
```
:::

### Select a model

You can specify one of the [available models](#available-models) for Weaviate to use, as shown in the following configuration example:

:::code-group{sync="languages"}
```python title="Python" {5-8}
from weaviate.classes.config import Configure

client.collections.create(
    "DemoCollection",
    reranker_config=Configure.Reranker.nvidia(
        model="nvidia/llama-3.2-nv-rerankqa-1b-v2",
        base_url="https://integrate.api.nvidia.com/v1",
    )
    # Additional parameters not shown
)
```

```typescript title="JavaScript/TypeScript" {3-6}
await client.collections.create({
  name: 'DemoCollection',
  reranker: weaviate.configure.reranker.nvidia({
    model: "nvidia/llama-3.2-nv-rerankqa-1b-v2",
    baseURL: "https://integrate.api.nvidia.com/v1"
  }),
});
```
:::

The [default model](#available-models) is used if no model is specified.

## Reranking query

Once the reranker is configured, Weaviate performs [reranking operations](../how-to-query-search/rerank.md) using the specified NVIDIA model.

More specifically, Weaviate performs an initial search, then reranks the results using the specified model.

Any search in Weaviate can be combined with a reranker to perform reranking operations.

![Reranker integration illustration](/assets/docs/weaviate/model-providers/_includes/integration_nvidia_reranker.png)

:::code-group{sync="languages"}
```python title="Python" {5-12}
from weaviate.classes.query import Rerank

collection = client.collections.use("DemoCollection")

response = collection.query.near_text(
    query="A holiday film",  # The model provider integration will automatically vectorize the query
    limit=2,
    rerank=Rerank(
        prop="title",                   # The property to rerank on
        query="A melodic holiday film"  # If not provided, the original query will be used
    )
)

for obj in response.objects:
    print(obj.properties["title"])
```

```typescript title="JavaScript/TypeScript" {5-11}
let myCollection = client.collections.use('DemoCollection');

const results = await myCollection.query.nearText(
  ['A holiday film'],
  {
    limit: 2,
    rerank: {
      property: 'title',                // The property to rerank on
      query: 'A melodic holiday film'   // If not provided, the original query will be used
    }
  }
);

for (const obj of results.objects) {
  console.log(obj.properties['title']);
}
```
:::

## References

### Available models

You can use any reranker model [on NVIDIA NIM APIs](https://build.nvidia.com/models) with Weaviate.

The default model is `nvidia/rerank-qa-mistral-4b`.

## Further resources

### Other integrations

- [NVIDIA text embedding models + Weaviate](nvidia-embeddings.md).
- [NVIDIA multimodal embedding models + Weaviate](nvidia-embeddings-multimodal.md).
- [NVIDIA generative models + Weaviate](nvidia-generative.md).

### Code examples

Once the integrations are configured at the collection, the data management and search operations in Weaviate work identically to any other collection. See the following model-agnostic examples:

- The [How-to: Manage collections](../how-to-manage-collections/index.md) and [How-to: Manage objects](../how-to-manage-objects/index.md) guides show how to perform data operations (i.e. create, read, update, delete collections and objects within them).
- The [How-to: Query & Search](../how-to-query-search/index.md) guides show how to perform search operations (i.e. vector, keyword, hybrid) as well as retrieval augmented generation.

### References

- [NVIDIA NIM API documentation](https://docs.api.nvidia.com/nim/)

## Questions and feedback

Have a question or feedback? Here's how to reach us.

::::card-grid
:::card{title="Community Forum" href="https://forum.weaviate.io/c/support" icon="messages-square"}
Ask questions and connect with other developers on our **Community forum**.
:::

:::card{title="Support" href="/guides/support-overview" icon="life-buoy"}
Weaviate Cloud user or customer? Find the right channel on the **Support page**.
:::
::::

## Related pages

- [Agents](./agents-index.md)
- [AI-assisted Weaviate code generation](./ai-assisted-vibe-coding-index.md)
- [APIs](./apis-index.md)
- [Authorization and authentication](./authorization-and-authentication-index.md)
- [Benchmarks](./benchmarks-index.md)
- [Best practices](./best-practices-index.md)
- [Client libraries](./clients-index.md)
- [Client Libraries / SDKs](./client-libraries-index.md)
- [Cloud](./cloud-index.md)
- [Cloud account management](./cloud-account-management-index.md)

# Agent Instructions

This portal answers questions programmatically. To receive a synthesized,
source-cited answer instead of crawling page by page, append the `?ask=`
query parameter to any page URL on this site:

    /guides/quickstart?ask=how+do+I+authenticate

Optional parameters:

- `&goal=<what-you-are-trying-to-do>` steers the answer toward your
  objective (e.g. `&goal=write+a+python+client`).
- `&version=<label>` scopes the answer to a mounted version when the
  portal publishes more than one.

The response is `text/markdown`: the answer followed by a `# Sources` list
of the portal pages it was grounded in. Status codes are the contract:

- `200` — the answer; `402` — the portal owner’s plan or answer credits are
  exhausted (surface this to your operator; do NOT retry); `429` — you are
  rate-limited; back off for the `Retry-After` seconds; `503` — the answer
  lane is temporarily unavailable; fall back to crawling the `.md` pages.

For the full corpus map read `llms.txt` at the site root; for the tool
surface (search + page fetch as MCP tools) see `/mcp`.
