# Reranker

Weaviate's integration with the Hugging Face Transformers library allows you to access their models' capabilities directly from Weaviate.

[Configure a Weaviate collection](#configure-the-reranker) to use Transformers integration, and [configure the Weaviate instance](#weaviate-configuration) with a model image, and Weaviate will use the specified model in the Transformers inference container to rerank search results.

This two-step process involves Weaviate first performing a search and then reranking the results using the specified model.

![Reranker integration illustration](/assets/docs/weaviate/model-providers/_includes/integration_transformers_reranker.png)

## Requirements

### Weaviate configuration

Your Weaviate instance must be configured with the Transformers reranker integration (`reranker-transformers`) module.

:::accordion{title="For Weaviate Cloud (WCD) users"}
This integration is not available for Weaviate Cloud (WCD) instances, as it requires spinning up a container with the Hugging Face model.
:::

#### Enable the integration module

- Check the [cluster metadata](../monitoring-and-logging/status.md#cluster-metadata) to verify if the module is enabled.
- Follow the [how-to configure modules](../how-to-configure-weaviate/modules.md) guide to enable the module in Weaviate.

#### Configure the integration

To use this integration, configure the container image of the Hugging Face Transformers model and the inference endpoint of the containerized model.

The following example shows how to configure the Hugging Face Transformers integration in Weaviate:

::::tabs{sync="deployments"}
:::tab{title="Docker"}
#### Docker Option 1: Use a pre-configured `docker-compose.yml` file

Follow the instructions on the [Weaviate Docker installation configurator](../installation/installation-guides-docker-installation.md#configurator) to download a pre-configured `docker-compose.yml` file with a selected model

#### Docker Option 2: Add the configuration manually

Alternatively, add the configuration to the `docker-compose.yml` file manually as in the example below.

```yaml
services:
  weaviate:
    # Other Weaviate configuration
    environment:
      RERANKER_INFERENCE_API: http://reranker-transformers:8080  # Set the inference API endpoint
  reranker-transformers:  # Set the name of the inference container
    image: cr.weaviate.io/semitechnologies/reranker-transformers:cross-encoder-ms-marco-MiniLM-L-6-v2
    environment:
      ENABLE_CUDA: 0  # Set to 1 to enable
```

- `RERANKER_INFERENCE_API` environment variable sets the inference API endpoint
- `reranker-transformers` is the name of the inference container
- `image` is the container image
- `ENABLE_CUDA` environment variable enables GPU usage

Set `image` from a [list of available models](#available-models) to specify a particular model to be used.
:::

:::tab{title="Kubernetes"}
Configure the Hugging Face Transformers integration in Weaviate by adding or updating the `reranker-transformers` module in the `modules` section of the Weaviate Helm chart values file. For example, modify the `values.yaml` file as follows:

```yaml
modules:

  reranker-transformers:

    enabled: true
    tag: cross-encoder-ms-marco-MiniLM-L-6-v2
    repo: semitechnologies/reranker-transformers
    registry: cr.weaviate.io
    envconfig:
      enable_cuda: true
```

See the [Weaviate Helm chart](https://github.com/weaviate/weaviate-helm/blob/master/weaviate/values.yaml) for an example of the `values.yaml` file including more configuration options.

Set `tag` from a [list of available models](#available-models) to specify a particular model to be used.
:::
::::

### Credentials

As this integration runs a local container with the transformers model, no additional credentials (e.g. API key) are required. Connect to Weaviate as usual, such as in the examples below.

:::code-group{sync="languages"}
```python title="Python"
```

```typescript title="JavaScript/TypeScript"
```
:::

## Configure the reranker

:::callout{intent="info" title="Reranker model integration mutable from `v1.25.23`, `v1.26.8` and `v1.27.1`"}
A collection's `reranker` model integration configuration is mutable from `v1.25.23`, `v1.26.8` and `v1.27.1`. See [this section](../how-to-manage-collections/generative-reranker-models.md#update-the-reranker-model-integration) for details on how to update the collection configuration.
:::

Configure a Weaviate collection to use a Transformer reranker model as follows:

:::code-group{sync="languages"}
```python title="Python" {5}
from weaviate.classes.config import Configure

client.collections.create(
    "DemoCollection",
    reranker_config=Configure.Reranker.transformers()
    # Additional parameters not shown
)
```

```typescript title="JavaScript/TypeScript" {3}
await client.collections.create({
  name: 'DemoCollection',
  reranker: weaviate.configure.reranker.transformers(),
});
```
:::

:::callout{intent="note" title="Choose a container image to select a model"}
To choose a model, select the [container image](#configure-the-integration) that hosts it.
:::

## Reranking query

Once the reranker is configured, Weaviate performs [reranking operations](../how-to-query-search/rerank.md) using the specified reranker model.

More specifically, Weaviate performs an initial search, then reranks the results using the specified model.

Any search in Weaviate can be combined with a reranker to perform reranking operations.

![Reranker integration illustration](/assets/docs/weaviate/model-providers/_includes/integration_transformers_reranker.png)

:::code-group{sync="languages"}
```python title="Python" {5-12}
from weaviate.classes.query import Rerank

collection = client.collections.use("DemoCollection")

response = collection.query.near_text(
    query="A holiday film",  # The model provider integration will automatically vectorize the query
    limit=2,
    rerank=Rerank(
        prop="title",                   # The property to rerank on
        query="A melodic holiday film"  # If not provided, the original query will be used
    )
)

for obj in response.objects:
    print(obj.properties["title"])
```

```typescript title="JavaScript/TypeScript" {5-11}
let myCollection = client.collections.use('DemoCollection');

const results = await myCollection.query.nearText(
  ['A holiday film'],
  {
    limit: 2,
    rerank: {
      property: 'title',                // The property to rerank on
      query: 'A melodic holiday film'   // If not provided, the original query will be used
    }
  }
);

for (const obj of results.objects) {
  console.log(obj.properties['title']);
}
```
:::

## References

### Available models

- `cross-encoder/ms-marco-MiniLM-L-6-v2`
- `cross-encoder/ms-marco-MiniLM-L-2-v2`
- `cross-encoder/ms-marco-TinyBERT-L-2-v2`

These pre-trained models are open-sourced on Hugging Face. The `cross-encoder/ms-marco-MiniLM-L-6-v2` model, for example, provides approximately the same benchmark performance as the largest model (L-12) when evaluated on [MS-MARCO](https://microsoft.github.io/msmarco/) (39.01 vs. 39.02).

We add new model support over time. For the latest list of available models, see the Docker Hub tags for the [reranker-transformers](https://hub.docker.com/r/semitechnologies/reranker-transformers/tags) container.

## Further resources

### Other integrations

- [Transformers embedding models + Weaviate](transformers-embeddings.md).
- [Transformers multi-modal embedding models + Weaviate](transformers-embeddings-multimodal.md).

### Code examples

Once the integrations are configured at the collection, the data management and search operations in Weaviate work identically to any other collection. See the following model-agnostic examples:

- The [How-to: Manage collections](../how-to-manage-collections/index.md) and [How-to: Manage objects](../how-to-manage-objects/index.md) guides show how to perform data operations (i.e. create, read, update, delete collections and objects within them).
- The [How-to: Query & Search](../how-to-query-search/index.md) guides show how to perform search operations (i.e. vector, keyword, hybrid) as well as retrieval augmented generation.

## Questions and feedback

Have a question or feedback? Here's how to reach us.

::::card-grid
:::card{title="Community Forum" href="https://forum.weaviate.io/c/support" icon="messages-square"}
Ask questions and connect with other developers on our **Community forum**.
:::

:::card{title="Support" href="/guides/support-overview" icon="life-buoy"}
Weaviate Cloud user or customer? Find the right channel on the **Support page**.
:::
::::

## Related pages

- [Agents](./agents-index.md)
- [AI-assisted Weaviate code generation](./ai-assisted-vibe-coding-index.md)
- [APIs](./apis-index.md)
- [Authorization and authentication](./authorization-and-authentication-index.md)
- [Benchmarks](./benchmarks-index.md)
- [Best practices](./best-practices-index.md)
- [Client libraries](./clients-index.md)
- [Client Libraries / SDKs](./client-libraries-index.md)
- [Cloud](./cloud-index.md)
- [Cloud account management](./cloud-account-management-index.md)

# Agent Instructions

This portal answers questions programmatically. To receive a synthesized,
source-cited answer instead of crawling page by page, append the `?ask=`
query parameter to any page URL on this site:

    /guides/quickstart?ask=how+do+I+authenticate

Optional parameters:

- `&goal=<what-you-are-trying-to-do>` steers the answer toward your
  objective (e.g. `&goal=write+a+python+client`).
- `&version=<label>` scopes the answer to a mounted version when the
  portal publishes more than one.

The response is `text/markdown`: the answer followed by a `# Sources` list
of the portal pages it was grounded in. Status codes are the contract:

- `200` — the answer; `402` — the portal owner’s plan or answer credits are
  exhausted (surface this to your operator; do NOT retry); `429` — you are
  rate-limited; back off for the `Retry-After` seconds; `503` — the answer
  lane is temporarily unavailable; fall back to crawling the `.md` pages.

For the full corpus map read `llms.txt` at the site root; for the tool
surface (search + page fetch as MCP tools) see `/mcp`.
