Skip to main content
Weaviate Docs (migrated from docs.weaviate.io) Docs

Search documentation

Type to search this documentation.

On this pageOverview

Reranker

Weaviate's integration with the Hugging Face Transformers library allows you to access their models' capabilities directly from Weaviate.

Configure a Weaviate collection to use Transformers integration, and configure the Weaviate instance with a model image, and Weaviate will use the specified model in the Transformers inference container to rerank search results.

This two-step process involves Weaviate first performing a search and then reranking the results using the specified model.

Reranker integration illustration

Your Weaviate instance must be configured with the Transformers reranker integration (reranker-transformers) module.

For Weaviate Cloud (WCD) users

This integration is not available for Weaviate Cloud (WCD) instances, as it requires spinning up a container with the Hugging Face model.

To use this integration, configure the container image of the Hugging Face Transformers model and the inference endpoint of the containerized model.

The following example shows how to configure the Hugging Face Transformers integration in Weaviate:

Docker Option 1: Use a pre-configured docker-compose.yml file

Follow the instructions on the Weaviate Docker installation configurator to download a pre-configured docker-compose.yml file with a selected model

Docker Option 2: Add the configuration manually

Alternatively, add the configuration to the docker-compose.yml file manually as in the example below.

YAML
services:
  weaviate:
    # Other Weaviate configuration
    environment:
      RERANKER_INFERENCE_API: http://reranker-transformers:8080  # Set the inference API endpoint
  reranker-transformers:  # Set the name of the inference container
    image: cr.weaviate.io/semitechnologies/reranker-transformers:cross-encoder-ms-marco-MiniLM-L-6-v2
    environment:
      ENABLE_CUDA: 0  # Set to 1 to enable
  • RERANKER_INFERENCE_API environment variable sets the inference API endpoint
  • reranker-transformers is the name of the inference container
  • image is the container image
  • ENABLE_CUDA environment variable enables GPU usage

Set image from a list of available models to specify a particular model to be used.

Configure the Hugging Face Transformers integration in Weaviate by adding or updating the reranker-transformers module in the modules section of the Weaviate Helm chart values file. For example, modify the values.yaml file as follows:

YAML
modules:

  reranker-transformers:

    enabled: true
    tag: cross-encoder-ms-marco-MiniLM-L-6-v2
    repo: semitechnologies/reranker-transformers
    registry: cr.weaviate.io
    envconfig:
      enable_cuda: true

See the Weaviate Helm chart for an example of the values.yaml file including more configuration options.

Set tag from a list of available models to specify a particular model to be used.

As this integration runs a local container with the transformers model, no additional credentials (e.g. API key) are required. Connect to Weaviate as usual, such as in the examples below.

Python
JavaScript/TypeScript

Configure a Weaviate collection to use a Transformer reranker model as follows:

Python
from weaviate.classes.config import Configureclient.collections.create(    "DemoCollection",    reranker_config=Configure.Reranker.transformers()    # Additional parameters not shown)
JavaScript/TypeScript
await client.collections.create({  name: 'DemoCollection',  reranker: weaviate.configure.reranker.transformers(),});

Once the reranker is configured, Weaviate performs reranking operations using the specified reranker model.

More specifically, Weaviate performs an initial search, then reranks the results using the specified model.

Any search in Weaviate can be combined with a reranker to perform reranking operations.

Reranker integration illustration

Python
from weaviate.classes.query import Rerankcollection = client.collections.use("DemoCollection")response = collection.query.near_text(    query="A holiday film",  # The model provider integration will automatically vectorize the query    limit=2,    rerank=Rerank(        prop="title",                   # The property to rerank on        query="A melodic holiday film"  # If not provided, the original query will be used    ))for obj in response.objects:    print(obj.properties["title"])
JavaScript/TypeScript
let myCollection = client.collections.use('DemoCollection');const results = await myCollection.query.nearText(  ['A holiday film'],  {    limit: 2,    rerank: {      property: 'title',                // The property to rerank on      query: 'A melodic holiday film'   // If not provided, the original query will be used    }  });for (const obj of results.objects) {  console.log(obj.properties['title']);}
  • cross-encoder/ms-marco-MiniLM-L-6-v2
  • cross-encoder/ms-marco-MiniLM-L-2-v2
  • cross-encoder/ms-marco-TinyBERT-L-2-v2

These pre-trained models are open-sourced on Hugging Face. The cross-encoder/ms-marco-MiniLM-L-6-v2 model, for example, provides approximately the same benchmark performance as the largest model (L-12) when evaluated on MS-MARCO (39.01 vs. 39.02).

We add new model support over time. For the latest list of available models, see the Docker Hub tags for the reranker-transformers container.

Once the integrations are configured at the collection, the data management and search operations in Weaviate work identically to any other collection. See the following model-agnostic examples:

Have a question or feedback? Here's how to reach us.

Suggest an edit

Propose a replacement for this page. The site team reviews it before applying any changes.

Export
Documentation menu