Weaviate's integration with Anyscale's Endpoints APIs allows you to access their models' capabilities directly from Weaviate.

[Configure a Weaviate collection](#configure-collection) to use a generative AI model with Anyscale. Weaviate will perform retrieval augmented generation (RAG) using the specified model and your Anyscale API key.

More specifically, Weaviate will perform a search, retrieve the most relevant objects, and then pass them to the Anyscale generative model to generate outputs.

![RAG integration illustration](/assets/docs/weaviate/model-providers/_includes/integration_anyscale_rag.png)

## Requirements

### Weaviate configuration

Your Weaviate instance must be configured with the Anyscale generative AI integration (`generative-anyscale`) module.

:::accordion{title="For Weaviate Cloud (WCD) users"}
This integration is enabled by default on Weaviate Cloud (WCD) instances.
:::

:::accordion{title="For self-hosted users"}
- Check the [cluster metadata](../monitoring-and-logging/status.md#cluster-metadata) to verify if the module is enabled.
- Follow the [how-to configure modules](../how-to-configure-weaviate/modules.md) guide to enable the module in Weaviate.
:::

### API credentials

You must provide a valid Anyscale API key to Weaviate for this integration. Go to [Anyscale](https://www.anyscale.com/) to sign up and obtain an API key.

Provide the API key to Weaviate using one of the following methods:

- Set the `ANYSCALE_APIKEY` environment variable that is available to Weaviate.
- Provide the API key at runtime, as shown in the examples below.

:::code-group{sync="languages"}
```python title="Python"
# Recommended: save sensitive data as environment variables
anyscale_key = os.getenv("ANYSCALE_API_KEY")
```

```typescript title="JavaScript/TypeScript"
const anyscaleApiKey = process.env.ANYSCALE_API_KEY || '';  // Replace with your inference API key
```
:::

## Configure collection

:::callout{intent="info" title="Generative model integration mutability"}
A collection's `generative` model integration configuration is mutable from `v1.25.23`, `v1.26.8` and `v1.27.1`. See [this section](../how-to-manage-collections/generative-reranker-models.md#update-the-generative-model-integration) for details on how to update the collection configuration.
:::

[Configure a Weaviate index](../how-to-manage-collections/generative-reranker-models.md#specify-a-generative-model-integration) as follows to use an Anyscale generative model:

:::code-group{sync="languages"}
```python title="Python" {5}
from weaviate.classes.config import Configure

client.collections.create(
    "DemoCollection",
    generative_config=Configure.Generative.anyscale()
    # Additional parameters not shown
)
```

```typescript title="JavaScript/TypeScript" {3}
await client.collections.create({
  name: 'DemoCollection',
  generative: weaviate.configure.generative.anyscale(),
  // Additional parameters not shown
});
```
:::

You can [specify](#generative-parameters) one of the [available models](#available-models) for Weaviate to use. The [default model](#available-models) is used if no model is specified.

### Select a model

You can specify one of the [available models](#available-models) for Weaviate to use, as shown in the following configuration example:

:::code-group{sync="languages"}
```python title="Python" {5-7}
from weaviate.classes.config import Configure

client.collections.create(
    "DemoCollection",
    generative_config=Configure.Generative.anyscale(
        model="mistralai/Mixtral-8x7B-Instruct-v0.1"
    )
    # Additional parameters not shown
)
```

```typescript title="JavaScript/TypeScript" {3-5}
await client.collections.create({
  name: 'DemoCollection',
  generative: weaviate.configure.generative.anyscale({
    model: 'mistralai/Mixtral-8x7B-Instruct-v0.1'
  }),
  // Additional parameters not shown
});
```
:::

### Generative parameters

Configure the following generative parameters to customize the model behavior.

:::code-group{sync="languages"}
```python title="Python" {5-9}
from weaviate.classes.config import Configure

client.collections.create(
    "DemoCollection",
    generative_config=Configure.Generative.anyscale(
        # # These parameters are optional
        # model="meta-llama/Llama-2-70b-chat-hf",
        # temperature=0.7,
    )
    # Additional parameters not shown
)
```

```typescript title="JavaScript/TypeScript" {3-7}
await client.collections.create({
  name: 'DemoCollection',
  generative: weaviate.configure.generative.anyscale({
    // These parameters are optional
    // model: 'meta-llama/Llama-2-70b-chat-hf',
    // temperature: 0.7,
  }),
  // Additional parameters not shown
});
```
:::

For further details on model parameters, see the [Anyscale Endpoints API documentation](https://docs.anyscale.com/endpoints/intro/).

## Select a model at runtime

Aside from setting the default model provider when creating the collection, you can also override it at query time.

:::code-group{sync="languages"}
```python title="Python" {9-15}
from weaviate.classes.config import Configure
from weaviate.classes.generate import GenerativeConfig

collection = client.collections.use("DemoCollection")
response = collection.generate.near_text(
    query="A holiday film",
    limit=2,
    grouped_task="Write a tweet promoting these two movies",
    generative_provider=GenerativeConfig.anyscale(
        # # These parameters are optional
        # base_url="https://api.anthropic.com",
        # model="meta-llama/Llama-2-70b-chat-hf",
        # temperature=0.7,
    ),
    # Additional parameters not shown
)
```

```typescript title="JavaScript/TypeScript"
import { generativeParameters } from 'weaviate-client';
```
:::

## Header parameters

You can provide the API key as well as some optional parameters at runtime through additional headers in the request. The following headers are available:

- `X-Anyscale-Api-Key`: The Anyscale API key.
- `X-Anyscale-Baseurl`: The base URL to use (e.g. a proxy) instead of the default Anyscale URL.

Any additional headers provided at runtime will override the existing Weaviate configuration.

Provide the headers as shown in the [API credentials examples](#api-credentials) above.

## Retrieval augmented generation

After configuring the generative AI integration, perform RAG operations, either with the [single prompt](#single-prompt) or [grouped task](#grouped-task) method.

### Single prompt

![Single prompt RAG integration generates individual outputs per search result](/assets/docs/weaviate/model-providers/_includes/integration_anyscale_rag_single.png)

To generate text for each object in the search results, use the single prompt method.

The example below generates outputs for each of the `n` search results, where `n` is specified by the `limit` parameter.

When creating a single prompt query, use braces `{}` to interpolate the object properties you want Weaviate to pass on to the language model. For example, to pass on the object's `title` property, include `{title}` in the query.

:::code-group{sync="languages"}
```python title="Python" {5-6}
collection = client.collections.use("DemoCollection")

response = collection.generate.near_text(
    query="A holiday film",  # The model provider integration will automatically vectorize the query
    single_prompt="Translate this into French: {title}",
    limit=2
)

for obj in response.objects:
    print(obj.properties["title"])
    print(f"Generated output: {obj.generated}")  # Note that the generated output is per object
```

```typescript title="JavaScript/TypeScript"
let response;
const myCollection = client.collections.use("DemoCollection");
```
:::

### Grouped task

![Grouped task RAG integration generates one output for the set of search results](/assets/docs/weaviate/model-providers/_includes/integration_anyscale_rag_grouped.png)

To generate one text for the entire set of search results, use the grouped task method.

In other words, when you have `n` search results, the generative model generates one output for the entire group.

:::code-group{sync="languages"}
```python title="Python" {5-6}
collection = client.collections.use("DemoCollection")

response = collection.generate.near_text(
    query="A holiday film",  # The model provider integration will automatically vectorize the query
    grouped_task="Write a fun tweet to promote readers to check out these films.",
    limit=2
)

print(f"Generated output: {response.generative.text}")  # Note that the generated output is per query
for obj in response.objects:
    print(obj.properties["title"])
```

```typescript title="JavaScript/TypeScript"
let response;
const myCollection = client.collections.use("DemoCollection");
```
:::

## References

### Available models

- `meta-llama/Llama-2-70b-chat-hf` (default)
- `meta-llama/Llama-2-13b-chat-hf`
- `meta-llama/Llama-2-7b-chat-hf`
- `codellama/CodeLlama-34b-Instruct-hf`
- `mistralai/Mistral-7B-Instruct-v0.1`
- `mistralai/Mixtral-8x7B-Instruct-v0.1`

## Further resources

### Code examples

Once the integrations are configured at the collection, the data management and search operations in Weaviate work identically to any other collection. See the following model-agnostic examples:

- The [How-to: Manage collections](../how-to-manage-collections/index.md) and [How-to: Manage objects](../how-to-manage-objects/index.md) guides show how to perform data operations (i.e. create, read, update, delete collections and objects within them).
- The [How-to: Query & Search](../how-to-query-search/index.md) guides show how to perform search operations (i.e. vector, keyword, hybrid) as well as retrieval augmented generation.

### References

- Anyscale [Endpoints API documentation](https://docs.anyscale.com/endpoints/intro/)

## Questions and feedback

Have a question or feedback? Here's how to reach us.

::::card-grid
:::card{title="Community Forum" href="https://forum.weaviate.io/c/support" icon="messages-square"}
Ask questions and connect with other developers on our **Community forum**.
:::

:::card{title="Support" href="/guides/support-overview" icon="life-buoy"}
Weaviate Cloud user or customer? Find the right channel on the **Support page**.
:::
::::

## Related pages

- [Agents](./agents-index.md)
- [AI-assisted Weaviate code generation](./ai-assisted-vibe-coding-index.md)
- [APIs](./apis-index.md)
- [Authorization and authentication](./authorization-and-authentication-index.md)
- [Benchmarks](./benchmarks-index.md)
- [Best practices](./best-practices-index.md)
- [Client libraries](./clients-index.md)
- [Client Libraries / SDKs](./client-libraries-index.md)
- [Cloud](./cloud-index.md)
- [Cloud account management](./cloud-account-management-index.md)

# Agent Instructions

This portal answers questions programmatically. To receive a synthesized,
source-cited answer instead of crawling page by page, append the `?ask=`
query parameter to any page URL on this site:

    /guides/quickstart?ask=how+do+I+authenticate

Optional parameters:

- `&goal=<what-you-are-trying-to-do>` steers the answer toward your
  objective (e.g. `&goal=write+a+python+client`).
- `&version=<label>` scopes the answer to a mounted version when the
  portal publishes more than one.

The response is `text/markdown`: the answer followed by a `# Sources` list
of the portal pages it was grounded in. Status codes are the contract:

- `200` — the answer; `402` — the portal owner’s plan or answer credits are
  exhausted (surface this to your operator; do NOT retry); `429` — you are
  rate-limited; back off for the `Retry-After` seconds; `503` — the answer
  lane is temporarily unavailable; fall back to crawling the `.md` pages.

For the full corpus map read `llms.txt` at the site root; for the tool
surface (search + page fetch as MCP tools) see `/mcp`.
