:::callout{intent="info" title="Looking for Azure OpenAI integration docs?"}
For Azure OpenAI integration docs, see [this page instead](openai-azure-generative.md).
:::

Weaviate's integration with OpenAI's APIs allows you to access their models' capabilities directly from Weaviate.

[Configure a Weaviate collection](#configure-collection) to use a generative AI model with OpenAI. Weaviate will perform retrieval augmented generation (RAG) using the specified model and your OpenAI API key.

More specifically, Weaviate will perform a search, retrieve the most relevant objects, and then pass them to the OpenAI generative model to generate outputs.

![RAG integration illustration](/assets/docs/weaviate/model-providers/_includes/integration_openai_rag.png)

## Requirements

### Weaviate configuration

Your Weaviate instance must be configured with the OpenAI generative AI integration (`generative-openai`) module.

:::accordion{title="For Weaviate Cloud (WCD) users"}
This integration is enabled by default on Weaviate Cloud (WCD) instances.
:::

:::accordion{title="For self-hosted users"}
- Check the [cluster metadata](../monitoring-and-logging/status.md#cluster-metadata) to verify if the module is enabled.
- Follow the [how-to configure modules](../how-to-configure-weaviate/modules.md) guide to enable the module in Weaviate.
:::

<!-- Docs note: the `OPENAI_ORGANIZATION` environment variable is not documented, as it is not the recommended way to provide the OpenAI organization parameter. -->

### API credentials

You must provide a valid OpenAI API key to Weaviate for this integration. Go to [OpenAI](https://openai.com/) to sign up and obtain an API key.

Provide the API key to Weaviate using one of the following methods:

- Set the `OPENAI_APIKEY` environment variable that is available to Weaviate.
- Provide the API key at runtime, as shown in the examples below.

:::code-group{sync="languages"}
```python title="Python"
# Recommended: save sensitive data as environment variables
openai_key = os.getenv("OPENAI_API_KEY")
```

```typescript title="JavaScript/TypeScript"
const openaiApiKey = process.env.OPENAI_API_KEY || '';  // Replace with your inference API key
```
:::

## Configure collection

:::callout{intent="info" title="Generative model integration mutability"}
A collection's `generative` model integration configuration is mutable from `v1.25.23`, `v1.26.8` and `v1.27.1`. See [this section](../how-to-manage-collections/generative-reranker-models.md#update-the-generative-model-integration) for details on how to update the collection configuration.
:::

[Configure a Weaviate index](../how-to-manage-collections/generative-reranker-models.md#specify-a-generative-model-integration) as follows to use an OpenAI generative AI model:

:::code-group{sync="languages"}
```python title="Python" {5}
from weaviate.classes.config import Configure

client.collections.create(
    "DemoCollection",
    generative_config=Configure.Generative.openai()
    # Additional parameters not shown
)
```

```typescript title="JavaScript/TypeScript" {3}
await client.collections.create({
  name: 'DemoCollection',
  generative: weaviate.configure.generative.openAI(),
  // Additional parameters not shown
});
```
:::

### Select a model

You can specify one of the [available models](#available-models) for Weaviate to use, as shown in the following configuration example:

:::code-group{sync="languages"}
```python title="Python" {5-7}
from weaviate.classes.config import Configure

client.collections.create(
    "DemoCollection",
    generative_config=Configure.Generative.openai(
        model="gpt-4-1106-preview"
    )
    # Additional parameters not shown
)
```

```typescript title="JavaScript/TypeScript" {3-5}
await client.collections.create({
  name: 'DemoCollection',
  generative: weaviate.configure.generative.openAI({
    model: 'gpt-4-1106-preview',
  }),
  // Additional parameters not shown
});
```
:::

You can [specify](#generative-parameters) one of the [available models](#available-models) for Weaviate to use. The [default model](#available-models) is used if no model is specified.

### Generative parameters

Configure the following generative parameters to customize the model behavior.

:::code-group{sync="languages"}
```python title="Python" {5-17}
from weaviate.classes.config import Configure

client.collections.create(
    "DemoCollection",
    generative_config=Configure.Generative.openai(
        # # These parameters are optional
        # model="gpt-4",
        # frequency_penalty=0,
        # max_tokens=500,
        # presence_penalty=0,
        # temperature=0.7,
        # top_p=0.7,
        # base_url="<custom_openai_url>",
        # # For reasoning models such as the gpt-5 family:
        # reasoning_effort="medium",  # One of "minimal", "low", "medium", "high"
        # verbosity="medium",  # One of "low", "medium", "high"
    )
    # Additional parameters not shown
)
```

```typescript title="JavaScript/TypeScript" {3-11}
await client.collections.create({
  name: 'DemoCollection',
  generative: weaviate.configure.generative.openAI({
    // These parameters are optional
    // model: 'gpt-4',
    // frequencyPenalty: 0,
    // maxTokens: 500,
    // presencePenalty: 0,
    // temperature: 0.7,
    // topP: 0.7,
  }),
  // Additional parameters not shown
});
```
:::

Two additional parameters are available for reasoning models such as the `gpt-5` family. They were added in `v1.33.0`, and backported to `v1.31.15` and `v1.32.9`:

- `reasoningEffort`: How much reasoning the model does before it answers. One of `minimal`, `low`, `medium`, or `high`. If not set, the model provider default applies.
- `verbosity`: How detailed the generated answer is. One of `low`, `medium`, or `high`. If not set, the model provider default applies.

The Python client exposes these as the `reasoning_effort` and `verbosity` arguments, both when you configure the collection and when you [select a model at runtime](#select-a-model-at-runtime). The TypeScript client does not expose them yet, so set them with another client or through the REST collection configuration API.

For further details on model parameters, see the [OpenAI API documentation](https://platform.openai.com/docs/api-reference/chat).

## Select a model at runtime

Aside from setting the default model provider when creating the collection, you can also override it at query time.

:::code-group{sync="languages"}
```python title="Python" {9-19}
from weaviate.classes.config import Configure
from weaviate.classes.generate import GenerativeConfig

collection = client.collections.use("DemoCollection")
response = collection.generate.near_text(
    query="A holiday film",
    limit=2,
    grouped_task="Write a tweet promoting these two movies",
    generative_provider=GenerativeConfig.openai(
        # # These parameters are optional
        # model="gpt-4",
        # frequency_penalty=0,
        # max_tokens=500,
        # presence_penalty=0,
        # temperature=0.7,
        # top_p=0.7,
        # base_url="<custom_openai_url>"
    ),
    # Additional parameters not shown
)
```

```typescript title="JavaScript/TypeScript"
import { generativeParameters } from 'weaviate-client';
```
:::

## Header parameters

You can provide the API key as well as some optional parameters at runtime through additional headers in the request. The following headers are available:

- `X-OpenAI-Api-Key`: The OpenAI API key.
- `X-OpenAI-Baseurl`: The base URL to use (e.g. a proxy) instead of the default OpenAI URL.
- `X-OpenAI-Organization`: The OpenAI organization ID.

Any additional headers provided at runtime will override the existing Weaviate configuration.

Provide the headers as shown in the [API credentials examples](#api-credentials) above.

:::callout{intent="note" title="How Weaviate builds the request URL"}
Use the `X-OpenAI-Baseurl` header, or the `baseURL` parameter, to target an OpenAI-compatible service. Weaviate builds the request URL by appending `/v1/chat/completions` to the base URL, or `/v1/completions` for the legacy `text-davinci-002` and `text-davinci-003` models.

The generative integration has no `endpoint` parameter to override this path, unlike the [OpenAI embeddings integration](openai-embeddings.md#header-parameters). If your provider expects a different path, point the base URL at a proxy that rewrites `/v1/chat/completions` to your provider's path.

The [Azure OpenAI integration](openai-azure-generative.md) builds a deployment-specific path and ignores the paths above.
:::

## Retrieval augmented generation

After configuring the generative AI integration, perform RAG operations, either with the [single prompt](#single-prompt) or [grouped task](#grouped-task) method.

### Single prompt

![Single prompt RAG integration generates individual outputs per search result](/assets/docs/weaviate/model-providers/_includes/integration_openai_rag_single.png)

To generate text for each object in the search results, use the single prompt method.

The example below generates outputs for each of the `n` search results, where `n` is specified by the `limit` parameter.

When creating a single prompt query, use braces `{}` to interpolate the object properties you want Weaviate to pass on to the language model. For example, to pass on the object's `title` property, include `{title}` in the query.

:::code-group{sync="languages"}
```python title="Python" {5-6}
collection = client.collections.use("DemoCollection")

response = collection.generate.near_text(
    query="A holiday film",  # The model provider integration will automatically vectorize the query
    single_prompt="Translate this into French: {title}",
    limit=2
)

for obj in response.objects:
    print(obj.properties["title"])
    print(f"Generated output: {obj.generated}")  # Note that the generated output is per object
```

```typescript title="JavaScript/TypeScript"
let response;
const myCollection = client.collections.use("DemoCollection");
```
:::

### Grouped task

![Grouped task RAG integration generates one output for the set of search results](/assets/docs/weaviate/model-providers/_includes/integration_openai_rag_grouped.png)

To generate one text for the entire set of search results, use the grouped task method.

In other words, when you have `n` search results, the generative model generates one output for the entire group.

:::code-group{sync="languages"}
```python title="Python" {5-6}
collection = client.collections.use("DemoCollection")

response = collection.generate.near_text(
    query="A holiday film",  # The model provider integration will automatically vectorize the query
    grouped_task="Write a fun tweet to promote readers to check out these films.",
    limit=2
)

print(f"Generated output: {response.generative.text}")  # Note that the generated output is per query
for obj in response.objects:
    print(obj.properties["title"])
```

```typescript title="JavaScript/TypeScript"
let response;
const myCollection = client.collections.use("DemoCollection");
```
:::

### RAG with images

You can also supply images as a part of the input when performing retrieval augmented generation in both single prompts and grouped tasks.

:::code-group{sync="languages"}
```python title="Python" {9-11,18}
import base64
import requests
from weaviate.classes.generate import GenerativeConfig, GenerativeParameters

src_img_path = "https://upload.wikimedia.org/wikipedia/commons/thumb/b/b0/Winter_forest_silver.jpg/960px-Winter_forest_silver.jpg"
base64_image = base64.b64encode(requests.get(src_img_path).content).decode('utf-8')

prompt = GenerativeParameters.grouped_task(
    prompt="Which movie is closest to the image in terms of atmosphere",
    images=[base64_image],      # A list of base64 encoded strings of the image bytes
    # image_properties=["img"], # Properties containing images in Weaviate
)

jeopardy = client.collections.use("DemoCollection")
response = jeopardy.generate.near_text(
    query="Movies",
    limit=5,
    grouped_task=prompt,
    generative_provider=GenerativeConfig.openai(
        max_tokens=1000
    ),
)

# Print the source property and the generated response
for o in response.objects:
    print(f"Title property: {o.properties['title']}")
print(f"Grouped task result: {response.generative.text}")
```

```typescript title="JavaScript/TypeScript"
import { generativeParameters } from 'weaviate-client';
```
:::

## References

### Available models

Weaviate does not validate the model name, so you can set any model that your OpenAI account can reach. Name validation was removed in `v1.33.0`, and backported to `v1.31.17` and `v1.32.10`.

The server default is `gpt-5-mini`. It changed in `v1.32.3`, and was backported to `v1.30.16` and `v1.31.10`. Earlier releases on each of those lines default to `gpt-3.5-turbo`.

See the [OpenAI model documentation](https://platform.openai.com/docs/models) for the list of available models.

Weaviate stores a token limit for the models below. The limit caps the `maxTokens` value you can set for those models; it does not restrict which models you can use.

:::accordion{title="Models with a stored token limit"}
- [gpt-5](https://platform.openai.com/docs/models/gpt-5)
- [gpt-5-mini](https://platform.openai.com/docs/models/gpt-5-mini) (server default)
- [gpt-5-nano](https://platform.openai.com/docs/models/gpt-5-nano)
- [gpt-3.5-turbo](https://platform.openai.com/docs/models/gpt-3-5) (previous server default)
- [gpt-3.5-turbo-16k](https://platform.openai.com/docs/models/gpt-3-5)
- [gpt-3.5-turbo-1106](https://platform.openai.com/docs/models/gpt-3-5)
- [gpt-4](https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turbo)
- [gpt-4-1106-preview](https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turbo)
- [gpt-4-32k](https://platform.openai.com/docs/models/gpt-4-and-gpt-4-turbo)
- [gpt-4o](https://platform.openai.com/docs/models#gpt-4o)
- [gpt-4o-mini](https://platform.openai.com/docs/models#gpt-4o-mini)

These older models also have a stored limit, but are not recommended:

- [davinci 002](https://platform.openai.com/docs/models/overview)
- [davinci 003](https://platform.openai.com/docs/models/overview)
:::

## Further resources

### Other integrations

- [OpenAI embedding models + Weaviate](openai-embeddings.md).

### Code examples

Once the integrations are configured at the collection, the data management and search operations in Weaviate work identically to any other collection. See the following model-agnostic examples:

- The [How-to: Manage collections](../how-to-manage-collections/index.md) and [How-to: Manage objects](../how-to-manage-objects/index.md) guides show how to perform data operations (i.e. create, read, update, delete collections and objects within them).
- The [How-to: Query & Search](../how-to-query-search/index.md) guides show how to perform search operations (i.e. vector, keyword, hybrid) as well as retrieval augmented generation.

### References

- OpenAI [Chat API documentation](https://platform.openai.com/docs/api-reference/chat)

## Questions and feedback

Have a question or feedback? Here's how to reach us.

::::card-grid
:::card{title="Community Forum" href="https://forum.weaviate.io/c/support" icon="messages-square"}
Ask questions and connect with other developers on our **Community forum**.
:::

:::card{title="Support" href="/guides/support-overview" icon="life-buoy"}
Weaviate Cloud user or customer? Find the right channel on the **Support page**.
:::
::::

## Related pages

- [Agents](./agents-index.md)
- [AI-assisted Weaviate code generation](./ai-assisted-vibe-coding-index.md)
- [APIs](./apis-index.md)
- [Authorization and authentication](./authorization-and-authentication-index.md)
- [Benchmarks](./benchmarks-index.md)
- [Best practices](./best-practices-index.md)
- [Client libraries](./clients-index.md)
- [Client Libraries / SDKs](./client-libraries-index.md)
- [Cloud](./cloud-index.md)
- [Cloud account management](./cloud-account-management-index.md)

# Agent Instructions

This portal answers questions programmatically. To receive a synthesized,
source-cited answer instead of crawling page by page, append the `?ask=`
query parameter to any page URL on this site:

    /guides/quickstart?ask=how+do+I+authenticate

Optional parameters:

- `&goal=<what-you-are-trying-to-do>` steers the answer toward your
  objective (e.g. `&goal=write+a+python+client`).
- `&version=<label>` scopes the answer to a mounted version when the
  portal publishes more than one.

The response is `text/markdown`: the answer followed by a `# Sources` list
of the portal pages it was grounded in. Status codes are the contract:

- `200` — the answer; `402` — the portal owner’s plan or answer credits are
  exhausted (surface this to your operator; do NOT retry); `429` — you are
  rate-limited; back off for the `Retry-After` seconds; `503` — the answer
  lane is temporarily unavailable; fall back to crawling the `.md` pages.

For the full corpus map read `llms.txt` at the site root; for the tool
surface (search + page fetch as MCP tools) see `/mcp`.
