# NVIDIA + Weaviate

<!-- Note: for images, use https://docs.google.com/presentation/d/15opIcJuaIjEEcs_1Zm8B6pccox2p7_MHSjCnRv4dPfU/edit?usp=sharing -->

NVIDIA NIM microservices offer a wide range of models for natural language processing and generation. Weaviate seamlessly integrates with NVIDIA, allowing users to leverage the inference engine within the Weaviate Database.

These integrations empower developers to build sophisticated AI-driven applications with ease.

## Integrations with NVIDIA

### Embedding models for vector search

![Embedding integration illustration](/assets/docs/weaviate/model-providers/_includes/integration_nvidia_embedding.png)

NVIDIA's embedding models transform text data into high-dimensional vector representations, capturing meaning and context.

[Weaviate integrates with NVIDIA's embedding models](nvidia-embeddings.md) to enable seamless vectorization of data. This integration allows users to perform semantic and hybrid search operations without the need for additional preprocessing or data transformation steps.

[NVIDIA embedding integration page](nvidia-embeddings.md)
[NVIDIA multimodal embedding integration page](nvidia-embeddings-multimodal.md)

### Generative AI models for RAG

![Single prompt RAG integration generates individual outputs per search result](/assets/docs/weaviate/model-providers/_includes/integration_nvidia_rag_single.png)

Generative AI models on NVIDIA can generate human-like text based on given prompts and contexts.

[Weaviate's generative AI integration](nvidia-generative.md) enables users to perform Retrieval Augmented Generation (RAG) directly from the Weaviate Database. This combines Weaviate's efficient storage and fast retrieval capabilities with generative AI models on NVIDIA to generate personalized and context-aware responses.

[NVIDIA generative AI integration page](nvidia-generative.md)

### Reranker models

![Reranker integration illustration](/assets/docs/weaviate/model-providers/_includes/integration_nvidia_reranker.png)

NVIDIA's reranker models are designed to improve the relevance and ranking of search results.

[The Weaviate reranker integration](nvidia-reranker.md) allows users to easily refine their search results by leveraging NVIDIA's reranker models.

[NVIDIA reranker integration page](nvidia-reranker.md)

## Summary

This integration enables developers to harness the power of NVIDIA's inference engine within Weaviate.

In turn, it simplifies the process of building AI-driven applications to speed up your development process, so that you can focus on creating innovative solutions.

## Get started

You must provide a valid NVIDIA API key to Weaviate for this integration. Go to [NVIDIA](https://build.nvidia.com/) to sign up and obtain an API key.

Then, go to the relevant integration page to learn how to configure Weaviate with the NVIDIA models and start using them in your applications.

- [Text Embeddings](nvidia-embeddings.md)
- [Multimodal Embeddings](nvidia-embeddings-multimodal.md)
- [Generative AI](nvidia-generative.md)

<!-- TODO - Add link back to reranker.md; removed because the link checker still sees it as a broken link -->

<!-- - [Reranker](nvidia.md) -->

## Questions and feedback

Have a question or feedback? Here's how to reach us.

::::card-grid
:::card{title="Community Forum" href="https://forum.weaviate.io/c/support" icon="messages-square"}
Ask questions and connect with other developers on our **Community forum**.
:::

:::card{title="Support" href="/guides/support-overview" icon="life-buoy"}
Weaviate Cloud user or customer? Find the right channel on the **Support page**.
:::
::::

## Related pages

- [Agents](./agents-index.md)
- [AI-assisted Weaviate code generation](./ai-assisted-vibe-coding-index.md)
- [APIs](./apis-index.md)
- [Authorization and authentication](./authorization-and-authentication-index.md)
- [Benchmarks](./benchmarks-index.md)
- [Best practices](./best-practices-index.md)
- [Client libraries](./clients-index.md)
- [Client Libraries / SDKs](./client-libraries-index.md)
- [Cloud](./cloud-index.md)
- [Cloud account management](./cloud-account-management-index.md)

# Agent Instructions

This portal answers questions programmatically. To receive a synthesized,
source-cited answer instead of crawling page by page, append the `?ask=`
query parameter to any page URL on this site:

    /guides/quickstart?ask=how+do+I+authenticate

Optional parameters:

- `&goal=<what-you-are-trying-to-do>` steers the answer toward your
  objective (e.g. `&goal=write+a+python+client`).
- `&version=<label>` scopes the answer to a mounted version when the
  portal publishes more than one.

The response is `text/markdown`: the answer followed by a `# Sources` list
of the portal pages it was grounded in. Status codes are the contract:

- `200` — the answer; `402` — the portal owner’s plan or answer credits are
  exhausted (surface this to your operator; do NOT retry); `429` — you are
  rate-limited; back off for the `Retry-After` seconds; `503` — the answer
  lane is temporarily unavailable; fall back to crawling the `.md` pages.

For the full corpus map read `llms.txt` at the site root; for the tool
surface (search + page fetch as MCP tools) see `/mcp`.
