Inference in Qdrant Managed Cloud

Inference is the process of creating vector embeddings from text, images, or other data types using a machine learning model.

Qdrant Managed Cloud allows you to use inference directly in the cloud, without the need to set up and maintain your own inference infrastructure. You can use embedding models hosted on Qdrant Cloud, or use externally hosted models.

Enabling/Disabling Inference

Inference is enabled by default for all new clusters created after July 7, 2025. You can enable it for existing clusters directly from the Inference tab of the Cluster Detail page in the Qdrant Cloud Console. Activating inference triggers a restart of your cluster to apply the new configuration.

Using Inference

Use inference through the Qdrant SDKs and the REST or gRPC APIs when upserting points and when querying the database. Refer to the Inference documentation for details.

Embedding Models

Clusters on Qdrant Managed Cloud can access embedding models that are hosted on Qdrant Cloud and models that are hosted externally by other providers.

Qdrant-Hosted Models

The following models are available:

Dense Models

ModelModalityDimensionsCost
sentence-transformers/all-minilm-l6-v2Text384Free
intfloat/multilingual-e5-smallText384Free
mixedbread-ai/mxbai-embed-large-v1Text1024Paid
qdrant/clip-vit-b-32-textText512Paid
qdrant/clip-vit-b-32-visionImage512Paid

The qdrant/clip-vit-b-32-text and qdrant/clip-vit-b-32-vision models share a vector space, so you can embed images with the vision model and search them with text queries embedded by the text model.

Sparse Models

ModelModalityCost
qdrant/bm25TextFree
prithivida/splade_pp_en_v1TextPaid

Multivector Models

ModelModalityDimensionsCost
answerdotai/answerai-colbert-small-v1Text96Free

Billing for Qdrant-Hosted Models

Usage of paid embedding models is billed based on the number of tokens processed by the model. The cost is calculated per 1,000,000 tokens. The price depends on the model and is displayed on the Inference tab of the Cluster Detail page. You can also see the current usage of each model there.

Free models are also available on free-tier clusters.

External Models

Qdrant Cloud can act as a proxy for the following external embedding providers:

  • OpenAI
  • Cohere
  • Jina AI
  • OpenRouter

This enables you to access any of the embedding models provided by these providers through the Qdrant API.

Billing for External Models

To use an external provider’s embedding model, you need an API key from that provider. Billing is managed directly through the external provider, based on API key usage. Refer to each external embedding model provider’s website for pricing details.

Was this page useful?

Thank you for your feedback! 🙏

We are sorry to hear that. 😔 You can edit this page on GitHub, or create a GitHub issue.