Skip to main content
Version: 0.12.0

Embed processor

Introduced 0.11.0

The embed processor generates vector embeddings from text, image, or multimodal content. It reads chunk arrays produced by the chunk or image_tiling processors and writes a dense vector into each chunk, making them searchable via kNN.

Six embedding providers are supported: AWS Bedrock, OpenAI, a generic HTTP endpoint for self-hosted models, GCP Vertex AI, Azure OpenAI, and Azure AI Vision (multimodal).

tip

The dimensions you set here must match both your chosen model's output width and the dimension of the knn_vector field in your index mapping. A mismatch causes indexing to fail. See End-to-end content processing for how the embed processor fits into a full extract → chunk → embed pipeline.

Syntax

{
"embed": {
"field": "chunks",
"model_id": "amazon.titan-embed-text-v2:0",
"provider": "bedrock",
"dimensions": 1024,
"provider_config": {
"region": "us-east-2"
}
}
}

Architecture

                        ┌──────────────────────┐
│ embed processor │
│ │
chunks[] ──────────────►│ For each chunk: │──────────► chunks[] + embedding
[text, image_data] │ 1. Detect content │ [text, embedding: [...]]
│ type (text/img/ │
│ multimodal) │
│ 2. Build input │
│ 3. Call provider │
│ 4. Store vector │
└──────────┬───────────┘

┌──────────────┼──────────────┐
▼ ▼ ▼
┌──────────┐ ┌──────────┐ ┌──────────┐
│ Bedrock │ │ OpenAI │ │ HTTP │
│ Provider │ │ Provider │ │ Provider │
└────┬─────┘ └────┬─────┘ └────┬─────┘
│ │ │
Titan Text text-embedding Your own
Titan Multi -3-small/large model
Cohere Embed ada-002 endpoint

Providers

AWS Bedrock

Calls Bedrock InvokeModel in your AWS account. Data never leaves your VPC.

ModelModel IDDimensionsContent types
Titan Text Embeddings V2amazon.titan-embed-text-v2:0256, 512, 1024Text
Titan Multimodal Embeddings G1amazon.titan-embed-image-v11024 (fixed)Text, image, multimodal
Cohere Embed English v3cohere.embed-english-v31024Text
Cohere Embed Multilingual v3cohere.embed-multilingual-v31024Text
note

Bedrock processes one input per API call (no batch API). For large document sets, consider using Titan Text v2 which is optimized for throughput.

Provider config:

KeyRequired/OptionalDescription
regionRequiredAWS region for the model (e.g., us-east-1, us-west-2). Set per pipeline and independent of the cluster's region and the source bucket's region.
access_keyOptionalAWS access key from keystore. When omitted, credentials resolve through the default AWS chain — including IRSA / STS web-identity (AssumeRoleWithWebIdentity) on EKS, ECS task roles, and EC2 instance profiles — so a workload identity needs no static keys.
secret_keyOptionalAWS secret key from keystore.
session_tokenOptionalSTS session token for temporary credentials.
note

Bring your own model. The model_id is passed straight to Bedrock InvokeModel, so any model your account can invoke works — the Titan and Cohere models above, plus custom or fine-tuned models imported via Bedrock Custom Model Import (referenced by ARN) and Bedrock Marketplace models. The table lists common examples, not an allowlist.

Cross-region and keyless. Because region is per pipeline and STS/IRSA resolves independently of it, you can embed with a model in one region while your data, bucket, and cluster live in others — with no credentials in the pipeline.

OpenAI

Calls the OpenAI Embeddings API. Supports batch processing (up to 100 inputs per call).

ModelModel IDDimensions
text-embedding-3-smalltext-embedding-3-small1536
text-embedding-3-largetext-embedding-3-large3072
text-embedding-ada-002text-embedding-ada-0021536
note

OpenAI provider supports text-only embeddings. For multimodal content, use Bedrock or HTTP.

Provider config:

KeyRequired/OptionalDescription
api_key_settingRequiredKeystore setting name (e.g., ingest.content.openai.api_key). Raw API keys are not allowed in pipeline configuration.
endpointOptionalCustom endpoint for Azure OpenAI or proxies. Default is https://api.openai.com/v1/embeddings.

HTTP (self-hosted)

Calls any HTTP embedding service. Use this for self-hosted models where data sovereignty requires all inference to stay within your network.

Provider config:

KeyRequired/OptionalDescription
endpointRequiredPOST endpoint URL (e.g., http://embedding-svc.internal:8080/embed).
response_embedding_pathOptionalJSONPath to extract embeddings from response. Default is $.embeddings.
auth_headerOptionalAuthorization header value (e.g., Bearer <token>).
max_batch_sizeOptionalMaximum inputs per request to the HTTP endpoint. This is the HTTP provider's own batch control and is separate from the top-level batch_size parameter. Default is 50.

Request format:

{
"inputs": [
{"text": "hello world"},
{"image": "<base64>", "image_mime_type": "image/jpeg"},
{"text": "caption", "image": "<base64>", "image_mime_type": "image/png"}
],
"model": "your-model-id"
}

Expected response:

{
"embeddings": [[0.1, 0.2, ...], [0.3, 0.4, ...]]
}

GCP Vertex AI

Embeds via Google Cloud Vertex AI — the model runs in your GCP project. Set model_id to the Vertex model (for example text-embedding-005, or multimodalembedding@001 for text+image). Prefer Workload Identity / Application Default Credentials (no key files): when gcp_access_token is omitted, the provider uses ADC; when it is set, the provider uses that static token.

Provider config:

KeyRequired/OptionalDescription
gcp_projectRequiredGCP project ID.
gcp_locationRequiredVertex region (for example us-central1).
endpointOptionalBase-URL override (for a private/VPC endpoint).
gcp_access_tokenOptionalA pre-fetched OAuth token — selects static credential mode. Omit it to use ADC / Workload Identity (recommended).

Azure OpenAI

Embeds via an Azure OpenAI deployment in your Azure resource. Set model_id (or deployment) to the deployment name. Prefer Microsoft Entra (managed identity / AKS workload identity): when api_key is omitted, the provider authenticates with Entra; when it is set, the provider uses the API key.

Provider config:

KeyRequired/OptionalDescription
endpointRequiredAzure OpenAI resource base URL (for example https://my-resource.openai.azure.com).
deploymentOptionalThe Azure deployment name (falls back to model_id).
api_versionOptionalAzure REST API version.
api_keyOptionalAzure OpenAI API key — selects key credential mode. Omit it to use Entra (recommended).

Azure AI Vision (multimodal)

Embeds text and images into the same vector space via Azure AI Vision 4.0 (dimension 1024). Credentials work like Azure OpenAI (api_key → key mode; omit → Entra).

Provider config:

KeyRequired/OptionalDescription
endpointRequiredAzure AI Vision resource base URL (for example https://my-resource.cognitiveservices.azure.com).
api_versionOptionalAzure REST API version.
model_versionOptionalAzure AI Vision multimodal model version.
api_keyOptionalAPI key — selects key mode. Omit it to use Entra.

Configuration parameters

ParameterData typeRequired/OptionalDescription
fieldStringOptionalSource field containing the chunk array. Default is chunks.
model_idStringRequiredModel identifier (provider-specific).
providerStringRequiredProvider name: bedrock, openai, http, vertex, azure, or azure_vision.
dimensionsIntegerOptionalEmbedding vector dimensions. Default is 1536. Must be exactly 1024 for Titan Multimodal.
content_typeStringOptionalProvider validation hint: text, image, or multimodal. Actual content type is auto-detected per chunk. Default is text.
batch_sizeIntegerOptionalInputs per provider call. Bedrock ignores this (always 1); OpenAI batches up to 100 per call; the HTTP provider uses its own max_batch_size (see below). Default is 50.
on_failure_actionStringOptionalskip (log warning, continue without embedding -- enables BM25 fallback) or fail (reject document). Default is skip.
block_typesArrayOptionalChunk types to embed. Non-matching chunks are skipped. Default embeds all types.
source_uri_fieldStringOptionalFor direct image embedding: document field containing image URI. Image is fetched transiently and not stored.
max_image_embed_bytesIntegerOptionalMaximum image size for transient fetch. Default is 26214400 (25 MB).
reference_configObjectOptionalReference resolver configuration for source_uri_field.
provider_configObjectOptionalProvider-specific configuration. See provider sections above.
descriptionStringOptionalA brief description of the processor.
tagStringOptionalAn identifier tag for the processor.

Content type detection

The processor auto-detects the content type of each chunk:

Chunk has text + image_data?  →  multimodal embedding
Chunk has text only? → text embedding
Chunk has image_data only? → image embedding
Chunk has neither? → skipped

Inline image data (image_data field in chunks) is limited to 5 MB after base64 decoding. For larger images, use source_uri_field for transient fetch.

Output structure

The processor adds an embedding field to each eligible chunk:

{
"chunks": [
{
"text": "First chunk of text...",
"chunk_index": 0,
"embedding": [0.0123, -0.0456, 0.0789, ...]
},
{
"text": "Second chunk...",
"chunk_index": 1,
"embedding": [0.0234, -0.0567, 0.0891, ...]
}
]
}

When using source_uri_field, the embedding is stored at the document root:

{
"source_uri": "s3://bucket/image.jpg",
"embedding": [0.0123, -0.0456, 0.0789, ...]
}

Security

warning

Raw API keys are not allowed in pipeline configuration. Use the Lucenia keystore to store credentials securely.

# Store an OpenAI API key
bin/lucenia-keystore add ingest.content.openai.api_key

# Store AWS credentials (alternative to instance profile)
bin/lucenia-keystore add ingest.content.bedrock.access_key
bin/lucenia-keystore add ingest.content.bedrock.secret_key

Then reference the keystore setting in your pipeline:

{
"embed": {
"provider_config": {
"api_key_setting": "ingest.content.openai.api_key"
}
}
}

Using the processor

Example 1: Text embeddings with Bedrock Titan

PUT _ingest/pipeline/text-embed
{
"processors": [
{
"embed": {
"field": "chunks",
"model_id": "amazon.titan-embed-text-v2:0",
"provider": "bedrock",
"dimensions": 1024,
"provider_config": {
"region": "us-east-2"
}
}
}
]
}

Example 2: Multimodal embeddings with Titan Multimodal

PUT _ingest/pipeline/multimodal-embed
{
"processors": [
{
"embed": {
"field": "chunks",
"model_id": "amazon.titan-embed-image-v1",
"provider": "bedrock",
"dimensions": 1024,
"content_type": "multimodal",
"provider_config": {
"region": "us-east-1"
}
}
}
]
}

Example 3: Self-hosted model via HTTP

PUT _ingest/pipeline/self-hosted-embed
{
"processors": [
{
"embed": {
"field": "chunks",
"model_id": "sentence-transformers/all-MiniLM-L6-v2",
"provider": "http",
"dimensions": 384,
"provider_config": {
"endpoint": "http://embedding-service.internal:8080/embed",
"max_batch_size": 32
}
}
}
]
}

Example 4: Direct image embedding from S3

Embed an image without storing it in the document -- the image is fetched transiently from S3, embedded, and discarded:

PUT _ingest/pipeline/image-embed
{
"processors": [
{
"embed": {
"source_uri_field": "image_uri",
"model_id": "amazon.titan-embed-image-v1",
"provider": "bedrock",
"dimensions": 1024,
"provider_config": {
"region": "us-east-1"
}
}
}
]
}
PUT /images/_doc/1?pipeline=image-embed
{
"image_uri": "s3://my-bucket/photos/landscape.jpg",
"title": "Mountain landscape"
}