Skip to main content

Enrichment processors

Three processors that add structure without calling a model.

concept_link links text to a controlled vocabulary, so a document is connected to the concepts it mentions rather than only to the words it contains.

{
"concept_link": {
"field": "content",
"vocabulary": "crop-types",
"target_field": "concepts",
"include_mentions": true
}
}
OptionRequiredDefaultMeaning
fieldyes—Text to scan.
vocabularyyes—Index holding the vocabulary to link against.
target_fieldnoconceptWhere linked concepts are written.
coverage_fieldnolinkedWhether this document linked to anything at all.
include_mentionsnofalseAlso record where in the text each concept was found.
ignore_missingnofalseRecord a document with no field as unlinked rather than failing it.

coverage_field holds a boolean, and it is written on every document — including the ones that matched nothing and the ones with no text at all. Writing it only on a match would be cheaper and would make the one question that matters unanswerable: what fraction of this corpus is linked at all? A search can then report what it found but never what it could never have found, which is precisely the failure a confident recall claim invites. Because it is an ordinary field on every document, coverage is an ordinary aggregation.

include_mentions writes to <target_field>_mentions — by default concept_mentions — giving the matched text and its offsets, so "why is this document about dogs?" has an answer a person can check rather than a score they have to trust.

Matching is by whole tokens, so cat does not match inside category, and where two labels overlap the longer one wins. A name shared by several concepts links to all of them: finding every candidate honestly is the dictionary's job, and choosing between them is disambiguation.

The vocabulary is an index, and it gets there through the ontology import endpoint. A pipeline may name a vocabulary that does not exist yet — the gazetteer is resolved per document, not when the pipeline is created, so importing and creating the pipeline can happen in either order.

include_mentions defaults to false because mentions are proportional to document length — useful for highlighting, expensive to store for every document by default.

Chunk​

chunk splits the blocks content_extract produced into passages small enough to embed and retrieve. Retrieval quality is decided here as much as by the model: a chunk that splits mid-argument answers half a question, and one that spans three topics matches all of them weakly.

{
"chunk": {
"field": "extracted.blocks",
"target_field": "chunks",
"algorithm": "recursive",
"chunk_size": 2000,
"chunk_overlap": 200
}
}
OptionDefaultMeaning
fieldextracted.blocksBlocks to split.
target_fieldchunksWhere the chunks are written.
algorithmrecursiverecursive, fixed, semantic or topic_shift.
chunk_size2000Target size, in characters.
chunk_overlap200Characters repeated between neighbours.
min_chunk_size100Below this, a fragment is merged rather than emitted.
block_types(all)Restrict to particular block types.
window_size3Sentences compared at a time — semantic and topic_shift only.
overlap_sentences0Sentences repeated between neighbours — semantic and topic_shift only.
provider, model_id, dimensions, provider_config—The embedding provider, for semantic and topic_shift.
on_failure_actionfallbackfallback to a non-model algorithm, or fail.

recursive and fixed need no model. semantic and topic_shift embed sentence windows to find the seams, so they need a provider — and default to falling back to a structural split rather than failing the document, because a pipeline that stops ingesting when a model endpoint blips is worse than a slightly worse chunk boundary.

Embed​

embed turns chunks into vectors through a configured provider. The vectors land on the chunks themselves, so one document carries many, which is what makes passage-level retrieval possible.

{
"embed": {
"field": "chunks",
"provider": "http",
"model_id": "bge-small-en",
"dimensions": 384,
"provider_config": { "endpoint": "https://models.internal:8080/embed" }
}
}

The http provider is shown throughout this page because it needs no vendor SDK and no secret in an example — an endpoint, a model name, and the same allowlist every other fetch obeys. openai works the same way but additionally needs an API key, and a key written into a pipeline is readable by anyone who can read that pipeline. Put it in the keystore and reference it by setting name rather than pasting it from a page.

OptionDefaultMeaning
fieldchunksChunks to embed.
provider—http or openai.
model_id—The model the provider should use.
dimensions1536Vector width; must match the field mapping.
content_typetexttext or image.
batch_size50Chunks per provider call.
source_uri_field—For image, the field holding the URI to fetch.
block_types(all)Restrict to particular block types.
provider_config{}Passed to the provider unchanged.
on_failure_actionskipskip the chunk, or fail the document.

dimensions must match the mapping of the field the vectors are indexed into. A mismatch is rejected at index time, not here, so it surfaces a long way from the pipeline that caused it.

An image embed fetches a URI out of the document, which is the same SSRF surface every other fetch is — so it obeys the same allowlist. See Source access.

Topic drift​

topic_drift scores how far a piece of analysis has moved from the goal it was given. It is meant for long-running agent work, where the failure is not a wrong answer but a slow slide into answering a different question.

{
"topic_drift": {
"field": "analysis",
"target_field": "drift",
"goal": "assess flood risk to the rail corridor",
"provider": "http",
"model_id": "bge-small-en",
"dimensions": 384,
"provider_config": { "endpoint": "https://models.internal:8080/embed" }
}
}

The provider is validated when the pipeline is created, not when the first document arrives — a missing endpoint or a malformed URL is rejected up front rather than becoming a per-document failure nobody sees until ingest is already running.

OptionDefaultMeaning
fieldanalysisText to score.
target_fielddriftWhere the score is written.
goal—The objective to measure against.
provider, model_id, provider_config—The embedding provider.
dimensions768Vector width.

Query embedding​

query_embedding is a search request processor: it embeds the query text at search time and injects the vector into the query before it reaches the shards. Without it, every caller has to embed their own query and send a raw vector, which means every client needs the model, the credentials and the same dimensions as the index.

PUT _search/pipeline/semantic
{
"request_processors": [
{
"query_embedding": {
"provider": "http",
"model_id": "bge-small-en",
"dimensions": 384,
"mode": "template",
"query_template": "{\"query\":{\"knn\":{\"chunks.vector\":{\"vector\":${embedding},\"k\":10}}}}",
"provider_config": { "endpoint": "https://models.internal:8080/embed" }
}
}
]
}
OptionDefaultMeaning
provider—http or openai.
model_id—The model the provider should use.
dimensions1536Vector width; must match the index.
modetemplatetemplate or path. Any other value is refused.
query_template—Required when mode is template. The query to substitute into.
embedding_path—Required when mode is path. Dot-path where the vector is injected.
provider_config, reference_config{}Passed to the provider and the fetch path unchanged.
on_failure_actionfailfail the search, or passthrough and search without a vector.

In template mode the query is a string, and ${embedding} is replaced with the vector as a JSON array. ${text} is also substituted, with the query text escaped — useful for a hybrid query that needs both the vector and the original words. Because the template is a string, its inner quotes are escaped, which is why the example above looks the way it does.

Which requirement applies is decided when the pipeline is created, not at search time: template without query_template, or path without embedding_path, is refused up front.

A raw api_key in provider_config is refused here by name. Anyone who can read a search pipeline could read the key, so the processor rejects it and directs you to api_key_setting, which names a secure setting in the keystore instead. That is why this page's examples use the http provider with an endpoint: an endpoint is not a credential.

on_failure_action defaults to fail rather than passthrough, unlike the ingest processors. A search that silently dropped its vector would return lexical results that look like a working semantic search — plausible, ranked, and wrong about what they are.

Rerank prepare​

rerank_prepare shapes chunks so a downstream reranker sees the context it needs. A reranker scores a query against a passage, and a passage stripped of its title and position usually scores worse than the same passage with them.

{
"rerank_prepare": {
"field": "chunks",
"document_title_field": "title",
"document_url_field": "url"
}
}
OptionDefaultMeaning
fieldchunksChunk array to prepare.
document_title_fieldtitleDocument title to carry onto each chunk.
document_url_fieldurlDocument URL to carry onto each chunk.
include_position_scoretrueCarry a position-derived prior onto each chunk.

include_position_score is on by default because position is genuinely informative in most document types — an abstract near the top of a paper is more likely to answer a general query than a chunk from the middle of an appendix. Turn it off for corpora where position means nothing, such as chat logs or record exports.

The reranker itself is not in this plugin; it calls a model endpoint and lives in the separate inference artifact. This processor prepares the input for it, which is why it is useful on its own.

Retrieval grounding​

retrieval_grounding is a search response processor, not an ingest processor. It goes in a search pipeline and annotates hits on the way out with where each result came from.

PUT _search/pipeline/grounded
{
"response_processors": [
{ "retrieval_grounding": { "include_spatial_context": true } }
]
}
OptionDefaultMeaning
target_field_groundingWhere grounding is attached on each hit.
include_provenancetrueSource document, offsets, and how the hit was produced.
include_spatial_contexttrueGeographic extent, when the hit has one.
include_chunk_contexttrueNeighbouring text around the matching chunk.
context_snippet_chars200Size of that surrounding snippet.
generate_preview_urlsfalseEmit a preview URL for an imagery tile.
preview_uri_fieldtile_uriField holding the tile URI to build a preview from.
preview_expiry_seconds3600Lifetime of a generated preview URL.

This exists for the answer-generation case. When a model composes an answer from retrieved passages, "where did this come from" has to survive the retrieval step — a citation reconstructed afterwards is a guess. Attaching provenance at response time means the grounding is produced by the thing that actually knows it.

generate_preview_urls is off by default because a preview URL is a credential with a lifetime: it grants access to a tile for preview_expiry_seconds. Turn it on deliberately, and keep the expiry short.