Enrichment processors
Three processors that add structure without calling a model.
Concept link
concept_link links text to a controlled vocabulary, so a document is connected to the concepts
it mentions rather than only to the words it contains.
{
"concept_link": {
"field": "content",
"vocabulary": "crop-types",
"target_field": "concepts",
"include_mentions": true
}
}
| Option | Required | Default | Meaning |
|---|---|---|---|
field | yes | — | Text to scan. |
vocabulary | yes | — | Index holding the vocabulary to link against. |
target_field | no | concept | Where linked concepts are written. |
coverage_field | no | linked | Whether this document linked to anything at all. |
include_mentions | no | false | Also record where in the text each concept was found. |
ignore_missing | no | false | Record a document with no field as unlinked rather than failing it. |
coverage_field holds a boolean, and it is written on every document — including the ones that
matched nothing and the ones with no text at all. Writing it only on a match would be cheaper and
would make the one question that matters unanswerable: what fraction of this corpus is linked at
all? A search can then report what it found but never what it could never have found, which is
precisely the failure a confident recall claim invites. Because it is an ordinary field on every
document, coverage is an ordinary aggregation.
include_mentions writes to <target_field>_mentions — by default concept_mentions — giving the
matched text and its offsets, so "why is this document about dogs?" has an answer a person can check
rather than a score they have to trust.
Matching is by whole tokens, so cat does not match inside category, and where two labels
overlap the longer one wins. A name shared by several concepts links to all of them: finding every
candidate honestly is the dictionary's job, and choosing between them is disambiguation.
The vocabulary is an index, and it gets there through the ontology import endpoint. A pipeline may name a vocabulary that does not exist yet — the gazetteer is resolved per document, not when the pipeline is created, so importing and creating the pipeline can happen in either order.
include_mentions defaults to false because mentions are proportional to document length — useful
for highlighting, expensive to store for every document by default.
Chunk
chunk splits the blocks content_extract produced into passages small enough to embed and retrieve.
Retrieval quality is decided here as much as by the model: a chunk that splits mid-argument answers
half a question, and one that spans three topics matches all of them weakly.
{
"chunk": {
"field": "extracted.blocks",
"target_field": "chunks",
"algorithm": "recursive",
"chunk_size": 2000,
"chunk_overlap": 200
}
}
| Option | Default | Meaning |
|---|---|---|
field | extracted.blocks | Blocks to split. |
target_field | chunks | Where the chunks are written. |
algorithm | recursive | recursive, fixed, semantic or topic_shift. |
chunk_size | 2000 | Target size, in characters. |
chunk_overlap | 200 | Characters repeated between neighbours. |
min_chunk_size | 100 | Below this, a fragment is merged rather than emitted. |
block_types | (all) | Restrict to particular block types. |
window_size | 3 | Sentences compared at a time — semantic and topic_shift only. |
overlap_sentences | 0 | Sentences repeated between neighbours — semantic and topic_shift only. |
provider, model_id, dimensions, provider_config | — | The embedding provider, for semantic and topic_shift. |
on_failure_action | fallback | fallback to a non-model algorithm, or fail. |
recursive and fixed need no model. semantic and topic_shift embed sentence windows to find the
seams, so they need a provider — and default to falling back to a structural split rather than
failing the document, because a pipeline that stops ingesting when a model endpoint blips is worse
than a slightly worse chunk boundary.
Embed
embed turns chunks into vectors through a configured provider. The vectors land on the chunks
themselves, so one document carries many, which is what makes passage-level retrieval possible.
{
"embed": {
"field": "chunks",
"provider": "http",
"model_id": "bge-small-en",
"dimensions": 384,
"provider_config": { "endpoint": "https://models.internal:8080/embed" }
}
}
The http provider is shown throughout this page because it needs no vendor SDK and no secret in an
example — an endpoint, a model name, and the same allowlist every other fetch obeys. openai works
the same way but additionally needs an API key, and a key written into a pipeline is readable by
anyone who can read that pipeline. Put it in the keystore and reference it by setting name rather
than pasting it from a page.
| Option | Default | Meaning |
|---|---|---|
field | chunks | Chunks to embed. |
provider | — | http or openai. |
model_id | — | The model the provider should use. |
dimensions | 1536 | Vector width; must match the field mapping. |
content_type | text | text or image. |
batch_size | 50 | Chunks per provider call. |
source_uri_field | — | For image, the field holding the URI to fetch. |
block_types | (all) | Restrict to particular block types. |
provider_config | {} | Passed to the provider unchanged. |
on_failure_action | skip | skip the chunk, or fail the document. |
dimensions must match the mapping of the field the vectors are indexed into. A mismatch is rejected
at index time, not here, so it surfaces a long way from the pipeline that caused it.
An image embed fetches a URI out of the document, which is the same SSRF surface every other fetch is — so it obeys the same allowlist. See Source access.
Topic drift
topic_drift scores how far a piece of analysis has moved from the goal it was given. It is meant for
long-running agent work, where the failure is not a wrong answer but a slow slide into answering a
different question.
{
"topic_drift": {
"field": "analysis",
"target_field": "drift",
"goal": "assess flood risk to the rail corridor",
"provider": "http",
"model_id": "bge-small-en",
"dimensions": 384,
"provider_config": { "endpoint": "https://models.internal:8080/embed" }
}
}
The provider is validated when the pipeline is created, not when the first document arrives — a missing endpoint or a malformed URL is rejected up front rather than becoming a per-document failure nobody sees until ingest is already running.
| Option | Default | Meaning |
|---|---|---|
field | analysis | Text to score. |
target_field | drift | Where the score is written. |
goal | — | The objective to measure against. |
provider, model_id, provider_config | — | The embedding provider. |
dimensions | 768 | Vector width. |
Query embedding
query_embedding is a search request processor: it embeds the query text at search time and
injects the vector into the query before it reaches the shards. Without it, every caller has to embed
their own query and send a raw vector, which means every client needs the model, the credentials and
the same dimensions as the index.
PUT _search/pipeline/semantic
{
"request_processors": [
{
"query_embedding": {
"provider": "http",
"model_id": "bge-small-en",
"dimensions": 384,
"mode": "template",
"query_template": "{\"query\":{\"knn\":{\"chunks.vector\":{\"vector\":${embedding},\"k\":10}}}}",
"provider_config": { "endpoint": "https://models.internal:8080/embed" }
}
}
]
}
| Option | Default | Meaning |
|---|---|---|
provider | — | http or openai. |
model_id | — | The model the provider should use. |
dimensions | 1536 | Vector width; must match the index. |
mode | template | template or path. Any other value is refused. |
query_template | — | Required when mode is template. The query to substitute into. |
embedding_path | — | Required when mode is path. Dot-path where the vector is injected. |
provider_config, reference_config | {} | Passed to the provider and the fetch path unchanged. |
on_failure_action | fail | fail the search, or passthrough and search without a vector. |
In template mode the query is a string, and ${embedding} is replaced with the vector as a JSON
array. ${text} is also substituted, with the query text escaped — useful for a hybrid query that
needs both the vector and the original words. Because the template is a string, its inner quotes are
escaped, which is why the example above looks the way it does.
Which requirement applies is decided when the pipeline is created, not at search time: template
without query_template, or path without embedding_path, is refused up front.
A raw api_key in provider_config is refused here by name. Anyone who can read a search
pipeline could read the key, so the processor rejects it and directs you to api_key_setting, which
names a secure setting in the keystore instead. That is why this page's examples use the http
provider with an endpoint: an endpoint is not a credential.
on_failure_action defaults to fail rather than passthrough, unlike the ingest processors. A
search that silently dropped its vector would return lexical results that look like a working
semantic search — plausible, ranked, and wrong about what they are.
Rerank prepare
rerank_prepare shapes chunks so a downstream reranker sees the context it needs. A reranker scores a
query against a passage, and a passage stripped of its title and position usually scores worse than
the same passage with them.
{
"rerank_prepare": {
"field": "chunks",
"document_title_field": "title",
"document_url_field": "url"
}
}
| Option | Default | Meaning |
|---|---|---|
field | chunks | Chunk array to prepare. |
document_title_field | title | Document title to carry onto each chunk. |
document_url_field | url | Document URL to carry onto each chunk. |
include_position_score | true | Carry a position-derived prior onto each chunk. |
include_position_score is on by default because position is genuinely informative in most document
types — an abstract near the top of a paper is more likely to answer a general query than a chunk
from the middle of an appendix. Turn it off for corpora where position means nothing, such as chat
logs or record exports.
The reranker itself is not in this plugin; it calls a model endpoint and lives in the separate inference artifact. This processor prepares the input for it, which is why it is useful on its own.
Retrieval grounding
retrieval_grounding is a search response processor, not an ingest processor. It goes in a
search pipeline and annotates hits on the way out with where each result came from.
PUT _search/pipeline/grounded
{
"response_processors": [
{ "retrieval_grounding": { "include_spatial_context": true } }
]
}
| Option | Default | Meaning |
|---|---|---|
target_field | _grounding | Where grounding is attached on each hit. |
include_provenance | true | Source document, offsets, and how the hit was produced. |
include_spatial_context | true | Geographic extent, when the hit has one. |
include_chunk_context | true | Neighbouring text around the matching chunk. |
context_snippet_chars | 200 | Size of that surrounding snippet. |
generate_preview_urls | false | Emit a preview URL for an imagery tile. |
preview_uri_field | tile_uri | Field holding the tile URI to build a preview from. |
preview_expiry_seconds | 3600 | Lifetime of a generated preview URL. |
This exists for the answer-generation case. When a model composes an answer from retrieved passages, "where did this come from" has to survive the retrieval step — a citation reconstructed afterwards is a guess. Attaching provenance at response time means the grounding is produced by the thing that actually knows it.
generate_preview_urls is off by default because a preview URL is a credential with a lifetime: it
grants access to a tile for preview_expiry_seconds. Turn it on deliberately, and keep the expiry
short.