Graph
A graph_edge field stores a directed graph in an ordinary index, and the graph_traversal query walks
it: every edge within some number of hops of a node, or every document whose node the walk reaches.
It works on any index. The graph is not part of the accelerator's codec, so it needs neither
index.lucenia.accelerator.enabled nor a license.
Map a graph
{
"settings": { "number_of_shards": 1 },
"mappings": {
"properties": {
"road": { "type": "graph_edge" },
"node": { "type": "long" }
}
}
}
The index must have one shard. A traversal can only see one shard's segments, so a graph split across shards would silently return partial answers. A multi-shard index is refused when the field is mapped, which is the last moment the fix is cheap.
Index edges
One document is one edge:
{ "road": { "source": 12345, "target": 67890 } }
- Both
sourceandtargetare required. A half edge would index cleanly and never appear in a traversal, so it is rejected. - One edge per document. An array of edges is rejected; index each edge as its own document.
- Any other field on the document — a weight, a name, a timestamp — is an ordinary field, updated without touching the graph.
Deleting, updating, merging, replicating and snapshotting edges are the index's normal behaviour.
Traverse
{
"query": {
"graph_traversal": { "field": "road", "source": 12345, "hops": 2 }
}
}
Without node_field, the edge documents come back. hops counts edges: one hop matches the edges
leaving source, two hops adds the edges leaving their targets.
With node_field, the documents whose node the walk reached come back instead. This is how a
corpus is filtered by a graph — for example, every document about a concept or anything beneath it:
{
"query": {
"graph_traversal": {
"field": "broader",
"source": 7160,
"node_field": "concept_id",
"direction": "reverse"
}
}
}
| Parameter | Required | Default | Meaning |
|---|---|---|---|
field | yes | — | The graph_edge field to walk. An unmapped field is named in the error. |
source | yes | — | The start node. Never defaults: zero is a real node id. |
hops | no | unbounded | How many edges to follow. The walk is bounded by max_nodes, not depth. |
node_field | no | — | Match documents whose node, held in this field, was reached. |
max_nodes | no | 65536 | How many nodes the walk may reach. Past this the query fails rather than returning a partial answer. |
direction | no | forward | forward follows edges as written; reverse follows them backwards. reverse requires node_field. |
boost, _name | no | — | As on any query. |
Unknown parameters are refused. The query is also refused when search.allow_expensive_queries is
false, because its cost is the size of the neighbourhood and cannot be bounded in advance.
Which direction
If edges are written from the specific to the general — dog → mammal → animal — then forward from
dog reaches its ancestors, and reverse from animal reaches everything beneath it. Neither answer
contains the other, and a walk in the wrong direction returns a plausible list of the wrong documents,
which is why the direction is never inferred.
How it is stored
Each edge is two numeric doc values, <field>.src and <field>.tgt. The first query that needs a
segment's adjacency derives it from those and caches it for as long as the segment lives, so a refresh
costs a segment's worth of work, not the whole graph's.
The derived adjacency is written to a scratch directory on each node, <data path>/lucenia-graph/, and
not into the index. It is not part of a commit, so snapshots and replicas never carry it and every
node derives its own. Two things follow for operators:
- The first traversal after a refresh or merge pays to derive the new segments.
- The scratch directory uses disk that does not appear in index size statistics. It is released as segments close and when an index is deleted, and emptied when a node starts.
Concept similarity
concept_similarity ranks documents by how alike their concept is to one you name. It is a scoring
clause, not a filter: it composes inside a bool alongside BM25 and vectors and contributes a
score.
{
"query": {
"concept_similarity": {
"field": "subclass_of",
"source": 278,
"node_field": "concept",
"measure": "wu_palmer",
"min_similarity": 0.4
}
}
}
| Option | Required | Default | Meaning |
|---|---|---|---|
field | yes | — | The graph_edge field describing the vocabulary. |
source | yes | — | The concept everything is scored against. |
node_field | yes | — | The field carrying each searched document's concept. |
measure | no | wu_palmer | wu_palmer or path. |
min_similarity | no | 0.0 | Score below which a concept is neither returned nor searched beneath. |
max_nodes | no | 65536 | How many concepts one scoring pass may consider. |
source is required even though it is a number, because a long defaults to zero and zero is a
perfectly plausible concept id — an omitted source would otherwise score against something you never
named. node_field must be an integral type with doc values.
min_similarity is not merely a filter. The search stops climbing once nothing beneath the current
ancestor could reach it, so raising it makes the query cheaper as well as narrower.
Why this is affordable
It looks like per-document graph work and is not. The vocabulary is scored against the source once, when the query is built, and the per-document query holds a sorted map from concept to score. Each document then costs a numeric doc-value read and a binary search — about what a range query costs.
Only structural measures
wu_palmer and path need nothing but the hierarchy's shape, and are computed from the shard's own
adjacency. The three corpus-frequency measures — Resnik, Lin and Jiang-Conrath — are refused with an
error, because the frequencies have to be recorded by a cluster action and a plugin installable on
managed OpenSearch cannot add one. graph_index is refused for the same reason: keep the vocabulary
and the documents it ranks in the same single-shard index.
What it does not return
Documents whose concept shares no ancestor with the source, at any threshold. They are not slightly similar, they are unrelated, and a vocabulary with disjoint roots is mostly made of them.
The query needs search.allow_expensive_queries to be enabled. If a pass hits max_nodes before
finishing it fails rather than returning a partial ranking — raise max_nodes or raise
min_similarity.
Ranking, honestly. Concept expansion is measured to recover a complete concept set where lexical and embedding retrieval do not. Adding
concept_similarityto an already-hybrid ranking did not measurably improve nDCG@10, and was reliably unhelpful on common terminology. Use it for recall of a set, not as a general relevance booster.
Not yet exposed
Contraction-hierarchy routing — NetworkDistanceQuery, for road-network shortest-path distance —
is implemented in the graph package but has no query exposed.
Walking a graph held in another index (graph_index) is refused. It needs a cluster action, which a
plugin installable on managed OpenSearch cannot add. Keep the graph and the documents it filters in the
same single-shard index.
Walking a graph held in another index (graph_index) is refused. It needs a cluster action, which a
plugin installable on managed OpenSearch cannot add. Keep the graph and the documents it filters in the
same single-shard index.
On Amazon OpenSearch Service
The graph_edge field and the graph_traversal query use only the mapper and search plugin extensions,
which Amazon OpenSearch Service accepts for custom plugins. The accelerator as a whole does not install
there, because its codec is an engine plugin; see Compatibility.