Skip to main content

Graph

A graph_edge field stores a directed graph in an ordinary index, and the graph_traversal query walks it: every edge within some number of hops of a node, or every document whose node the walk reaches.

It works on any index. The graph is not part of the accelerator's codec, so it needs neither index.lucenia.accelerator.enabled nor a license.

Map a graph​

{
"settings": { "number_of_shards": 1 },
"mappings": {
"properties": {
"road": { "type": "graph_edge" },
"node": { "type": "long" }
}
}
}

The index must have one shard. A traversal can only see one shard's segments, so a graph split across shards would silently return partial answers. A multi-shard index is refused when the field is mapped, which is the last moment the fix is cheap.

Index edges​

One document is one edge:

{ "road": { "source": 12345, "target": 67890 } }
  • Both source and target are required. A half edge would index cleanly and never appear in a traversal, so it is rejected.
  • One edge per document. An array of edges is rejected; index each edge as its own document.
  • Any other field on the document — a weight, a name, a timestamp — is an ordinary field, updated without touching the graph.

Deleting, updating, merging, replicating and snapshotting edges are the index's normal behaviour.

Traverse​

{
"query": {
"graph_traversal": { "field": "road", "source": 12345, "hops": 2 }
}
}

Without node_field, the edge documents come back. hops counts edges: one hop matches the edges leaving source, two hops adds the edges leaving their targets.

With node_field, the documents whose node the walk reached come back instead. This is how a corpus is filtered by a graph — for example, every document about a concept or anything beneath it:

{
"query": {
"graph_traversal": {
"field": "broader",
"source": 7160,
"node_field": "concept_id",
"direction": "reverse"
}
}
}
ParameterRequiredDefaultMeaning
fieldyes—The graph_edge field to walk. An unmapped field is named in the error.
sourceyes—The start node. Never defaults: zero is a real node id.
hopsnounboundedHow many edges to follow. The walk is bounded by max_nodes, not depth.
node_fieldno—Match documents whose node, held in this field, was reached.
max_nodesno65536How many nodes the walk may reach. Past this the query fails rather than returning a partial answer.
directionnoforwardforward follows edges as written; reverse follows them backwards. reverse requires node_field.
boost, _nameno—As on any query.

Unknown parameters are refused. The query is also refused when search.allow_expensive_queries is false, because its cost is the size of the neighbourhood and cannot be bounded in advance.

Which direction​

If edges are written from the specific to the general — dog → mammal → animal — then forward from dog reaches its ancestors, and reverse from animal reaches everything beneath it. Neither answer contains the other, and a walk in the wrong direction returns a plausible list of the wrong documents, which is why the direction is never inferred.

How it is stored​

Each edge is two numeric doc values, <field>.src and <field>.tgt. The first query that needs a segment's adjacency derives it from those and caches it for as long as the segment lives, so a refresh costs a segment's worth of work, not the whole graph's.

The derived adjacency is written to a scratch directory on each node, <data path>/lucenia-graph/, and not into the index. It is not part of a commit, so snapshots and replicas never carry it and every node derives its own. Two things follow for operators:

  • The first traversal after a refresh or merge pays to derive the new segments.
  • The scratch directory uses disk that does not appear in index size statistics. It is released as segments close and when an index is deleted, and emptied when a node starts.

Concept similarity​

concept_similarity ranks documents by how alike their concept is to one you name. It is a scoring clause, not a filter: it composes inside a bool alongside BM25 and vectors and contributes a score.

{
"query": {
"concept_similarity": {
"field": "subclass_of",
"source": 278,
"node_field": "concept",
"measure": "wu_palmer",
"min_similarity": 0.4
}
}
}
OptionRequiredDefaultMeaning
fieldyes—The graph_edge field describing the vocabulary.
sourceyes—The concept everything is scored against.
node_fieldyes—The field carrying each searched document's concept.
measurenowu_palmerwu_palmer or path.
min_similarityno0.0Score below which a concept is neither returned nor searched beneath.
max_nodesno65536How many concepts one scoring pass may consider.

source is required even though it is a number, because a long defaults to zero and zero is a perfectly plausible concept id — an omitted source would otherwise score against something you never named. node_field must be an integral type with doc values.

min_similarity is not merely a filter. The search stops climbing once nothing beneath the current ancestor could reach it, so raising it makes the query cheaper as well as narrower.

Why this is affordable​

It looks like per-document graph work and is not. The vocabulary is scored against the source once, when the query is built, and the per-document query holds a sorted map from concept to score. Each document then costs a numeric doc-value read and a binary search — about what a range query costs.

Only structural measures​

wu_palmer and path need nothing but the hierarchy's shape, and are computed from the shard's own adjacency. The three corpus-frequency measures — Resnik, Lin and Jiang-Conrath — are refused with an error, because the frequencies have to be recorded by a cluster action and a plugin installable on managed OpenSearch cannot add one. graph_index is refused for the same reason: keep the vocabulary and the documents it ranks in the same single-shard index.

What it does not return​

Documents whose concept shares no ancestor with the source, at any threshold. They are not slightly similar, they are unrelated, and a vocabulary with disjoint roots is mostly made of them.

The query needs search.allow_expensive_queries to be enabled. If a pass hits max_nodes before finishing it fails rather than returning a partial ranking — raise max_nodes or raise min_similarity.

Ranking, honestly. Concept expansion is measured to recover a complete concept set where lexical and embedding retrieval do not. Adding concept_similarity to an already-hybrid ranking did not measurably improve nDCG@10, and was reliably unhelpful on common terminology. Use it for recall of a set, not as a general relevance booster.

Not yet exposed​

Contraction-hierarchy routing — NetworkDistanceQuery, for road-network shortest-path distance — is implemented in the graph package but has no query exposed.

Walking a graph held in another index (graph_index) is refused. It needs a cluster action, which a plugin installable on managed OpenSearch cannot add. Keep the graph and the documents it filters in the same single-shard index.

Walking a graph held in another index (graph_index) is refused. It needs a cluster action, which a plugin installable on managed OpenSearch cannot add. Keep the graph and the documents it filters in the same single-shard index.

On Amazon OpenSearch Service​

The graph_edge field and the graph_traversal query use only the mapper and search plugin extensions, which Amazon OpenSearch Service accepts for custom plugins. The accelerator as a whole does not install there, because its codec is an engine plugin; see Compatibility.