Skip to main content

Importing a vocabulary

Loads a vocabulary — a taxonomy, thesaurus, ontology, or a domain graph written as an edge list — into an index, as concepts concept_link can match and edges a graph traversal can walk.

The readers and the concept_link processor both predate this endpoint. What was missing was the step between them: nothing outside a test could put a vocabulary into the cluster, so a linker that worked had nothing to work on.

Route​

POST /_plugins/_ontology/{vocabulary}/_import

{vocabulary} is the index written to, and the name a concept_link processor refers to it by. It is created if absent.

Request fields​

FieldTypeDescriptionRequired
uristringWhere to fetch from: https:// or http://yes
formatstringrdf (default) for Turtle/N-Triples, or edge_listno
profilestringskos, owl, or default (both). RDF onlyno
predicatestringEdge field the rows become. Required for edge_listno
delimiterstringSingle character; omit to infer from the first rowno
reference_configobjectPassed to the resolver unchanged — see belowno
max_conceptsintegerCap on concepts held in memory. Default 2000000no
batch_sizeintegerDocuments per bulk request. Default 1000no

reference_config is the same per-request resolver configuration the ingest processors take. This plugin reads authorization (sent as the HTTP Authorization header) and read_deadline_seconds.

Example​

POST /_plugins/_ontology/concepts/_import
{
"uri": "https://vocab.example.com/taxonomy.ttl",
"format": "rdf",
"profile": "default"
}
{
"vocabulary": "concepts",
"concepts": 2847,
"relations": 391,
"predicates": ["broader"],
"took_in_millis": 412
}

Counts rather than an acknowledgement. An import that matched no predicate the profile knows still succeeds at every step, so "acknowledged": true would sit over an empty vocabulary. "concepts": 0 is how that becomes visible without going to look.

Fetching obeys the same allowlist as everything else​

A vocabulary URI is an ordinary fetch, governed by octane.content.source.allowed_hosts exactly like any URI a document points at — see Source access. Nothing is fetchable until an operator says so, and an unallowed host is refused before a connection is attempted.

There is no file:// resolver. A vocabulary on local disk cannot be imported; it has to be somewhere fetchable. That follows from governance being about fetchable sources, but it will surprise someone.

What gets written​

Concept documents carry id, iri, label and alt_labels — exactly the fields the linker reads. Edge documents carry one object field named for the predicate:

{ "broader": { "source": 12345, "target": 67890 } }

Both share one index. The linker skips any document without a numeric id, so edges are invisible to it by construction and a caller names one index rather than keeping two in step.

An edge field is mapped before any document uses it. A document reaching an unmapped field is mapped dynamically as an object of two longs; a field's type cannot be changed afterwards; and the result indexes cleanly, reports success, and traverses as nothing at all.

Edges need the accelerator​

graph_edge is a field type the octane-accelerator plugin registers, not this one. A vocabulary of concepts and labels imports without it. A vocabulary with a hierarchy fails when the first edge field is mapped, and the failure names the plugin rather than reading as an unknown-field-type error nobody can act on.

Single shard, necessarily​

The index is created with number_of_shards: 1. A graph_edge field refuses any other count — a traversal walks one shard's segments, and a graph spread across shards answers partially while appearing complete. A vocabulary therefore lives on one shard however large it is. Importing into an existing multi-shard index takes the concepts and refuses the edges.

Re-importing is repeatable​

Identities are derived — a concept's from its IRI, an edge's from the two concepts it joins — so re-importing rewrites the same documents in place rather than appending. A vocabulary published in fragments imports fragment by fragment, in any order.

A bulk failure stops the import rather than continuing past it: a vocabulary quietly missing whatever the cluster was too busy to accept is worse than one that failed loudly. Run it again.

Limits​

  • Memory is concepts-sized, not file-sized. A concept's labels are scattered through an RDF file, so concepts are held until it closes while relations stream. max_concepts bounds this and fails by name rather than by OutOfMemoryError.
  • The label cache is per node. The linker holds a vocabulary from first use, and an import drops that cache only on the node that ran the import. Re-importing a vocabulary already in use elsewhere goes on matching by the old labels on other nodes until they restart.

Licensing​

Importing a vocabulary requires a license, and the refusal is a 403 naming the content feature.

It is gated on the same entitlement as concept_link, because an imported vocabulary is the data that processor matches against — the capability and its loader are one paid feature, so a cluster licensed for content can load the vocabularies it is entitled to use.

Reading an already imported vocabulary is never gated. Searching it, traversing its edges and linking against it keep working if a license lapses, in line with the rule that a license gates starting work and never finishing or undoing it — see Licensing. What stops is importing a new one.

A 30-day trial can be started from the cluster:

POST _lucenia/license/_start_trial?issued_to=<your-org>