Skip to main content

Content

octane-content adds processors to the ingest pipeline: it turns files into indexable text, redacts what must not be stored, and turns imagery into tiles and geometry. Everything happens before a document is indexed, so what lands in the index is already clean.

Processors​

TypeStageWhat it does
content_extractingestPDF, DOCX, HTML, images → normalized text blocks (Apache Tika)
chunkingestSplits extracted blocks into overlapping chunks, four ways
embedingestTurns chunks into vectors through an embedding provider
complianceingestDetects and redacts PII/PHI/PCI/secrets by policy
image_tilingingestStreams a Cloud Optimized GeoTIFF and emits tiles
vectorizeingestCategorical raster → indexable geometry, no model required
reprojectingestConverts a shape from one coordinate reference system to another
concept_linkingestLinks text to a controlled vocabulary
topic_driftingestScores how far an analysis has drifted from the goal it was given
rerank_prepareingestShapes chunks for a downstream reranker
query_embeddingsearch requestEmbeds the query at search time and injects the vector
retrieval_groundingsearch responseAttaches provenance and previews to hits
compliance_redactsearch responseRedacts sensitive values in results for non-exempt callers

Three of these are not ingest processors. query_embedding is a search request processor and retrieval_grounding and compliance_redact are search response processors — OpenSearch keeps those in an entirely separate hierarchy, so they go in a search pipeline, not an ingest pipeline.

Loading the vocabulary that concept_link matches against is its own endpoint — see Importing a vocabulary.

Endpoints​

EndpointWhat it does
POST /_plugins/_ontology/{vocabulary}/_importLoads a vocabulary as concepts and graph edges
PUT|GET|DELETE /_plugins/_compliance/policies/{name}Manages named compliance policies — see Compliance

concept_link needs a vocabulary in the cluster before it can link anything, and the import endpoint is how one gets there. See Ontology import.

A pipeline, end to end​

   document  ─▶  content_extract  ─▶  compliance  ─▶  index
│ │
│ └─ redacts in place, records what it found
└─ fetches referenced URIs ─┐
▼
source access control
(deny by default)
PUT _ingest/pipeline/documents
{
"processors": [
{ "content_extract": { "input_mode": "reference", "source_uri_field": "url", "target_field": "extracted" } },
{ "compliance": { "fields": ["extracted.text"], "profile": "gdpr" } }
]
}

Order matters. Extraction must run first — compliance can only redact text that exists — and both must run before indexing, because a redaction applied afterwards has already been written to disk.

Fetching remote content is deny-by-default​

Several processors can fetch a URI named inside a document. That is an SSRF surface: whoever can index a document could otherwise make your cluster fetch your metadata service.

No host is reachable until you list it:

octane.content.source.allowed_hosts: ["data.example.com", "imagery.example.com"]

See Source access. This is not redundant with OpenSearch's java agent — the agent can answer "may this code open sockets?" but not "should this URL be fetched", because it cannot know the URL came from an untrusted document.

What is not here​

Three processors that call a remote model endpoint — ocr, image_segment and multimodal_rerank — are deliberately not in this plugin. Each needs the inference stack and the dependency cone that comes with it, including a per-provider cloud SDK. That belongs in its own artifact with its own answer about credentials and egress, not smuggled into a plugin that already carries Tika and a geospatial toolkit.

chunk, embed, query_embedding and topic_drift are here. They reach a model through the generic HTTP and OpenAI-compatible providers, which need no vendor SDK — an endpoint, a model name, and the same source allowlist every other fetch obeys. That is the line: a provider you configure, not a provider you compile in.

vectorize sits on the other side of it entirely: for categorical rasters the labels are already in the pixels, so no model is involved at all.

Licensing​

Creating a pipeline that uses these processors requires a license. Running one never does — an ingest pipeline sits in front of indexing, and a processor that refused documents on a lapsed license would stop the cluster accepting writes. See Licensing.