Content
octane-content adds processors to the ingest pipeline: it turns files into indexable text,
redacts what must not be stored, and turns imagery into tiles and geometry. Everything happens
before a document is indexed, so what lands in the index is already clean.
Processors
| Type | Stage | What it does |
|---|---|---|
content_extract | ingest | PDF, DOCX, HTML, images → normalized text blocks (Apache Tika) |
chunk | ingest | Splits extracted blocks into overlapping chunks, four ways |
embed | ingest | Turns chunks into vectors through an embedding provider |
compliance | ingest | Detects and redacts PII/PHI/PCI/secrets by policy |
image_tiling | ingest | Streams a Cloud Optimized GeoTIFF and emits tiles |
vectorize | ingest | Categorical raster → indexable geometry, no model required |
reproject | ingest | Converts a shape from one coordinate reference system to another |
concept_link | ingest | Links text to a controlled vocabulary |
topic_drift | ingest | Scores how far an analysis has drifted from the goal it was given |
rerank_prepare | ingest | Shapes chunks for a downstream reranker |
query_embedding | search request | Embeds the query at search time and injects the vector |
retrieval_grounding | search response | Attaches provenance and previews to hits |
compliance_redact | search response | Redacts sensitive values in results for non-exempt callers |
Three of these are not ingest processors. query_embedding is a search request processor and
retrieval_grounding and compliance_redact are search response processors — OpenSearch keeps
those in an entirely separate hierarchy, so they go in a search pipeline, not an ingest pipeline.
Loading the vocabulary that concept_link matches against is its own endpoint — see
Importing a vocabulary.
Endpoints
| Endpoint | What it does |
|---|---|
POST /_plugins/_ontology/{vocabulary}/_import | Loads a vocabulary as concepts and graph edges |
PUT|GET|DELETE /_plugins/_compliance/policies/{name} | Manages named compliance policies — see Compliance |
concept_link needs a vocabulary in the cluster before it can link anything, and the import endpoint
is how one gets there. See Ontology import.
A pipeline, end to end
document ─▶ content_extract ─▶ compliance ─▶ index
│ │
│ └─ redacts in place, records what it found
└─ fetches referenced URIs ─┐
▼
source access control
(deny by default)
PUT _ingest/pipeline/documents
{
"processors": [
{ "content_extract": { "input_mode": "reference", "source_uri_field": "url", "target_field": "extracted" } },
{ "compliance": { "fields": ["extracted.text"], "profile": "gdpr" } }
]
}
Order matters. Extraction must run first — compliance can only redact text that exists — and both must run before indexing, because a redaction applied afterwards has already been written to disk.
Fetching remote content is deny-by-default
Several processors can fetch a URI named inside a document. That is an SSRF surface: whoever can index a document could otherwise make your cluster fetch your metadata service.
No host is reachable until you list it:
octane.content.source.allowed_hosts: ["data.example.com", "imagery.example.com"]
See Source access. This is not redundant with OpenSearch's java agent — the agent can answer "may this code open sockets?" but not "should this URL be fetched", because it cannot know the URL came from an untrusted document.
What is not here
Three processors that call a remote model endpoint — ocr, image_segment and multimodal_rerank
— are deliberately not in this plugin. Each needs the inference stack and the dependency cone that
comes with it, including a per-provider cloud SDK. That belongs in its own artifact with its own
answer about credentials and egress, not smuggled into a plugin that already carries Tika and a
geospatial toolkit.
chunk, embed, query_embedding and topic_drift are here. They reach a model through the
generic HTTP and OpenAI-compatible providers, which need no vendor SDK — an endpoint, a model name,
and the same source allowlist every other fetch obeys. That is the line: a provider you configure,
not a provider you compile in.
vectorize sits on the other side of it entirely: for categorical rasters the labels are already in
the pixels, so no model is involved at all.
Licensing
Creating a pipeline that uses these processors requires a license. Running one never does — an ingest pipeline sits in front of indexing, and a processor that refused documents on a lapsed license would stop the cluster accepting writes. See Licensing.