Content extraction processor
content_extract turns a file into indexable text. It uses Apache Tika, so it handles PDF, DOCX,
HTML, plain text and images, and it emits a normalized array of content blocks rather than one
undifferentiated string.
{
"content_extract": {
"input_mode": "reference",
"source_uri_field": "url",
"target_field": "extracted"
}
}
Four input modes
The mode decides where the bytes come from, and each has its own size ceiling because each has a different blast radius.
inline the field already holds the text or base64 1 MB
attachment a base64 attachment on the document 10 MB
stream a multipart stream supplied with the request 100 MB
reference a URI named in the document; the node fetches it 250 MB
| Mode | Option that names the input | Default limit |
|---|---|---|
inline (default) | field | max_inline_bytes — 1 MB |
attachment | field | max_attachment_bytes — 10 MB |
stream | field | max_stream_bytes — 100 MB |
reference | source_uri_field | max_reference_bytes — 250 MB |
The ceilings differ by an order of magnitude at each step because the cost of a mistake does. An oversized inline field wastes a request; an unbounded reference fetch turns one indexing call into arbitrary outbound traffic from your cluster.
Options
| Option | Default | Meaning |
|---|---|---|
field | content | Source field for inline, attachment and stream modes. |
target_field | extracted | Where the extracted blocks are written. |
input_mode | inline | One of inline, attachment, stream, reference. |
source_uri_field | source_uri | (reference mode) Field holding the URI to fetch. |
region_field | — | (reference mode) Field naming a sub-region to extract, for large rasters. |
mime_type_field | — | Field holding the MIME type, if the document knows it. |
mime_type | — | Fallback MIME type when detection and mime_type_field both come up empty. |
preserve_image_data | false | Keep image bytes on image/* blocks for a downstream processor. |
max_inline_bytes | 1 MB | |
max_attachment_bytes | 10 MB | |
max_stream_bytes | 100 MB | |
max_reference_bytes | 250 MB |
Reference mode is guarded
In reference mode the processor fetches a URI named inside the document. Anyone who can index a
document could otherwise point your cluster at your own metadata service.
Every fetch goes through the source access controller, which denies all hosts until you list them:
octane.content.source.allowed_hosts: ["docs.example.com"]
Redirects are not followed automatically, and the destination is authorised again if one is handled — otherwise an allowed host could redirect to a denied one and launder the request.
Output
target_field receives an array of content blocks with the extracted text, plus document-level
metadata such as language, page count and detected MIME type. For image/* sources the block can
carry the image bytes when preserve_image_data is on, which is what lets
image_tiling work from the same extraction rather than fetching the file twice.
Downstream processors consume the blocks — compliance scans their text, and the chunking and
embedding processors (in the separate inference artifact) consume the same shape.
Missing fields
If the configured source field is absent, the document fails. Use the usual ingest-pipeline
on_failure handling if you index heterogeneous documents where the field is genuinely optional.