Skip to main content

Content extraction processor

content_extract turns a file into indexable text. It uses Apache Tika, so it handles PDF, DOCX, HTML, plain text and images, and it emits a normalized array of content blocks rather than one undifferentiated string.

{
"content_extract": {
"input_mode": "reference",
"source_uri_field": "url",
"target_field": "extracted"
}
}

Four input modes​

The mode decides where the bytes come from, and each has its own size ceiling because each has a different blast radius.

  inline      the field already holds the text or base64                 1 MB
attachment a base64 attachment on the document 10 MB
stream a multipart stream supplied with the request 100 MB
reference a URI named in the document; the node fetches it 250 MB
ModeOption that names the inputDefault limit
inline (default)fieldmax_inline_bytes — 1 MB
attachmentfieldmax_attachment_bytes — 10 MB
streamfieldmax_stream_bytes — 100 MB
referencesource_uri_fieldmax_reference_bytes — 250 MB

The ceilings differ by an order of magnitude at each step because the cost of a mistake does. An oversized inline field wastes a request; an unbounded reference fetch turns one indexing call into arbitrary outbound traffic from your cluster.

Options​

OptionDefaultMeaning
fieldcontentSource field for inline, attachment and stream modes.
target_fieldextractedWhere the extracted blocks are written.
input_modeinlineOne of inline, attachment, stream, reference.
source_uri_fieldsource_uri(reference mode) Field holding the URI to fetch.
region_field—(reference mode) Field naming a sub-region to extract, for large rasters.
mime_type_field—Field holding the MIME type, if the document knows it.
mime_type—Fallback MIME type when detection and mime_type_field both come up empty.
preserve_image_datafalseKeep image bytes on image/* blocks for a downstream processor.
max_inline_bytes1 MB
max_attachment_bytes10 MB
max_stream_bytes100 MB
max_reference_bytes250 MB

Reference mode is guarded​

In reference mode the processor fetches a URI named inside the document. Anyone who can index a document could otherwise point your cluster at your own metadata service.

Every fetch goes through the source access controller, which denies all hosts until you list them:

octane.content.source.allowed_hosts: ["docs.example.com"]

Redirects are not followed automatically, and the destination is authorised again if one is handled — otherwise an allowed host could redirect to a denied one and launder the request.

Output​

target_field receives an array of content blocks with the extracted text, plus document-level metadata such as language, page count and detected MIME type. For image/* sources the block can carry the image bytes when preserve_image_data is on, which is what lets image_tiling work from the same extraction rather than fetching the file twice.

Downstream processors consume the blocks — compliance scans their text, and the chunking and embedding processors (in the separate inference artifact) consume the same shape.

Missing fields​

If the configured source field is absent, the document fails. Use the usual ingest-pipeline on_failure handling if you index heterogeneous documents where the field is genuinely optional.