Skip to main content

Source access control

Several processors fetch a URI that is named inside a document: content_extract in reference mode, image_tiling, and vectorize. That is a server-side request forgery surface. Whoever can index a document can otherwise choose what your cluster connects to.

Octane's answer is a deny-by-default allowlist, enforced at the one place every fetch passes through.

Configure it​

No host is reachable until you name it:

octane.content.source.allowed_hosts: ["imagery.example.com", "docs.example.com"]

With the setting unset or empty, every fetch is refused:

fetching content from [https://internal.example/x] is not permitted:
no source hosts are allowed. Set octane.content.source.allowed_hosts

An allowed list that does not include the host gives the more specific message:

fetching content from [https://169.254.169.254/latest/meta-data/] is not permitted:
host [169.254.169.254] is not in octane.content.source.allowed_hosts

Host matching is exact​

A host must match exactly. There is no wildcard and no suffix matching — example.com does not admit evil.example.com.attacker.net, and it does not admit data.example.com either. List each host you intend to allow.

This is deliberately inconvenient. Suffix matching is the mechanism behind a long history of allowlist bypasses, and an allowlist that is easy to write loosely is one that will be.

Redirects are re-checked​

Redirects are not followed automatically, and when one is handled the destination is authorised again:

   document names https://allowed.example.com/a
│
├─ authorise allowed.example.com ✓
├─ fetch → 302 → https://169.254.169.254/
└─ authorise 169.254.169.254 ✗ refused

Without the second check, an allowed host could redirect anywhere and launder the request. One authorisation would have been a guard against typos rather than against an adversary.

Why the java agent does not cover this​

OpenSearch 3.x replaced the SecurityManager with a java agent that mediates socket connections and a set of file operations. It is a useful guard, and it does not help here.

The agent can answer "may this code open sockets?" — and for an ingest plugin that fetches content, the answer is permanently yes, or the feature cannot work at all. It has no way to know that the URL came from an untrusted document.

"Should this URL be fetched" is an application-layer question. The agent structurally cannot answer it, which is exactly the gap this controller fills, and why deny-by-default is the right posture on the COG path rather than a precaution.

Operational advice​

  • Name hosts, not networks. If you find yourself wanting a range, you probably want an egress proxy with its own policy, and one allowed host pointing at it.
  • Back it with network policy. An allowlist is a control in your cluster; it is not a substitute for the node being unable to reach the metadata service at all. Defense in depth means both.
  • Review it like a firewall rule. It is one, and it is easier to add a host than to notice one that should have been removed.