Runtime detection and observability

Multi-tenant OTLP ingest with Alloy, the secure way

Your OTLP endpoint accepts logs from anyone who can find it, and the tenant is whatever the sender says it is. That is not a telemetry pipeline. That is a public comment box with your SIEM on the other end.

The short answer

Give each tenant its own OTLP receiver in Alloy, each with TLS 1.3 and its own basic-auth htpasswd file, a memory limiter, and an exporter that sets X-Scope-OrgID to that tenant. The tenant then comes from which credential was accepted, never from a header or attribute the client controls.

Updated Houssam Hammoudi, CTOTested with Grafana Alloy 1.19.2, Loki 3.7.8

On this page
  1. What goes wrong
  2. What the docs say
  3. The secure configuration
  4. Prove it
  5. Mistakes people make
  6. Checklist

What goes wrong

An OpenTelemetry receiver listens on 4317 (gRPC) and 4318 (HTTP). By default it has no TLS and no authentication. Anyone who can reach it can send logs, metrics and traces, and many setups expose it so that apps in other clusters or browsers can report.

For security telemetry that is a real risk. An attacker can flood the pipeline to hide their tracks, or write fake events into another team's tenant to cause confusion or bury a real alert.

Multi-tenancy makes it worse when the tenant comes from the client. If the pipeline forwards a tenant header or resource attribute that the sender set, every sender can write into every tenant.

What the docs say

Your OTel Collector configuration should include encryption and authentication.

Source: OpenTelemetry docs, Collector configuration best practices

Try to always use specific interfaces, such as a pod's IP, or localhost instead of 0.0.0.0.

Source: OpenTelemetry docs, Collector configuration best practices

otelcol.auth.basic exposes a handler that other otelcol components can use to authenticate requests using basic authentication.

Source: Grafana Alloy docs, otelcol.auth.basic

The docs cover authentication and TLS. They leave tenant mapping to you. The simplest mapping that cannot be forged is structural: one receiver per tenant, so the credential that was accepted decides the tenant.

The secure configuration

hcl
// config.alloy: one authenticated OTLP entry point per tenant.
// The tenant comes from WHICH receiver accepted the data, never from anything the client sends.

// ---- tenant-a -------------------------------------------------------------
otelcol.auth.basic "tenant_a" {
  htpasswd {
    file = "/etc/alloy/secrets/tenant-a.htpasswd"   // htpasswd entries, one per sender
  }
}

otelcol.receiver.otlp "tenant_a" {
  http {
    endpoint = "0.0.0.0:4318"
    auth     = otelcol.auth.basic.tenant_a.handler
    tls {
      cert_file   = "/etc/alloy/tls/server.crt"
      key_file    = "/etc/alloy/tls/server.key"
      min_version = "1.3"
    }
  }
  output {
    logs = [otelcol.processor.memory_limiter.tenant_a.input]
  }
}

// Drop data rather than run out of memory when one tenant floods the endpoint.
otelcol.processor.memory_limiter "tenant_a" {
  check_interval = "1s"
  limit          = "150MiB"
  output {
    logs = [otelcol.processor.batch.tenant_a.input]
  }
}

otelcol.processor.batch "tenant_a" {
  output {
    logs = [otelcol.exporter.otlphttp.loki_tenant_a.input]
  }
}

otelcol.exporter.otlphttp "loki_tenant_a" {
  client {
    endpoint = "http://loki:3100/otlp"
    headers  = { "X-Scope-OrgID" = "tenant-a" }   // set here, by the pipeline
  }
}

// ---- tenant-b: the same shape on its own port and credentials ---------------
otelcol.auth.basic "tenant_b" {
  htpasswd {
    file = "/etc/alloy/secrets/tenant-b.htpasswd"
  }
}

otelcol.receiver.otlp "tenant_b" {
  http {
    endpoint = "0.0.0.0:4319"
    auth     = otelcol.auth.basic.tenant_b.handler
    tls {
      cert_file   = "/etc/alloy/tls/server.crt"
      key_file    = "/etc/alloy/tls/server.key"
      min_version = "1.3"
    }
  }
  output {
    logs = [otelcol.processor.memory_limiter.tenant_b.input]
  }
}

otelcol.processor.memory_limiter "tenant_b" {
  check_interval = "1s"
  limit          = "150MiB"
  output {
    logs = [otelcol.processor.batch.tenant_b.input]
  }
}

otelcol.processor.batch "tenant_b" {
  output {
    logs = [otelcol.exporter.otlphttp.loki_tenant_b.input]
  }
}

otelcol.exporter.otlphttp "loki_tenant_b" {
  client {
    endpoint = "http://loki:3100/otlp"
    headers  = { "X-Scope-OrgID" = "tenant-b" }
  }
}

Notes for production:

  • In Kubernetes, bind to the pod IP (endpoint = string.format("%s:4318", sys.env("POD_IP"))) instead of 0.0.0.0, and expose each tenant port through its own Service or gateway route.
  • Keep the htpasswd files and the TLS key in Kubernetes Secrets or a secret store, mounted read-only. Give every sender its own entry so one leaked credential can be removed without touching the others.
  • The exporter to Loki sets X-Scope-OrgID. Loki must still sit behind its own authenticating gateway or network policy; see the Loki page.
  • Add the same pipelines for metrics and traces if the tenant sends them.

Prove it

The configuration passes alloy fmt:

text
config.alloy: OK

Sends to the tenant-a receiver: no credentials, a wrong password, and tenant-b's valid credentials are all refused. A valid tenant-a sender that also sets X-Scope-OrgID: tenant-b is accepted, and plain HTTP to the TLS port is rejected:

bash
curl -s -o /dev/null -w "%{http_code}" --cacert ca.crt -H "Content-Type: application/json" \
  -u shipper-a:*** --data-binary @log.json https://alloy:4318/v1/logs
text
no credentials: HTTP 401
wrong password: HTTP 401
tenant-b credentials on the tenant-a port: HTTP 401
shipper-a, valid: HTTP 200
shipper-a claiming X-Scope-OrgID tenant-b: HTTP 200
shipper-b, valid: HTTP 200
plain HTTP to the TLS port: HTTP 400

What landed in Loki, per tenant. The forged header changed nothing; the line went to tenant-a, because the tenant-a pipeline set the header:

bash
curl -s -G -H "X-Scope-OrgID: tenant-a" --data-urlencode 'query={service_name="app"}' http://loki:3100/loki/api/v1/query_range
text
-- tenant-a
  shipper-a, valid
  shipper-a claiming X-Scope-OrgID tenant-b
-- tenant-b
  shipper-b, valid

The test (certificates and passwords generated per run) is secure-tests/multi-tenant-otlp-ingest-alloy/run.sh.

Mistakes people make

An open receiver on the internet

0.0.0.0:4318 without auth accepts anything from anyone who can route to it. Add TLS and authentication before the first external sender connects.

Taking the tenant from the sender

Forwarding a client header (include_metadata with a header setter) or a tenant resource attribute lets every sender pick any tenant. Map tenants from the authenticated identity.

One credential shared by every sender

Rotation then means reconfiguring every sender at once, and a leak from one host is a leak for all. One htpasswd entry per sender.

No memory limiter

A single noisy or hostile sender can push Alloy out of memory, and every tenant loses telemetry. Put otelcol.processor.memory_limiter first in each pipeline.

TLS without checking the client side

Senders must verify the server certificate (ca_file in the exporter config), or anyone on the path can pose as the collector and collect the credentials.

Checklist

  • Every OTLP receiver has TLS with min_version = "1.3".
  • Every OTLP receiver has an auth handler.
  • Each tenant has its own receiver and credentials.
  • The tenant ID is set by the exporter, never taken from client input.
  • Each sender has its own credential entry.
  • Each pipeline starts with a memory limiter.
  • Receivers bind to a specific interface, not 0.0.0.0, in production.
  • Loki behind Alloy is itself protected by a gateway or network policy.
  • A test sends with wrong credentials and a forged tenant header and checks both fail.

A telemetry pipeline decides what your security team will believe. Make sure it only believes senders who can prove who they are.

H2-CTDE

Learn it on a live range

The detection pipeline and SIEM, in Runtime Detection and Response: a real host in your browser, and every objective checked on the machine.

Start free

The Secure Way

More on runtime detection and observability

Tetragon, alerting as code, multi-tenant logs and knowing when a sensor goes quiet.

All runtime detection and observability guides