UNCLASSIFIED - NO CUI

fix(fluentbit): replace Loki label denylist with explicit allowlist

Package Merge Request

Package Changes

Fixes fluentbit issue big-bang/product/packages/fluentbit#244 (closed)— Fluent Bit pods becoming unready when ECK log streams exceed Loki's default 15-label limit.

Root Cause

The previous design used auto_kubernetes_labels on (promotes all Kubernetes pod labels to Loki) combined with a Lua denylist (remove_labels.lua) that stripped known-bad patterns. ECK attaches 10+ labels to Elasticsearch data and master pods (elasticsearch.k8s.elastic.co/*, common.k8s.elastic.co/*). After filtering, ECK pods landed at 16 Loki labels, triggering repeated HTTP 400 rejections from Loki. Those errors failed the Fluent Bit /api/v2/health endpoint, leaving the DaemonSet permanently unready during upgrades.

The denylist model is fundamentally reactive - it has already required two previous patches (for Vault) and would need another update for every new operator that attaches labels to pods.

Change

Replaces the denylist with an explicit seven-label allowlist:

Label Source
job static
cluster Helm template value
container $kubernetes['container_name']
pod $kubernetes['pod_name']
namespace $kubernetes['namespace_name']
node_name $kubernetes['host']
app_kubernetes_io_name $kubernetes['labels']['app.kubernetes.io/name']
  • auto_kubernetes_labels set to off
  • remove_labels.lua and its filter block removed entirely
  • cluster label now set directly from the Helm template value

Do We Lose Anything?

All existing Loki dashboard queries continue to work. An audit of all Grafana dashboards in the repo confirmed namespace, pod, container, cluster, and app_kubernetes_io_name (required by bbctl dashboards) are all retained.

Labels dropped that were previously auto-promoted:

Dropped label Impact
app_kubernetes_io_instance No dashboard queries it via LogQL
app_kubernetes_io_version No dashboard queries it; was high-cardinality (new series per release)
controller-revision-hash No dashboard queries it; was very high-cardinality (new series every rollout) - latent Loki degradation bug
pod-template-hash No dashboard queries it; same cardinality issue
All ECK/Vault/Istio operator labels No dashboard queries them; operational metadata

The high-cardinality removals are an incidental correctness improvement - those were silently creating new Loki series on every deployment cluster-wide. The allowlist hard-caps at 7 labels today, leaving 8 of the 15-label budget unused for future use.

Testing

  • Helm unit test added verifying auto_kubernetes_labels off and cluster= present in rendered Loki output when loki.enabled: true
  • Validated on a local cluster with ECK (data + master pods) - all 4 Fluent Bit pods 2/2 Ready, zero HTTP 400 label-limit errors in logs

Package MR

N/A - change is in the umbrella chart (bigbang/chart/templates/fluentbit/values.yaml), not the fluentbit package repo.

For Issue

Closes big-bang/product/packages/fluentbit#244 (closed)

Upgrade Notices

Operators querying Loki by app_kubernetes_io_instance, app_kubernetes_io_version, controller-revision-hash, or pod-template-hash labels will need to switch to pod or app_kubernetes_io_name instead. No built-in BB dashboards are affected.

Edited by Kirby Liu

Merge request reports

Loading
Loading