fix(fluentbit): replace Loki label denylist with explicit allowlist
Package Merge Request
Package Changes
Fixes fluentbit issue big-bang/product/packages/fluentbit#244 (closed)— Fluent Bit pods becoming unready when ECK log streams exceed Loki's default 15-label limit.
Root Cause
The previous design used auto_kubernetes_labels on (promotes all Kubernetes pod labels to Loki) combined with a Lua denylist (remove_labels.lua) that stripped known-bad patterns. ECK attaches 10+ labels to Elasticsearch data and master pods (elasticsearch.k8s.elastic.co/*, common.k8s.elastic.co/*). After filtering, ECK pods landed at 16 Loki labels, triggering repeated HTTP 400 rejections from Loki. Those errors failed the Fluent Bit /api/v2/health endpoint, leaving the DaemonSet permanently unready during upgrades.
The denylist model is fundamentally reactive - it has already required two previous patches (for Vault) and would need another update for every new operator that attaches labels to pods.
Change
Replaces the denylist with an explicit seven-label allowlist:
| Label | Source |
|---|---|
job |
static |
cluster |
Helm template value |
container |
$kubernetes['container_name'] |
pod |
$kubernetes['pod_name'] |
namespace |
$kubernetes['namespace_name'] |
node_name |
$kubernetes['host'] |
app_kubernetes_io_name |
$kubernetes['labels']['app.kubernetes.io/name'] |
auto_kubernetes_labelsset tooffremove_labels.luaand its filter block removed entirelyclusterlabel now set directly from the Helm template value
Do We Lose Anything?
All existing Loki dashboard queries continue to work. An audit of all Grafana dashboards in the repo confirmed namespace, pod, container, cluster, and app_kubernetes_io_name (required by bbctl dashboards) are all retained.
Labels dropped that were previously auto-promoted:
| Dropped label | Impact |
|---|---|
app_kubernetes_io_instance |
No dashboard queries it via LogQL |
app_kubernetes_io_version |
No dashboard queries it; was high-cardinality (new series per release) |
controller-revision-hash |
No dashboard queries it; was very high-cardinality (new series every rollout) - latent Loki degradation bug |
pod-template-hash |
No dashboard queries it; same cardinality issue |
| All ECK/Vault/Istio operator labels | No dashboard queries them; operational metadata |
The high-cardinality removals are an incidental correctness improvement - those were silently creating new Loki series on every deployment cluster-wide. The allowlist hard-caps at 7 labels today, leaving 8 of the 15-label budget unused for future use.
Testing
- Helm unit test added verifying
auto_kubernetes_labels offandcluster=present in rendered Loki output whenloki.enabled: true - Validated on a local cluster with ECK (data + master pods) - all 4 Fluent Bit pods
2/2 Ready, zero HTTP 400 label-limit errors in logs
Package MR
N/A - change is in the umbrella chart (bigbang/chart/templates/fluentbit/values.yaml), not the fluentbit package repo.
For Issue
Closes big-bang/product/packages/fluentbit#244 (closed)
Upgrade Notices
Operators querying Loki by app_kubernetes_io_instance, app_kubernetes_io_version, controller-revision-hash, or pod-template-hash labels will need to switch to pod or app_kubernetes_io_name instead. No built-in BB dashboards are affected.