Sizing and performance considerations

Kubernetes resource sizing guidance for Cloudera Streaming Analytics Operator for Kubernetes control plane components, Flink job clusters, and Cloudera SQL Stream Builder.

Control plane components

Size the namespace that hosts the Flink Kubernetes Operator and Cloudera SQL Stream Builder for steady-state overhead plus headroom for webhook admission and reconciliation bursts. The Helm chart default CPU and memory requests increased in recent releases; override them in values.yaml if your cluster is capacity constrained. Consider the following:

  • Flink Kubernetes Operator: reserve enough memory for informers when watching multiple namespaces.
  • Cloudera SQL Stream Builder SSE: scale vertically before adding replicas; the metadata database must sustain concurrent console users and job submissions.
  • Embedded PostgreSQL (default): use an external database for production deployments. The bundled database is intended for proof of concept and development.

Flink job clusters

Flink TaskManager count, CPU, memory, and disk depend on parallelism, state size, and checkpoint interval. Enable checkpointing to durable storage for production jobs. Use the Flink autoscaler configuration when jobs have variable load.

Storage and networking

Provision fast persistent volumes for checkpoint and savepoint directories. For Longhorn (Rancher/RKE2) and OpenStack Cinder (Taikun), set ssb.database.pod.securityContext.fsGroup: 999 when using the embedded PostgreSQL volume.

Taikun CloudWorks minimum cluster

For Technical Preview installs on Taikun, use at least 4 CPU and 16 GB RAM per master, bastion, and worker node (two workers). See Install Cloudera Streaming Analytics Operator for Kubernetes in Taikun CloudWorks [Technical Preview] for prerequisites.