Starting, stopping, restarting, and refreshing Cloudera Embedded Container Service Clusters

Provides information about how you can manage your Cloudera Data Services on premises Cloudera Embedded Container Service clusters.

Starting a Cloudera Embedded Container Service Cluster

Procedure to start an Cloudera Embedded Container Service cluster.

  1. Navigate to the ECS service, Home > Status tab, click the Actions menu to the right of the Embedded Container Service cluster name and select Start.
  2. Click the Start button that appears in the next screen to confirm. The Command Details window shows the progress of starting services.
When the All services successfully started message appears, the task is complete and you can close the Command Details window.

Stopping a Cloudera Embedded Container Service Cluster

Procedure to stop an Cloudera Embedded Container Service cluster.

  1. Navigate to the ECS service, Home > Status tab, click the Actions menu to the right of the Cloudera Embedded Container Service cluster name and select Stop.
  2. Click the Stop button in the confirmation screen. The Command Details window shows the progress of stopping services.
When the All services successfully stopped message appears, the task is complete and you can close the Command Details window.

Stopping the Cloudera Embedded Container Service Agent causes yunikorn deployment to scale down to 0

To overcome this scenario, you must stop the ECS Agent at the role level.

  1. Log into your Cloudera Manager instance.
  2. Navigate to the ECS service.
  3. Click the Instances tab.
  4. Locate the instance that needs to be stopped and click ECS Agent in the Role Type column.
  5. On the ECS Role page, open the Actions menu.
  6. Select Stop this ECS Agent.

Before restarting the High Availability enabled Cloudera Embedded Container Service Cluster

You must review the following information prior to restarting your High Availability (HA) enabled Cloudera Embedded Container Service cluster.

About Full Cluster Restart

When a full cluster restart is required (rolling restart is not recommended), do not restart all roles at once. Employ one of the procedures.

Option A: Start Primary Master First, then All Others

  1. Start the primary master (the bootstrap node that existed before other nodes were added).
  2. Start all other roles (For example, additional masters, workers, and others).

Option B: Restart Non-Master First, then Rolling Restart Masters

  1. Restart all roles except master nodes (For example, workers, agents).
  2. Perform a rolling restart after selecting all the master nodes.
Table 1. Summary
Scenario Procedure
Full cluster start Start primary master first, then all other roles
Full cluster restart Restart non-masters first, then rolling restart after selecting all master nodes

Rolling Restart of an Cloudera Embedded Container Service Cluster

Procedure to perform the rolling restart.

  1. Navigate to the ECS service, Home > Status tab, click the Actions menu to the right of the cluster name and select Rolling Restart.
  2. Click the Rolling Restart button that appears in the next screen to confirm. On this screen, you can select the services (Docker or /and ECS), Roles (Workers only, Non-workers only, All Roles).
    The Command Details window shows the progress of rolling restart of a batch of nodes. Here, batch size refers to the number of worker roles that can be restarted in parallel. The Batch size is 1 by default.
  3. Click Actions > Unseal Vault

Restarting a Cloudera Embedded Container Service Cluster

Procedure to restart your Cloudera Embedded Container Service cluster.

  1. Navigate to the ECS service, Home > Status tab, click the Actions menu to the right of the cluster name and select Restart.
  2. Click the Restart button that appears in the next screen to confirm.
  3. Click Actions > Unseal Vault

Configuring Restart for an Cloudera Embedded Container Service cluster

Follow these instructions to configure a cluster restart.

  1. Navigate to the ECS service, Home > Configuration tab.
  2. Select Node Readiness Timeout OR Drain Node Timeout configurations to configure overall restart time for an Cloudera Embedded Container Service cluster.
    Table 2. Cloudera Embedded Container Service Cluster Actions and Performance Impact
    Cloudera Embedded Container Service Cluster action Affects Availability Impact Speed
    Stop and Start Yes All nodes Fastest
    Restart Least one node at a time Slowest
    Rolling restart (if the agent batch size is 1 then it is similar to Restart).

    Inversely proportional to the batch size.

    Depends on batch size. Speed is proportional to the batch size.That is, higher the batch size, higher the speed.

Refreshing a Cloudera Embedded Container Service Cluster

Refresh your Cloudera Embedded Container Service cluster with a single step process.

Navigate to the ECS service, Home > Status tab, click the Actions menu to the right of the cluster name and select Refresh Cluster.

Terminate All RKE2 Processes on Cloudera Embedded Container Service Hosts

The Terminate all RKE2 processes commands stop RKE2 completely on one Cloudera Embedded Container Service host or on all Cloudera Embedded Container Service hosts in a service. Cloudera Manager runs the same rke2-killall.sh script shipped with the Cloudera Embedded Container Service parcel, with confirmation and logging.

Use these commands when a normal Stop or Bring Down does not fully terminate RKE2—for example, when containerd shims, Pods, or systemd units leave processes running.

Terminate all RKE2 processes (single host)
  • Scope: One Cloudera Embedded Container Service Server or ECS Agent role
  • UI path: Clusters → cluster → ECS → role → Actions → Terminate all RKE2 processes.
  • What it does:
    1. Brings down the selected role (bring-down failures are ignored).
    2. On the role’s host, stops RKE2 systemd units if present (Cloudera Embedded Container Service 1.5.5 SP4+).
    3. Runs rke2-killall.sh to terminate all remaining RKE2-related processes.
Terminate all RKE2 processes on all Cloudera Embedded Container Service hosts (service-wide)
  • Scope: Entire Cloudera Embedded Container Service.
  • UI path: Clusters → cluster → ECS → Actions → Terminate all RKE2 processes on all ECS hosts
  • What it does:
    1. Brings down the Cloudera Embedded Container Service (bring-down failures are ignored).
    2. Runs the per-role killall command on every Cloudera Embedded Container Service server and agent host in parallel.
After running the command:
  • Affected roles and the Cloudera Embedded Container Service are stopped in Cloudera Manager.
  • RKE2 is not running on the affected host(s).
  • To restore the cluster, run Bring Up on the Cloudera Embedded Container Service (or individual roles) through Cloudera Manager.