Troubleshooting Ozone replication policies

The troubleshooting scenarios in this topic help you to troubleshoot the Ozone replication policies in Replication Manager.

How to skip replicated files during subsequent Ozone replication policy runs?

Some subsequent Ozone replication policy runs might fall back to bootstrap replication if the incremental replication fails. You can choose to skip to replicate the already replicated files during these subsequent Ozone replication policy runs.

Problem

Subsequent bootstrap replication job runs replicate the files that were already replicated.

Cause

The first Ozone replication policy job is always a bootstrap job where all the specified data is replicated from the source cluster to the target cluster. Some subsequent runs might also fall back to bootstrap jobs if the incremental replication fails. In this scenario, the bootstrap replication jobs might replicate the files that were already replicated if the modification time is different for a file on the source and the target cluster. This scenario is unavoidable when the target filesystem does not support setting the modification time, and hence results in a performance issue.

Solution

You can choose one of the following methods to skip the already replicated files during subsequent Ozone replication policy runs.
  • Ignore the modification time to skip replicated files.

    Add the following advanced configuration snippet. The subsequent replication jobs for the specified Ozone replication policies ensure that the already-replicated files are skipped by the jobs. The advanced configuration snippet considers the relative file path, file name, and file size, and ignores the modification time to determine the replicated files to skip in subsequent job runs.
    1. Go to the target Cloudera Manager > Clusters > [*** OZONE SERVICE ***] > Configuration tab.
    2. Search for Ozone Replication Advanced Configuration Snippet (Safety Valve) for core-site.xml.
    3. Add com.cloudera.enterprise.distcp.ozone-schedules-with-unsafe-equality-check = [***ENTER COMMA-SEPARATED LIST OF OZONE REPLICATION POLICIES’ ID or ENTER all TO APPLY TO ALL OZONE REPLICATION POLICIES***] key-value pair.
    4. Save the changes.
  • Consider the modification time to skip replicated files.

    When both of the advanced configuration snippets in the following task are configured, subsequent Ozone replication policy runs skip replicating a file when the relative file path, file name, and file size are equal, and when the source file’s modification time is less than or equal to the target file’s modification time. This reduces the risk of data loss, because if a file is modified on the source after a bootstrap replication job, the source file’s modification time would be higher than the target file’s modification time, therefore the file is not skipped during the subsequent replication job run.

    1. Go to the target Cloudera Manager > Clusters > [*** OZONE SERVICE ***] > Configuration tab.
    2. Search for Ozone Replication Advanced Configuration Snippet (Safety Valve) for core-site.xml.
    3. Add com.cloudera.enterprise.distcp.ozone-schedules-with-unsafe-equality-check = [***ENTER COMMA-SEPARATED LIST OF OZONE REPLICATION POLICIES’ ID or ENTER all TO APPLY TO ALL OZONE REPLICATION POLICIES***] key-value pair.
    4. Add com.cloudera.enterprise.distcp.require-source-before-target-modtime-in-unsafe-equality-check = [***ENTER true OR false***]
    5. Save the changes.

Incremental replication for Ozone replication policies fail with a 403 Forbidden error.

Incremental replication for Ozone replication policies fail with a 403 Forbidden error because of insufficient administrative privileges for the users performing the replication.

Problem

Incremental replication for Ozone replication policies fail with a 403 Forbidden error. The logs display the Only bucket owners and Ozone admins can create snapshots. message indicating that the replication process was unable to create or list snapshots on the source Ozone buckets.

Cause

The failure might occur because of insufficient administrative privileges for the users performing the replication. The following user identities might be affected:

  • Cloudera Manager peer user — The Cloudera Manager peer user configured for the replication job might not have the necessary administrative role to run the privileged API commands on the source cluster.
  • Service user (Hive) — The hive user, specified in the Run as Username and Run on Peer as Username fields in the Ozone replication policy, might not be assigned the Ozone administrator role.
  • Ozone administrator rights — These requests might have been rejected at the Ozone service level because of missing administrator rights. In this scenario, the failures might not appear as Denied entries in the Ranger audit logs.

Solution

  1. Update the Ozone service configurations by adding the hive user to the Ozone Administrator list in Cloudera Manager > Clusters > Ozone service > Configuration > Ozone Administrator.
  2. Add the source cluster as a peer to the target cluster, and select the Create User With Admin Role option.
  3. During the Ozone replication policy creation process, add the hive user in both the Run as Username and Run on Peer as Username fields.