Copy curated HDFS directories from a source cluster to Apache Ozone on a target
cluster without
altering
source data, using Cloudera Replication Manager-backed executions.
Planning completes with Cloudera Replication Manager replication policies copying
every selected directory into deployed Ozone layout targets.
The
source
HDFS data remains intact.
Both source (HDFS) and target (Ozone) clusters
must
be displayed on the
Clusters
page.
For
instructions, see
Registering
clusters.
The target cluster must
host
Apache Ozone with the service running.
You
must
model
and deploy Ozone volumes, buckets, and folders in the
Storage
Layout
Editor.
Without deployed
layout destinations, executions have nowhere to persist data.
For
instructions, see
Managing Ozone
storage.
CMA Agents
must operate on source and target clusters.
Cloudera Replication Manager
must
be reachable. Cloudera Migration Assistant submits policies
through Cloudera Replication Manager APIs.
Secure clusters authenticate with
Kerberos, so you
must
confirm that
the migration user principal works for filesystem access.
Whenever
Discovery trees look stale before
mapping, rerun the steps
in Scanning
clusters.
Collections exported from scans help bound execution scopes
. For more information, see
Creating collections for migration if you segregate phased cutovers.
Step 1 — Scan the
source
cluster
(HDFS)
Open Clusters, then
select
the source cluster.
Click Start Scanning.
Select
Hdfs data scan.
Optional: Constrain
the scan path,
for example,
/tmp or
/user,
to shorten runs on massive namespaces.
Click Scan selected and watch the progress drawer until
completion.
After success,
Discovery lists HDFS trees with
sizing and ownership snapshots.
Step 2 — Scan the
target
cluster
(Ozone)
Confirm Apache Ozone is healthy in Cloudera Manager before scanning.
From
Clusters,
open the target cluster.
Click Start Scanning.
Select Ozone data scan.
Click Scan selected and wait until the status turns green.
Set Name, select the source HDFS cluster, and the target Ozone
cluster.
Complete the wizard
(Create) and land on the plan detail
surface.
Step 4 — Map HDFS
directories
to Ozone
paths
on the
Data
tab
The
Data tab
displays
source HDFS on the left and target storage on the right.
Go to
Plan > Data.
Switch the target browser mode from HDFS to Ozone using the service picker.
Go to your target volume, bucket, and folder staging area.
On the source side, go to the parent directory that contains data you want to migrate.
Hover the directory and click
Copy.
Cloudera Migration Assistant persists a structural mapping.
Switch to
the Review
view any
time to
view a list of your
current
structure
mappings.
Optional: When Ozone misses a folder UI-side, create it before copying structures.
In the target browser, go to the parent bucket or folder.
Click Create Folder.
Provide the folder name and confirm creation.
Cloudera Migration Assistant updates its metadata
immediately
and
executions push the filesystem change during processing.
Repeat the mapping workflow for additional directories whenever you broaden scope.
You can
map
each directory to
separate
buckets or combine
them
into shared
buckets.
The
layout must
meet the
Cloudera Replication Manager requirements for your tenancy model.
Step 5 — Create and
run
an
execution
Open Execution.
Select
Create Execution.
Complete
the steps in the
configuration
wizard (General, Scope,
Review),
and
then click
Finish.