JSON file
The JSON file is an optional component. Cloudera Lakehouse Optimizer uses the threshold values in the file during the evaluation phase to compare the current and expected stats. The action arguments are supported by a corresponding Spark action.
{
"expireSnapshot": {
"expireOlderThan": 432000000,
"retainLast": 5,
"cleanExpiredFiles": true
},
"rewriteManifest" : {
"useCaching": true,
"fileCountMax": 100,
"manifestFileSize": 8388608,
"smallFileRatioMax": 0.5
},
"rewriteDataFiles" : {
"targetFileSize": 536870912,
"maxConcurrentRewriteFileGroups": 5,
"minInputFiles": 5,
"partialProgressEnabled": true,
"partialProgressMaxCommits": 10,
"deleteFileThreshold": 2000000,
"useStartingSequenceNumber": false,
"rewriteAll": false
},
"rewritePositionDelete" : {
"enabled" : false,
"targetFileSize": 67108864,
"maxConcurrentGroupRewrite": 5,
"minInputFiles": 6
"partialProgressMaxCommits": 10,
"partialProgressEnabled": true
},
"deleteOrphanFiles" : {
"olderThan": 259200000
},
"priority": "MEDIUM",
"description": "An example policy constant",
"cron": "0 4 * ? * *"
}compressDataFiles block in the JSON
file:"compressDataFiles": {
"enabled": true,
"method": "zstd",
"level": 12,
"filter": "1 = 0"
}For more information, see Data file compression for Cloudera Lakehouse Optimizer.
| Action argument | Value | Description |
|---|---|---|
| expireSnapshot | ||
enabled |
Default is true | Determines whether to evaluate the actions and generate the action arguments. |
cleanExpiredFiles |
Default is true | Removes the expired snapshots permanently. |
expireOlderThan |
Default is 120 * 3600 * 1000 ms, that is 5
days. Minimum is 10 seconds |
Deletes the snapshot when the snapshot is older than the set
time. For example, a snapshot is deleted after 5 days by default. |
retainLast |
Default is 5 Minimum is 1 |
Deletes the last snapshot when the number of snapshots exceeds the set
value. For example, by default the first snapshot gets deleted automatically after the sixth snapshot is created. |
expireSnapshotId |
No default value | Expires the specified snapshot. |
| rewriteManifest | ||
enabled |
Default is true | Determines whether to evaluate the actions and generate the action arguments. |
useCaching |
Default is true | Uses cache during the rewrite manifest file operation process. |
targetFileSize |
Default is 8388608 bytes | Specifies the target manifest file size in bytes. |
| rewriteDataFiles | ||
enabled |
Default is true | Determines whether to evaluate the actions and generate the action arguments. |
targetFileSize |
Default is 512
MB Minimum is 1 KB Maximum is 64 GB |
Determines the target output file size after compaction. |
maxConcurrentRewriteFileGroups |
Default is 5 Minimum is 1 Maximum is 1000 |
Defines the maximum number of file groups to be simultaneously rewritten. |
minInputFiles |
Default is
5 Minimum is 1 |
Rewrites a file group when the file group exceeds the specified number of files, regardless of other criteria. For example, the number of small files tolerated per partition. |
partialProgressMaxCommits |
Default is
10 Minimum is 1 |
Defines the maximum number of commits that the rewrite action is allowed to commit when partial progress is enabled. |
deleteFileThreshold |
Default is
2000000 Minimum is 1 |
Defines the minimum number of deletes that must be associated with a data file for it to be considered for the rewriting action. |
partialProgressEnabled |
Default is false. | Defines the maximum number of commits that are allowed during the rewrite operation. This ensures that the changes are committed and snapshots are created even while the rewrite operation is in progress. If a table is not updated frequently, retain the value as false. |
use-starting-sequence-number |
Default is false. | Specifies the sequence number of the snapshot at compaction operation start time instead of the newly produced snapshot. |
rewrite-all |
Default is false. | Force rewrites all the files overriding other options. Ensures full compaction of the tables. |
| deleteOrphanFiles | ||
enabled |
Default is true | Determines whether to evaluate the actions and generate the action arguments. |
olderThan in ms |
Default is 72 * 3600 * 1000 that is 3
days. Minimum is 10 in seconds. |
Removes orphan files created before the specified time. |
| rewritePositionDelete | ||
enabled |
Default is true | Determines whether to evaluate the actions and generate the action arguments. |
targetFileSize |
Default is 64 MB Minimum is 1 KB |
Determines the target output file size after the rewrite positional delete operation. |
maxConcurrentGroupRewrite |
Default is 5 Minimum is 1 Maximum is 1000 |
Defines the maximum number of file groups to be simultaneously rewritten. |
minInputFiles
|
Default is
5 Minimum is 1 |
Rewrites a file group when the file group exceeds the specified number of files regardless of other criteria. |
partial-progress.max-commits |
Default is 10 | Defines the maximum number of commits that are allowed during the rewrite operation. This ensures that the changes are committed and snapshots are created even while the rewrite operation is in progress. |
partialProgressEnabled |
Default is true | Enables committing groups of files before the rewrite operation completes. For more information, see Partial Progress Enabled. |
rewrite-job-order |
No default value | Forces the rewrite job order based on the chosen value. You can choose one
of the following values:
|
| computeTableStats | ||
enabled |
No default; must be true to run the action | Opt-in gate for the compute-table-stats action that writes Puffin statistics to table metadata. |
columns |
No default; omit for all columns | Array of column names to compute NDV statistics for. When omitted or empty, Cloudera Lakehouse Optimizer computes statistics for all primitive columns in the table schema. Maximum of 512 columns, column names must not exceed 1024 characters, and duplicate names are rejected. |
snapshotId |
No default; omit for current snapshot | Snapshot ID to compute statistics against. |
| compressDataFiles | ||
enabled |
Default is false | Must be true to run the data file compression action. |
filter |
No default; required when enabled | SQL WHERE clause that scopes which rows or files to rewrite. Use 1 =
0 at the namespace level to opt in tables individually with ALTER
TABLE … SET TBLPROPERTIES. Cloudera Lakehouse Optimizer skips the
action when the filter is absent or blank. |
method |
Default is the table format default | Compression codec. For supported values, see Supported codecs for data file compression. |
level |
Default is the codec default | Codec-specific compression level. |
validateFilterColumns |
Default is true | When true, Cloudera Lakehouse Optimizer validates that filter column names exist in the table schema before dispatching the task. |
| priority | ||
priority |
Default is Medium Valid values are Critical, Urgent, High, Medium, Low, and Minimal. |
Defines the priority level for maintenance tasks that are configured for a policy. Cloudera Lakehouse Optimizer schedules higher-priority tasks before lower-priority tasks when multiple policies compete for cluster resources. |
Puffin statistics policy example
Puffin statistics require a separate custom policy. Do not add
computeTableStats to the default ClouderaAdaptive policy constants.
Use a dedicated JEXL script and JSON constants file. The action runs only when
computeTableStats.enabled is true in the JSON constants
(the default is false for this action).
JEXL
// Dry-run policy: compute table stats using the same merge rules as the former ClouderaAdaptive.jexl block.
// Opt-in via computeTableStats.enabled in policy constants (default false per PolicyConstants for this action).
#pragma dlm.cron "0 0 13 ? * *"
#pragma dlm.statistics false
const actions = [...]
if ($constants.computeTableStats.enabled) {
const computeStats = dlm:computeTableStats($table);
let colsTable = $table['dlm.computeTableStats.columns'];
if (colsTable != null && colsTable != '') {
computeStats.addColumns(colsTable);
} else if (colsTable == null && $constants.computeTableStats.columns != null) {
computeStats.addColumns($constants.computeTableStats.columns);
}
computeStats.validate($table);
let resolvedSnapshotId = $table['dlm.computeTableStats.snapshotId'] ?? $constants.computeTableStats.snapshotId;
if (resolvedSnapshotId != null) {
computeStats.snapshot(resolvedSnapshotId);
}
actions.add(computeStats);
}
actions
JSON
{
"computeTableStats": {
"enabled": true,
"columns": ["id", "data"],
"snapshotId": 424242
},
"description": "Constants for compute table stats"
}
Omit snapshotId to use the current snapshot. Upload the JEXL and JSON files
when you define the policy resources. For more information, see Defining Cloudera
Lakehouse Optimizer resources using REST APIs.
