Configuring HDFS File System Check read timeout

When running HDFS File System Check (hdfs fsck) commands on large clusters or directories with a high volume of corrupt blocks, the command may fail with a java.net.SocketTimeoutException: Read timed out error. This error occurs when the default client timeout period is insufficient for the NameNode to process the request and return results. You can prevent these timeouts by manually configuring the dfs.client.fsck.read.timeout property in Cloudera Manager.

Commands such as hdfs fsck / -list-corruptfileblocks or hdfs fsck / -delete fail with the following exception:

Exception in thread "main" java.net.SocketTimeoutException: Read timed out
    at java.net.SocketInputStream.socketRead0(Native Method)
    ...
    at org.apache.hadoop.hdfs.tools.DFSck.doWork(DFSck.java:381)

To resolve this issue, configure dfs.client.fsck.read.timeout property and increase the read timeout.

  1. Sign in to Cloudera Manager.
  2. In the left navigation, click Clusters.
  3. Click the HDFS cluster.
  4. Click Configuration.
  5. Search for HDFS Client Advanced Configuration Snippet (Safety Valve) for hdfs-site.xml.
  6. Click the + icon and add the following property:
    • Name: dfs.client.fsck.read.timeout
    • Value: 1200000

      Value is in milliseconds; 1200000 represents 20 minutes. Adjust this value based on your cluster size and requirements.

  7. Click Save Changes.
  8. Restart the HDFS service and deploy the client configuration.