Known Issues in Apache Kafka

Learn about the known issues in Kafka, the impact or changes to the functionality, and the workaround.

CDPD-60862: Rolling restart fails during ZDU when DDL operations are in progress

During a Zero Downtime Upgrade (ZDU), the rolling restart of services that support Data Definition Language (DDL) statements might fail if DDL operations are in progress during the upgrade. As a result, ensure that you do not run DDL statements during ZDU.

The following services support DDL statements:

Impala
Hive – using HiveQL
Spark – using SparkSQL
HBase
Phoenix
Kafka

Data Manipulation Lanaguage (DML) statements are not impacted and can be used during ZDU. Following the successful upgrade, you can resume running DDL statements.

None. Cloudera recommends modifying applications to not use DDL statements for the duration of the upgrade. If the upgrade is already in progress, and you have experienced a service failure, you can remove the DDLs in-flight and resume the upgrade from the point of failure.

CDPD-60489: Jackson-dataformat-yaml 2.12.7 and Snakeyaml 2.0 are not compatible.

You must not use Jackson-dataformat-yaml through the platform for YAML parsing.

OPSAPS-59553: SMM's bootstrap server config should be updated based on Kafka's listeners

SMM does not show any metrics for Kafka or Kafka Connect when multiple listeners are set in Kafka.

SMM cannot identify multiple listeners and still points to bootstrap server using the default broker port (9093 for SASL_SSL). You would have to override bootstrap server URL (hostname:port as set in the listeners for broker) in the following path:

Cloudera Manager > SMM > Configuration > Streams Messaging Manager Rest Admin Server Advanced Configuration Snippet (Safety Valve) for streams-messaging-manager.yaml > Save Changes > Restart SMM.

The offsets.topic.replication.factor property must be less than or equal to the number of live brokers: The offsets.topic.replication.factor broker configuration is now enforced upon auto topic creation. Internal auto topic creation will fail with a GROUP_COORDINATOR_NOT_AVAILABLE error until the cluster size meets this replication factor requirement.; None

Requests fail when sending to a nonexistent topic with auto.create.topics.enable set to true: The first few produce requests fail when sending to a nonexistent topic with auto.create.topics.enable set to true.; Increase the number of retries in the producer configuration setting retries.

Performance degradation when SSL Is enabled: In some configuration scenarios, significant performance degradation can occur when SSL is enabled. The impact varies depending on your CPU, JVM version, Kafka configuration, and message size. Consumers are typically more affected than producers.; Configure brokers and clients with ssl.secure.random.implementation = SHA1PRNG. It often reduces this degradation drastically, but its effect is CPU and JVM dependent.

OPSAPS-43236: Kafka garbage collection logs are written to the process directory: By default Kafka garbage collection logs are written to the agent process directory. Changing the default path for these log files is currently unsupported.; None

RANGER-3809: Idempotent Kafka producer fails to initialize due to an authorization failure: Kafka producers that have idempotence enabled require the Idempotent Write permission to be set on the cluster resource in Ranger. If permission is not given, the client fails to initialize and an error similar to the following is thrown:
org.apache.kafka.common.KafkaException: Cannot execute transactional method because we are in an error state at org.apache.kafka.clients.producer.internals.TransactionManager.maybeFailWithError(TransactionManager.java:1125) at org.apache.kafka.clients.producer.internals.TransactionManager.maybeAddPartition(TransactionManager.java:442) at org.apache.kafka.clients.producer.KafkaProducer.doSend(KafkaProducer.java:1000) at org.apache.kafka.clients.producer.KafkaProducer.send(KafkaProducer.java:914) at org.apache.kafka.clients.producer.KafkaProducer.send(KafkaProducer.java:800) . . . Caused by: org.apache.kafka.common.errors.ClusterAuthorizationException: Cluster authorization failed.
Idempotence is enabled by default for clients in Kafka 3.0.1, 3.1.1, and any version after 3.1.1. This means that any client updated to 3.0.1, 3.1.1, or any version after 3.1.1 is affected by this issue.; This issue has two workarounds, do either of the following:

Explicitly disable idempotence for the producers. This can be done by setting enable.idempotence to false.

Update your policies in Ranger and ensure that producers have Idempotent Write permission on the cluster resource.
CDPD-45183: Kafka Connect active topics might be visible to unauthorised users: The Kafka Connect active topics endpoint (/connectors/[***CONNECTOR NAME***]/topics) and the Connect Cluster page on the SMM UI disregard the user permissions configured for the Kafka service in Ranger. As a result, all active topics of connectors might become visible to users who do not have permissions to view them. Note that user permission configured for Kafka Connect in Ranger are not affected by this issue and are correctly applied.; None.
CDPD-49304: AvroConverter does not support composite default values: AvroConverter cannot handle schemas containing a STRUCT type default value.; None.
DBZ-4990: The Debezium Db2 Source connector does not support schema evolution: The Debezium Db2 Source connector does not support the evolution (updates) of schemas. In addition, schema change events are not emitted to the schema change topic if there is a change in the schema of a table that is in capture mode. For more information, see DBZ-4990.; None.
CFM-3532: The Stateless NiFi Source, Stateless NiFi Sink, and HDFS Stateless Sink connectors cannot use Snappy compression: This issue only affects Stateless NiFi Source and Sink connectors if the connector is running a dataflow that uses a processor that uses Hadoop libraries and is configured to use Snappy compression. The HDFS Stateless Sink connector is only affected if the Compression Codec or Compression Codec for Parquet properties are set to SNAPPY.
If you are affected by this issue, errors similar to the following will be present in the logs.
Failed to write to HDFS due to java.lang.UnsatisfiedLinkError: org.apache.hadoop.util.NativeCodeLoader.buildSupportsSnappy()
Failed to write to HDFS due to java.lang.RuntimeException: native snappy library not available: this version of libhadoop was built without snappy support.; Download and deploy missing libraries. Show Me How

important
Ensure that you complete steps 1-11 on all Kafka Connect hosts. Additionally, ensure that the advanced configuration snippet in step 12 is configured for all Kafka Connect role instances.

Create the /opt/nativelibs directory.
mkdir /opt/nativelibs

Change the owner to kafka.
chown kafka:kafka /opt/nativelibs

Locate the directory containing the Hadoop native libraries and copy its contents to the directory you created.
cp /opt/cloudera/parcels/CDH/lib/hadoop/lib/native/* /opt/nativelibs

Verify that libsnappy.so was copied to the directory you created.

Remove the following from /opt/nativelibs.
libhadoop.a libhadoop.so libhadoop.so.1.0.0

Run the following command.
hadoop version
The command returns the Hadoop version running in the cluster. Note down the first three digits in the version.

Go to https://archive.apache.org/dist/hadoop/common/ and download the Hadoop version that matches the first three digits of the version running in the cluster.
For example, if your Hadoop version is 3.1.1.7.1.9.0-296, then you need to download Hadoop 3.1.1.

Extract the downloaded archive.

Copy the following libraries from the downloaded archive to /opt/nativelibs on the cluster host.
libhadoop.a libhadoop.so.1.0.0
The libraries are located in hadoop-[***VERSION***]/lib/native.

Create a symlink named libhadoop.so and point it to /opt/nativelibs/libhadoop.so.1.0.0.
ln -s /opt/nativelibs/libhadoop.so.1.0.0 /opt/nativelibs/libhadoop.so

Change the owner of every entry within /opt/nativelibs to kafka.
chown -h kafka:kafka /opt/nativelibs/*

In Cloudera Manager, go to Kafka service > Configuration.

Add the following key-value pair to Kafka Connect Environment Advanced Configuration Snippet (Safety Valve).

Key: LD_LIBRARY_PATH

Value: /opt/nativelibs

Click Save Changes.

Restart the Kafka service.
OPSAPS-69481: Some Kafka Connect metrics missing from Cloudera Manager due to conflicting definitions: The metric definitions for kafka_connect_connector_task_metrics_batch_size_avg and kafka_connect_connector_task_metrics_batch_size_max in recent Kafka CSDs conflict with previous definitions in other CSDs. This prevents Cloudera Manager from registering these metrics. It also results in SMM returning an error. The metrics also cannot be monitored in Cloudera Manager chart builder or queried using the Cloudera Manager API.; Contact Cloudera support for a workaround.

OPSAPS-71258:Zstandard and Snappy compression do not support /tmp mounted as noexec: Kafka cannot process messages compressed with Zstandard or Snappy if /tmp is mounted as noexec.; The workaround steps for Zstandard compression are the following. You need to complete these steps for all the collected hosts.

Identify the hosts running Kafka.

Verify the Zstandard version by checking the classpath of the running process.
ps aux
By default, Kafka uses version libzstd-jni-1.5.2-1.
The Zstandard JAR file is located at /opt/cloudera/parcels/CDH/jars/zstd-jni-1.5.2-1.jar. This file includes the architecture and OS-specific binaries. Check Zstandard JNI GitHub repository for the full list of supported native libraries.

Extract the binary corresponding to the target system similar to the following example.
/usr/java/default/bin/jar xf /opt/cloudera/parcels/CDH/jars/zstd-jni-1.5.2-1.jar linux/amd64/libzstd-jni-1.5.2-1.so

Copy the extracted binary to a location included in java.library.path similar to the following example.
cp linux/amd64/libzstd-jni-1.5.2-1.so /lib

Restart Kafka.

The workaround steps for Snappy compression are the following:

In Cloudera Manager, select the Kafka service and go to Configuration.

Add the following line to Additional Broker Java Options.
-org.xerial.snappy.tempdir==[***PATH***]
Where [***PATH***] is a directory that is not /tmp.

Restart Kafka.

Unsupported Features🔗

: The following Kafka features are not supported in Cloudera Data Platform:

Only Java and .Net based clients are supported. Clients developed with C, C++, Python, and other languages are currently not supported.

The Kafka default authorizer is not supported. This includes setting ACLs and all related APIs, broker functionality, and command-line tools.

SASL/SCRAM is only supported for delegation token based authentication. It is not supported as a standalone authentication mechanism.

Kafka KRaft in this release of Cloudera Runtime is in technical preview and does not support the following:

Deployments with multiple log directories. This includes deployments that use JBOD for storage.

Delegation token based authentication.

Migrating an already running Kafka service from ZooKeeper to KRaft.

Atlas Integration.

Limitations🔗

Collection of Partition Level Metrics May Cause Cloudera Manager’s Performance to Degrade: If the Kafka service operates with a large number of partitions, collection of partition level metrics may cause Cloudera Manager's performance to degrade.

If you are observing performance degradation and your cluster is operating with a high number of partitions, you can choose to disable the collection of partition level metrics.
important
If you are using SMM to monitor Kafka or Cruise Control for rebalancing Kafka partitions, be aware that both SMM and Cruise Control rely on partition level metrics. If partition level metric collection is disabled, SMM will not be able to display information about partitions. In addition, Cruise Control will not operate properly.

Complete the following steps to turn off the collection of partition level metrics:

Obtain the Kafka service name:

In Cloudera Manager, Select the Kafka service.

Select any available chart, and select Open in Chart Builder from the configuration icon drop-down.

Find $SERVICENAME= near the top of the display.
The Kafka service name is the value of $SERVICENAME.

Turn off the collection of partition level metrics:

Go to Hosts > Hosts Configuration.

Find and configure the Cloudera Manager Agent Monitoring Advanced Configuration Snippet (Safety Valve) configuration property.
Enter the following to turn off the collection of partition level metrics:
[KAFKA_SERVICE_NAME]_feature_send_broker_topic_partition_entity_update_enabled=false
Replace [KAFKA_SERVICE_NAME] with the service name of Kafka obtained in step 1. The service name should always be in lower case.

Click Save Changes.

Known Issues in Apache Kafka

Unsupported Features🔗

Limitations🔗

We want your opinion

How can we improve this page?