Known Issues in Apache Hive

Learn about the known issues in Hive, the impact or changes to the functionality, and the workaround.

OPSAPS-58664: Hive on Tez LDAP configurations are not pushed to hive-site.xml by Cloudera Manager: After setting up LDAP properties in the Hive on Tez service, the settings are not pushed into hive-site.xml for Hive on Tez service even after a restart. The issue is due to HiveOnTezServiceHandler re-using definitions from HiveConfigFileDefinitions. The definitions are not including any roletypes other than HiveServiceHandler's roletypes.
OPSAPS-59928: INSERT INTO from SELECT using hive (hbase) table returns an error under certain conditions.: Users who upgraded to a Kerberized CDP cluster from HDP and enabled AutoTLS have reported this problem. For more information, see Cloudera Community article: ERROR: "FAILED: Execution Error, return code 2" when the user is unable to issue INSERT INTO from SELECT using hive (hbase) table.; In Cloudera Manager > TEZ > Configurations, find the tez.cluster.additional.classpath.prefix Safety Valve, and set the value to /etc/hbase/conf.
CDPD-21365: Performing a drop catalog operation drops the catalog from the CTLGS table. The DBS table has a foreign key reference on CTLGS for CTLG_NAME. Because of this, the DBS table is locked and creates a deadlock.: You must create an index in the DBS table on CTLG_NAME: CREATE INDEX CTLG_NAME_DBS ON DBS(CTLG_NAME);.
CDPD-26229: Hive TPCDS test query is timing out after 3 hours.: None.
CDPD-26556 After an upgrade, querying a CTAS table under certain conditions might throw an exception: If you upgrade your Hive cluster from CDH 6 to CDP 7, create a CTAS table in the CDP cluster from a table you upgraded from CDH, you might see the following exception when you query the new table:
class org.apache.hadoop.io.IntWritable cannot be cast to class org.apache.hadoop.hive.serde2.objectinspector.StandardUnionObjectInspector$StandardUnion
This issue involves CDH-based tables having columns of complex types ARRAY, MAP, and STRUCT.

OPSAPS-54299 Installing Hive on Tez and HMS in the incorrect order causes HiveServer failure: You need to install Hive on Tez and HMS in the correct order; otherwise, HiveServer fails. You need to install additional HiveServer roles to Hive on Tez, not the Hive service; otherwise, HiveServer fails.; Follow instructions on Installing Hive on Tez.

CDPD-23506: OutOfMemoryError in LLAP: Long running spark-shell applications can leave sessions in interactive Hiveserver2 until the Spark application finishes (user exists from spark-shell), causing memory pressure in case of a high number of queries in the same shell (1000+).; You must close spark-shell so that sessions are closed. Add the owner of the database or the tables as a user with read or read/write access to the tables directly.
CDPD-23041: DROP TABLE on a table having an index does not work: If you migrate a Hive table to CDP having an index, DROP TABLE does not drop the table. Hive no longer supports indexes (HIVE-18448). A foreign key constraint on the indexed table prevents dropping the table. Attempting to drop such a table results in the following error:
java.sql.BatchUpdateException: Cannot delete or update a parent row: a foreign key constraint fails ("hive"."IDXS", CONSTRAINT "IDXS_FK1" FOREIGN KEY ("ORIG_TBL_ID") REFERENCES "TBLS ("TBL_ID")); There are two workarounds:

Drop the foreign key "IDXS_FK1" on the "IDXS" table within the metastore. You can also manually drop indexes, but do not cascade any drops because the IDXS table includes references to "TBLS".

Launch an older version of Hive, such as Hive 2.3 that includes IDXS in the DDL, and then drop the indexes as described in Language Manual Indexing.; Apache Issue: Hive-24815

CDPD-20636 and DWX-6163: SHOW TABLES command does not produce a list of tables that are owned by the current user: When you run the SHOW TABLES command against a Hive Virtual Warehouse, tables are only returned if you have explicit read or read/write access to the table, or if you belong to a group that has read or read/write access. If you only have access to the tables because you are the owner of the objects, you can query the table content, but the table names do not appear in the SHOW TABLES command output.; Add the owner of the database or the tables as a user with read or read/write access to the tables directly.
CDPD-17766: Queries fail when using spark.sql.hive.hiveserver2.jdbc.url.principal in the JDBC URL to invoke Hive.: Do not specify spark.sql.hive.hiveserver2.jdbc.url.principal in the JDBC URL to invoke Hive remotely.; Workaround: specify principal=hive.server2.authentication.kerberos.principal as shown in the following syntax:
jdbc:hive://<host>:<port>/<dbName>;principal=hive.server2.authentication.kerberos.principal;<otherSessionConfs>?<hiveConfs>#<hiveVars>

HIVE-24271: Problem creating an ACID table in legacy table mode: In site-level, legacy CREATE TABLE mode, the CREATE MANAGED TABLE command might not work as expected to override the legacy behavior and create a managed ACID table. The command works only at the session level.; Workaround: Include table properties in a CREATE TABLE that specify a transactional table. For example:
CREATE TABLE T2(a int, b int) STORED AS ORC TBLPROPERTIES ('transactional'='true');
CDPD-13636: Hive job fails with OutOfMemory exception in the Azure DE cluster: Set the parameter hive.optimize.sort.dynamic.partition.threshold=0. Add this parameter in Cloudera Manager (Hive Service Advanced Configuration Snippet (Safety Valve) for hive-site.xml)
CDPD-16802: Autotranslate assertion failure.: The exception is not triggered when it is executed from Spark-Shell. This is from Hive in the getJdoFilterPushdownParam parameter of ExpressionTree.java, which checks the partition column as only String and not any other type.; This can be disabled by setting hive.metastore.integral.jdo.pushdown to true.

HiveServer Web UI displays incorrect data: If you enabled auto-TLS for TLS encryption, the HiveServer2 Web UI does not display the correct data in the following tables: Active Sessions, Open Queries, Last Max n Closed Queries

CDPD-11890: Hive on Tez cannot run certain queries on tables stored in encryption zones: This problem occurs when the Hadoop Key Management Server (KMS) connection is SSL-encrypted and a self signed certificate is used. SSLHandshakeException might appear in Hive logs.; Use one of the workarounds:

Install a self signed SSL certificate into cacerts file on all hosts.

Copy ssl-client.xml to a directory that is available in all hosts. In Cloudera Manager, in Clusters > Hive on Tez > Configuration. In Hive Service Advanced Configuration Snippet for hive-site.xml, click +, and add the name tez.aux.uris and valuepath-to-ssl-client.xml.
OPSAPS-72159: Data corruption issue in PostgreSQL databases after OS upgrade: Upgrading to CentOS 8 or Red Hat Enterprise Linux (RHEL) 8 results in PostgreSQL database corruption due to changes in the GNU C Library (glibc). The corruption affects B-Tree indexes on text columns, causing Hive metadata issues, duplicate tables, and failed drop table operations.
For details and resolution steps, see the Knowledge Base article Data corruption issue in PostgreSQL databases after OS upgrade to CentOS 8 or RHEL 8.

Technical Service Bulletins

TSB 2021-501: JOIN queries return wrong result for join keys with large size in Hive: JOIN queries return wrong results when performing joins on large size keys (larger than 255 bytes). This happens when the fast hash table join algorithm is enabled, which is enabled by default.
Knowledge article: For the latest update on this issue see the corresponding Knowledge article: TSB 2021-501: JOIN queries return wrong result for join keys with large size in Hive

TSB 2021-529: Ranger RMS leads to HMS Connection leak and increased heap memory usage in NameNode process: After enabling Ranger Resource Mapping Service (RMS), RMS connects to Hive MetaStore (HMS) every 30 seconds to fetch the notification event. However, for each request, RMS creates two HMS connections and only closes one of them.
Knowledge article: For the latest update on this issue see the corresponding Knowledge article: TSB 2021-529: Ranger RMS leads to HMS Connection leak and increased heap memory usage in NameNode process

TSB 2021-532: HWC fails to write empty DataFrame to orc files: HWC writes fail when an empty DataFrame write is attempted. That is because the writer does not create an orc file if no records are present in the DataFrame. This causes the HWC write commit validation to fail.
Knowledge article: For the latest update on this issue see the corresponding Knowledge article: TSB 2021-532: HWC fails to write empty DataFrame to orc files

TSB 2022-526: A Hive query may produce wrong results for some vectorized built-in functions with compound expression in PARTITION BY or ORDER BY clause: Vectorized functions with PARTITION BY and/or ORDER BY clauses where the partition or order by expression is compound (example: cast string to integer) and not just a simple column reference may be broken.
The query may fail or output wrong results, depending on the compound expression. For example:

Cast integer to string results in query failure with a NullPointerExpression

Cast string to integer outputs wrong results
Knowledge article: For the latest update on this issue see the corresponding Knowledge article: TSB 2022-526: A Hive query may produce wrong results for some vectorized built-in functions with compound expression in PARTITION BY or ORDER BY clause

TSB 2022-567: Potential Data Loss due to CTLT HBaseStorageHandler failure dropping underlying HBase table while rollback: If the create table target_table like source table command (CTLT) fails and the source table is HBaseStorageHandler-based table, the HBaseMetaHook rollback logic deletes the underlying HBase table, resulting in potential data loss.
Upstream JIRA: HIVE-25989
Knowledge article: For the latest update on this issue, see the corresponding Knowledge article: TSB 2022-567: Potential Data Loss due to CTLT HBaseStorageHandler failure dropping underlying HBase table while rollback

TSB 2023-627: IN/OR predicate on binary column returns wrong result: An IN or an OR predicate involving a binary datatype column may produce wrong results. The OR predicate is converted to an IN due to the setting hive.optimize.point.lookup which is true by default. Only binary data types are affected by this issue. See https://issues.apache.org/jira/browse/HIVE-26235 for example queries which may be affected.
Upstream JIRA: HIVE-26235
Knowledge article: For the latest update on this issue, see the corresponding Knowledge article: TSB 2023-627: IN/OR predicate on binary column returns wrong result