Support for Trino Lakehouse connector
The Lakehouse connector provides a unified interface to query and write data stored across multiple open table formats and metastore services. You can use a single catalog to manage tables in Apache Iceberg, and Apache Hive formats without maintaining separate connector configurations.
Overview of the Trino Lakehouse connector
The Lakehouse connector combines the features of the Hive, Iceberg, Delta Lake, and Hudi connectors into a single unified connector. Configuration properties, session properties, table properties, and behaviors originate from the underlying engine connectors.
The Lakehouse connector supports connectivity to the Hive Metastore (HMS) for metastore management in deployments.
Lakehouse connector configuration properties
To configure the Lakehouse connector within the DWX Federation Connectors, use the following configuration property structure. This configuration integrates the unified metastore, native storage caching layer, Ranger authorization, Kerberos authentication, and inherited Hive engine behaviors:
# 1. Core Identity
connector.name=lakehouse
lakehouse.table-type=ICEBERG
# 2. Unified Metastore Connection
hive.metastore.uri=thrift://metastore-service.{{ .Values.warehouseId }}.svc.cluster.local:9083
# 3. Consolidated Security
hive.security={{ .Values.authorizationMode }}
# 4. Storage & Caching Layer (Native FileSystem SPI)
fs.cache.enabled=true
fs.cache.directories=/data/trino/caches/lakehouse
fs.cache.max-disk-usage-percentages=30
fs.cache.ttl=7d
# 5. Inherited Hive Engine Behaviors
hive.collect-column-statistics-on-write=false
hive.non-managed-table-writes-enabled=true
hive.temporary-staging-directory-enabled={{ if and .Values.isPrivateCloud .Values.ozone .Values.ozone.enabled }}false{{ else }}true{{ end }}
# 6. Ranger Properties
ranger.hadoop.config.resource=/etc/trino/core-site.xml
ranger.plugin.config.resource=/etc/trino/ranger-trino-security.xml,/etc/trino/ranger-trino-audit.xml,/etc/trino/ranger-trino-policymgr-ssl.xml
ranger.service.name={{ .Values.rangerHiveSvcName }}
# 7. Kerberos default values
hive.metastore.authentication.type={{ .Values.hmsAuthenticationType }}
hive.metastore.client.keytab=/etc/security/keytabs/hive.service.keytab
hive.metastore.client.principal={{ .Values.krbPrincipal }}
hive.metastore.service.principal={{ .Values.metastoreKrbPrincipal }}
Limitations of the Trino Lakehouse connector
The Trino Lakehouse connector has specific setup requirements when deployed on clusters enabled with Kerberos and Apache Ranger.
- Kerberos properties requirement for UI catalog creation - Creating a Lakehouse catalog by using the UI on a cluster enabled with Kerberos authentication and Apache Ranger authorization requires manual property configuration. The UI does not automatically set the default Kerberos configuration values.Workaround: Manually add the required Kerberos configuration properties to the catalog configuration when creating a Lakehouse catalog by using the UI.
Configuration parameters reference
The Lakehouse connector requires a Hive Metastore. The hive.metastore.uri property simultaneously defines the Iceberg catalog location; therefore, you must not set iceberg.catalog.type.
The following table shows the configuration property for the Lakehouse connector:
| Parameter | Description | Default |
|---|---|---|
| lakehouse.table-type | Specifies the default table type for newly created tables when the type table property is omitted. Valid values: HIVE, ICEBERG. | ICEBERG |
Feature support matrix
The following table shows the capabilities of the Lakehouse connector:
| Category | Feature | Status | Details |
|---|---|---|---|
| Engine Interop | Apache Iceberg | Full | Supports v1 and DML-heavy v2 tables with row-level modifications. |
| Engine Interop | Delta Lake | Full | Supports deletion vectors and advanced indexing. |
| Engine Interop | Apache Hive | Full | Processes Parquet, ORC, Avro, and Text storage formats. |
| Engine Interop | Apache Hudi | Read-Only | Restricted to Snapshot and Read-Optimized queries. |
| Metastores | Self-Hosted HMS | Full | Direct Thrift metadata access for setups. |
| SQL Capabilities | SELECT / JOIN / DML | Full | Cross-format table joins and routed DML commands (INSERT, UPDATE, DELETE). |
Creating tables using the Lakehouse connector
When creating a table through the Lakehouse connector, specify the target table format by setting the type property in the WITH clause.
To create an Iceberg table, run the following SQL command:
CREATE TABLE iceberg_table (
c1 INTEGER,
c2 DATE,
c3 DOUBLE
)
WITH (
type = 'ICEBERG',
format = 'PARQUET',
partitioning = ARRAY['c1', 'c2'],
sorted_by = ARRAY['c3']
);
To create a Hive table, run the following SQL command:
CREATE TABLE hive_page_views (
view_time TIMESTAMP,
user_id BIGINT,
page_url VARCHAR,
ds DATE,
country VARCHAR
)
WITH (
type = 'HIVE',
format = 'ORC',
partitioned_by = ARRAY['ds', 'country'],
bucketed_by = ARRAY['user_id'],
bucket_count = 50
);
