Skip to main content
Version: 1.6.0

Dremio

Overview​

Dremio is a DataHub utility or metadata-focused integration.

The DataHub integration for Dremio covers metadata entities and operational objects relevant to this connector. It also captures table- and column-level lineage, usage statistics, data profiling, ownership, and stateful deletion detection.

Concept Mapping​

Source ConceptDataHub ConceptNotes
Physical Dataset/TableDatasetSubtype: Table
Virtual Dataset/ViewsDatasetSubtype: View
SpacesContainerMapped to DataHub’s Container aspect. Subtype: Space
FoldersContainerMapped as a Container in DataHub. Subtype: Folder
SourcesContainerRepresented as a Container in DataHub. Subtype: Source

Module dremio​

Certified

Important Capabilities​

CapabilityStatusNotes
Asset Containers✅Enabled by default. Supported for types - Dremio Space, Dremio Source.
Column-level Lineage✅Extract column-level lineage. Supported for types - Table.
Data Profiling✅Optionally enabled via configuration.
Dataset Usage✅Enabled by default to get usage stats.
Descriptions✅Enabled by default.
Detect Deleted Entities✅Enabled by default via stateful ingestion.
Domains✅Supported via the domain config field.
Extract Ownership✅Enabled by default.
Operation Capture✅Optionally enabled via include_query_lineage; generated from Dremio job history.
Platform Instance✅Enabled by default.
Table-Level Lineage✅Enabled by default. Supported for types - Table.

Overview​

The dremio module ingests metadata from Dremio into DataHub. It is intended for production ingestion workflows and module-specific capabilities are documented below.

This plugin integrates with Dremio to extract and ingest metadata into DataHub. The following types of metadata are extracted:

  • Metadata for Spaces, Folders, Sources, and Datasets:
    • Includes physical and virtual datasets, with detailed information about each dataset.
    • Extracts metadata about Dremio's organizational hierarchy: Spaces (top-level), Folders (sub-level), and Sources (external data connections). *Schema and Column Information:
    • Column types and schema metadata associated with each physical and virtual dataset.
    • Extracts column-level metadata, such as names, data types, and descriptions, if available.
  • Lineage Information:
    • Dataset-level and column-level lineage tracking:
      • Dataset-level lineage shows dependencies and relationships between physical and virtual datasets.
      • Column-level lineage tracks transformations applied to individual columns across datasets.
    • Lineage information helps trace the flow of data and transformations within Dremio.
  • Ownership and Glossary Terms:
    • Metadata related to ownership of datasets, extracted from Dremio’s ownership model.
    • Glossary terms and business metadata associated with datasets, providing additional context to the data.
    • Note: Ownership information will only be available for the Cloud and Enterprise editions, it will not be available for the Community edition.
  • Optional SQL Profiling (if enabled):
    • Table, row, and column statistics can be profiled and ingested via optional SQL queries.
    • Extracts statistics about tables and columns, such as row counts and data distribution, for better insight into the dataset structure.

Prerequisites​

Before running ingestion, ensure network connectivity to the source, valid authentication credentials, and read permissions for metadata APIs required by this module.

  1. Generate an API Token:

    • Log in to your Dremio instance.
    • Navigate to your user profile in the top-right corner.
    • Select Generate API Token to create an API token for programmatic access.
  2. Permissions:

    • The token should have read-only or admin permissions that allow it to:
      • View all datasets (physical and virtual).
      • Access all spaces, folders, and sources.
      • Retrieve dataset and column-level lineage information.
  3. Verify External Data Source Permissions:

    • If Dremio is connected to external data sources (e.g., AWS S3, relational databases), ensure that Dremio has access to the credentials required for querying those sources.

Install the Plugin​

pip install 'acryl-datahub[dremio]'

Starter Recipe​

Check out the following recipe to get started with ingestion! See below for full configuration options.

For general pointers on writing and running a recipe, see our main recipe guide.

source:
type: dremio
config:
# Coordinates
hostname: localhost
port: 9047
tls: true

# Credentials with personal access token(recommended)
authentication_method: PAT
password: pass
# OR Credentials with basic auth
# authentication_method: password
# username: user
# password: pass

#For cloud instance
#is_dremio_cloud: True
#dremio_cloud_project_id: <project_id>

include_query_lineage: True

ingest_owner: true

#Optional
source_mappings:
- platform: s3
source_name: samples

#Optional
schema_pattern:
allow:
- "<source_name>.<table_name>"

sink:
# sink configs

Config Details​

Note that a . is used to denote nested fields in the YAML recipe.

FieldDescription
authentication_method
One of string, null
Authentication method: 'password' or 'PAT' (Personal Access Token)
Default: PAT
bucket_duration
Enum
One of: "DAY", "HOUR"
disable_certificate_verification
One of boolean, null
Disable TLS certificate verification
Default: False
domain
One of string, null
Domain for all source objects.
Default: None
dremio_cloud_project_id
One of string, null
ID of Dremio Cloud Project. Found in Project Settings in the Dremio Cloud UI
Default: None
dremio_cloud_region
Enum
One of: "US", "EU"
Default: US
end_time
string(date-time)
Latest date of lineage/usage to consider. Default: Current time in UTC
hostname
One of string, null
Hostname or IP Address of the Dremio server
Default: None
include_query_lineage
boolean
Whether to include query-based lineage information.
Default: False
ingest_owner
boolean
Ingest Owner from source. This will override Owner info entered from UI
Default: True
is_dremio_cloud
boolean
Whether this is a Dremio Cloud instance
Default: False
max_workers
integer
Number of worker threads to use for parallel processing
Default: 50
password
One of string(password), null
Dremio password or Personal Access Token
Default: None
path_to_certificates
string
Path to SSL certificates
Default: /private/tmp/datahub-v1.6.0.1rc2/metadata-ingestio...
platform_instance
One of string, null
The instance of the platform that all assets produced by this recipe belong to. This should be unique within the platform. See https://docs.datahub.com/docs/platform-instances/ for more details.
Default: None
port
integer
Port of the Dremio REST API
Default: 9047
start_time
string(date-time)
Earliest date of lineage/usage to consider. Default: Last full day in UTC (or hour, depending on bucket_duration). You can also specify relative time with respect to end_time such as '-7 days' Or '-7d'.
Default: None
tls
boolean
Whether the Dremio REST API port is encrypted
Default: True
username
One of string, null
Dremio username
Default: None
env
string
The environment that all assets produced by this connector belong to
Default: PROD
dataset_pattern
AllowDenyPattern
A class to store allow deny regexes
dataset_pattern.ignoreCase
One of boolean, null
Whether to ignore case sensitivity during pattern matching.
Default: True
profile_pattern
AllowDenyPattern
A class to store allow deny regexes
profile_pattern.ignoreCase
One of boolean, null
Whether to ignore case sensitivity during pattern matching.
Default: True
schema_pattern
AllowDenyPattern
A class to store allow deny regexes
schema_pattern.ignoreCase
One of boolean, null
Whether to ignore case sensitivity during pattern matching.
Default: True
source_mappings
One of array, null
Mappings from Dremio sources to DataHub platforms and datasets.
Default: None
source_mappings.DremioSourceMapping
DremioSourceMapping
source_mappings.DremioSourceMapping.platform ❓
string
Source connection made by Dremio (e.g. S3, Snowflake)
source_mappings.DremioSourceMapping.source_name ❓
string
Alias of platform in Dremio connection
source_mappings.DremioSourceMapping.platform_instance
One of string, null
The instance of the platform that all assets produced by this recipe belong to. This should be unique within the platform. See https://docs.datahub.com/docs/platform-instances/ for more details.
Default: None
source_mappings.DremioSourceMapping.env
string
The environment that all assets produced by this connector belong to
Default: PROD
usage
BaseUsageConfig
usage.bucket_duration
Enum
One of: "DAY", "HOUR"
usage.end_time
string(date-time)
Latest date of lineage/usage to consider. Default: Current time in UTC
usage.format_sql_queries
boolean
Whether to format sql queries
Default: False
usage.include_operational_stats
boolean
Whether to display operational stats.
Default: True
usage.include_read_operational_stats
boolean
Whether to report read operational stats. Experimental.
Default: False
usage.include_top_n_queries
boolean
Whether to ingest the top_n_queries.
Default: True
usage.start_time
string(date-time)
Earliest date of lineage/usage to consider. Default: Last full day in UTC (or hour, depending on bucket_duration). You can also specify relative time with respect to end_time such as '-7 days' Or '-7d'.
Default: None
usage.top_n_queries
integer
Number of top queries to save to each table.
Default: 10
usage.user_email_pattern
AllowDenyPattern
A class to store allow deny regexes
usage.user_email_pattern.ignoreCase
One of boolean, null
Whether to ignore case sensitivity during pattern matching.
Default: True
profiling
ProfileConfig
profiling.enabled
boolean
Whether profiling should be done.
Default: False
profiling.include_field_distinct_count
boolean
Whether to profile for the number of distinct values for each column.
Default: True
profiling.include_field_distinct_value_frequencies
boolean
Whether to profile for distinct value frequencies.
Default: False
profiling.include_field_histogram
boolean
Whether to profile for the histogram for numeric fields.
Default: False
profiling.include_field_max_value
boolean
Whether to profile for the max value of numeric columns.
Default: True
profiling.include_field_mean_value
boolean
Whether to profile for the mean value of numeric columns.
Default: True
profiling.include_field_min_value
boolean
Whether to profile for the min value of numeric columns.
Default: True
profiling.include_field_null_count
boolean
Whether to profile for the number of nulls for each column.
Default: True
profiling.include_field_quantiles
boolean
Whether to profile for the quantiles of numeric columns.
Default: False
profiling.include_field_sample_values
boolean
Whether to profile for the sample values for all columns.
Default: True
profiling.include_field_stddev_value
boolean
Whether to profile for the standard deviation of numeric columns.
Default: True
profiling.limit
One of integer, null
Max number of documents to profile. By default, profiles all documents.
Default: None
profiling.max_workers
integer
Number of worker threads to use for profiling. Set to 1 to disable.
Default: 50
profiling.method
Enum
One of: "ge", "sqlalchemy"
Default: sqlalchemy
profiling.offset
One of integer, null
Offset in documents to profile. By default, uses no offset.
Default: None
profiling.profile_table_level_only
boolean
Whether to perform profiling at table-level only, or include column-level profiling as well.
Default: False
profiling.query_timeout
integer
Time before cancelling Dremio profiling query
Default: 300
profiling.operation_config
OperationConfig
profiling.operation_config.lower_freq_profile_enabled
boolean
Whether to do profiling at lower freq or not. This does not do any scheduling just adds additional checks to when not to run profiling.
Default: False
profiling.operation_config.profile_date_of_month
One of integer, null
Number between 1 to 31 for date of month (both inclusive). If not specified, defaults to Nothing and this field does not take affect.
Default: None
profiling.operation_config.profile_day_of_week
One of integer, null
Number between 0 to 6 for day of week (both inclusive). 0 is Monday and 6 is Sunday. If not specified, defaults to Nothing and this field does not take affect.
Default: None
stateful_ingestion
One of StatefulStaleMetadataRemovalConfig, null
Default: None
stateful_ingestion.enabled
boolean
Whether or not to enable stateful ingest. Default: True if a pipeline_name is set and either a datahub-rest sink or datahub_api is specified, otherwise False
Default: False
stateful_ingestion.fail_safe_threshold
number
Prevents large amount of soft deletes & the state from committing from accidental changes to the source configuration if the relative change percent in entities compared to the previous state is above the 'fail_safe_threshold'.
Default: 75.0
stateful_ingestion.remove_stale_metadata
boolean
Soft-deletes the entities present in the last successful run but missing in the current run with stateful_ingestion enabled.
Default: True

Capabilities​

Use the Important Capabilities table above as the source of truth for supported features and whether additional configuration is required.

Limitations​

Module behavior is constrained by source APIs, permissions, and metadata exposed by the platform. Refer to capability notes for unsupported or conditional features.

Troubleshooting​

If ingestion fails, validate credentials, permissions, connectivity, and scope filters first. Then review ingestion logs for source-specific errors and adjust configuration accordingly.

Code Coordinates​

  • Class Name: datahub.ingestion.source.dremio.dremio_source.DremioSource
  • Browse on GitHub
Questions?

If you've got any questions on configuring ingestion for Dremio, feel free to ping us on our Slack.

💡 Contributing to this documentation

This page is auto-generated from the underlying source code. To make changes, please edit the relevant source files in the metadata-ingestion directory.

Tip: For quick typo fixes or documentation updates, you can click the ✏️ Edit icon directly in the GitHub UI to open a Pull Request. For larger changes and PR naming conventions, please refer to our Contributing Guide.