HPE Ezmeral Data Fabric 6.2 is In Maintenance and transitions to "End of Maintenance" in June 2024. Please see the latest documentation.

About Release 6.2
This site contains documentation for HPE Ezmeral Data Fabric release 6.2 including installation, configuration, administration, and reference content, as well as content for the associated bundled ecosystem components and drivers.
6.2 Installation
This section contains information about installing and upgrading HPE Ezmeral Data Fabric software. It also contains information about how to migrate data and applications from an Apache Hadoop cluster to a HPE Ezmeral Data Fabric cluster.
6.2 Data Fabric
HPE Ezmeral Data Fabric is the industry-leading data platform for AI and analytics that solves enterprise business needs.
6.2 Administration
This section describes how to manage the nodes and services that make up a cluster.
6.2 Development
This section contains information related to application development for Ezmeral ecosystem components and HPE Ezmeral Data Fabric products, including the file system, Database (Key-Value and JSON), and Event Streams.
- Application Development Process
  Before you start developing applications on the HPE Ezmeral Data Fabric platform, consider how you will get the data into the platform, the storage format of the data, the type of processing or modeling that is required, and how the data will be accessed.
- File Store and Apps
  The following sections provide information about accessing the File Store with C and Java applications.
- HPE Ezmeral Data Fabric Database and Apps
  This section contains information about developing client applications for JSON and key-value tables.
- HPE Ezmeral Data Fabric Streams and Apps
  HPE Ezmeral Data Fabric Streams brings integrated publish and subscribe messaging to HPE Ezmeral Data Fabric.
- MapReduce and Apps
  This section contains information associated with developing YARN applications.
- Kubernetes Interfaces for Data Fabric
  This section describes how to leverage the capabilities of the Kubernetes Interfaces for Data Fabric.
- Ecosystem Components
  The following sections provide information about each open-source project that is supported by the HPE Ezmeral Data Fabric.
  - Ezmeral Ecosystem Packs
  - Apache Airflow
    This topic provides an overview of Apache Airflow on HPE Ezmeral Data Fabric.
  - AsyncHBase
  - Cascading
  - Apache Drill
    - Drill Tutorial
    - Drill-on-YARN
    - Configuring Drill
      Lists the data-fabric-specific configuration for Drill.
    - Working with Drill
      - Connecting Drill to Data Sources
        Choose and configure storage plugins to enable Drill to connect to a data source.
        Drill Storage and Format Plugin Support Matrix
        maprdb Format Plugin for Drill
        Drill supports access to HPE Ezmeral Data Fabric Database JSON and binary tables through the maprdb format plugin.
        Configuring the Hive Storage Plugin
        Configuring the Kafka Storage Plugin
        The Kafka storage plugin is not officially supported for Drill; however, if you choose to configure Kafka as a data source in Drill, you must update the <drill_home>/jars/3rdParty directory such that it contains the required JAR files and then restart Drill before you configure the kafka storage plugin in the Drill Web UI.
      - Start the Drill Web UI
        The Drill Web UI is one of several client interfaces that you can use to access Drill.
      - Start the Drill Shell (SQLLine)
        SQLLine is a JDBC application packaged with Drill that serves as the Drill shell. When you issue queries from the SQLLine, the SQLLine client sends the queries to the connected Drillbit (Drill node).
      - Hive to Drill Type Mapping
    - Securing Drill
      An administrator can install Drill with the default security configuration or manually configure custom security for Drill.
    - Drill Drivers
      HPE Ezmeral Data Fabric provides Drill ODBC and JDBC drivers that you can download and use to connect Drill to BI tools. The drivers are updated periodically to include support for new functionality in Drill.
    - Drill Configuration Files
      The Drill installation includes configuration files with start-up options that you can modify prior to starting Drill.
    - Monitoring Drill Metrics
    - Optimizing Queries with Indexes
      HPE Ezmeral Data Fabric Database provides a highly scalable key-value database platform on which you can run SQL queries using Drill. As of the 6.0 release of the MapR Data Platform, HPE Ezmeral Data Fabric Database natively supports indexes on secondary fields in JSON tables.
    - Drill Limitations
      Provides information about Drill limitations and solutions where applicable.
    - Vulnerability Reports
      Provides vulnerability information in relation to Drill.
  - Flume
  - Hadoop
  - HBase
  - HBase Client and HPE Ezmeral Data Fabric Database Binary Tables
  - HCatalog
  - Hive
  - HttpFS
  - Hue
  - Impala
  - Livy
    Apache Livy is primarily used to provide integration between Hue and Spark.
  - HPE Ezmeral Data Fabric Streams Clients and Tools
    Describes the supported HPE Ezmeral Data Fabric Streams tools and clients.
  - S3 Gateway
    The S3 gateway is a service that provides an S3-compatible interface to expose data in HPE Ezmeral Data Fabric as objects. The S3 gateway manages all inbound S3 API requests to put data into and get data out of cloud storage.
  - Oozie
  - Pig
  - Sentry
  - Apache Spark
  - Sqoop
  - YARN
- Maven and the HPE Ezmeral Data Fabric
  This section discusses topics associated with Maven and the HPE Ezmeral Data Fabric.
- Developer's Reference
  This section contains in-depth information for the developer.
- API Documentation
  HPE Ezmeral Data Fabric supports public APIs for file system, HPE Ezmeral Data Fabric Database, and HPE Ezmeral Data Fabric Streams. These APIs are available for application-development purposes.
Other Docs
This section contains release-independent information, including: Installer documentation, Ecosystem release notes, interoperability matrices, security vulnerabilities, and links to other data-fabric version documentation.
Glossary
Definitions for commonly used terms in MapR Converged Data Platform environments.

Drill Storage and Format Plugin Support Matrix

You can deploy Drill without Hadoop in a standalone configuration on a single node, however multi-node standalone cluster deployments of Drill are not supported. Note that Drill itself does not require Hadoop.

The following table lists the supported and unsupported data sources and formats in Drill:

Data Source	Storage Plugin Type	Formats	Supported
file system	dfs	Text (CSV, TSV, PSV)	Yes
		Parquet	Yes
		JSON	Yes
		Avro	No
HPE Ezmeral Data Fabric Database	dfs	Binary	Yes
		JSON	Yes
HBase	hbase	Binary	No (as of Drill 1.11 and Core 6.0)
Hive	hive	Text (CSV, TSV, PSV)	Yes
		Parquet	Yes
		JSON	Yes
		Avro	Yes
		Other Hive built-in SerDes	Yes (Not recommended due to the memory overhead and performance implications.)
S3	s3	Supports the same formats as the dfs storage plugin.	Yes
MongoDB	mongodb	N/A	No
RDBMS	jdbc	N/A	No
Kudu	kudu	N/A	No
Kafka	kafka	JSON	No NOTE The kafka storage plugin on the Streams is in the Alpha testing phase and not officially supported. See Configuring the Kafka Storage Plugin for more information.
OpenTSDB	openTSDB	N/A	NOTE The openTSDB storage plugin is not officially supported. See OpenTSDB Storage Plugin for more information.

NOTE As of the Core 6.0 and Drill 1.11, HBase is no longer supported, therefore the communication path between Drill and HBase is also not supported. If you have an hbase storage plugin configured in Drill, you should disable it.