impala-reviews mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Lars Volker (Code Review)" <>
Subject [Impala-ASF-CR] IMPALA-2523: Make HdfsTableSink aware of clustered input
Date Thu, 17 Nov 2016 15:25:47 GMT
Hello Alex Behm, Tim Armstrong,

I'd like you to reexamine a change.  Please visit

to look at the new patch set (#15).

Change subject: IMPALA-2523: Make HdfsTableSink aware of clustered input

IMPALA-2523: Make HdfsTableSink aware of clustered input

IMPALA-2521 introduced clustering for insert statements. This change
makes the HdfsTableSink aware of clustered inputs, so that partitions
are opened, written, and closed one by one.

This change also adds/modifies tests in several ways:

- clustered insert tests switch from selecting all rows from
  alltypessmall to alltypes. Together with varying settings for
  batch_size, this results in a larger number of row batches being
- clustered insert tests select from alltypes instead of
  functional.alltypes to make sure we also select from various input
- clustered insert tests have been added to select from alltypestiny to
  create inserts with 1 and 2 rows per partition respectively.
- exhaustive insert tests now use different values for batch_size: 1,
  16, 0 (meaning default, 1024). This is limited to uncompressed parquet
  files, to maintain a reasonable runtime. On my machine execution of
  test.insert took 1778 seconds, compared to 1002 seconds with the just
  default row batch size.
- There is additional testing in to make sure
  that insertion over several row batches only creates one file per
- It renames the test_insert method to make it unique in the file and
  allow for effective filtering with -k.

Change-Id: Ibeda0bdabbfe44c8ac95bf7c982a75649e1b82d0
M be/src/exec/
M be/src/exec/
M be/src/exec/hbase-table-writer.h
M be/src/exec/
M be/src/exec/hdfs-avro-table-writer.h
M be/src/exec/
M be/src/exec/hdfs-parquet-table-writer.h
M be/src/exec/
M be/src/exec/hdfs-sequence-table-writer.h
M be/src/exec/
M be/src/exec/hdfs-table-sink.h
M be/src/exec/
M be/src/exec/hdfs-table-writer.h
M be/src/exec/
M be/src/exec/hdfs-text-table-writer.h
M bin/
M common/thrift/DataSinks.thrift
M fe/src/main/java/org/apache/impala/analysis/
M fe/src/main/java/org/apache/impala/analysis/
M fe/src/main/java/org/apache/impala/analysis/
M fe/src/main/java/org/apache/impala/planner/
M fe/src/main/java/org/apache/impala/planner/
M fe/src/test/java/org/apache/impala/analysis/
M testdata/workloads/functional-query/queries/QueryTest/insert.test
M tests/query_test/
M tests/query_test/
26 files changed, 424 insertions(+), 138 deletions(-)

  git pull ssh:// refs/changes/63/4863/15
To view, visit
To unsubscribe, visit

Gerrit-MessageType: newpatchset
Gerrit-Change-Id: Ibeda0bdabbfe44c8ac95bf7c982a75649e1b82d0
Gerrit-PatchSet: 15
Gerrit-Project: Impala-ASF
Gerrit-Branch: master
Gerrit-Owner: Lars Volker <>
Gerrit-Reviewer: Alex Behm <>
Gerrit-Reviewer: Lars Volker <>
Gerrit-Reviewer: Marcel Kornacker <>
Gerrit-Reviewer: Tim Armstrong <>
Gerrit-Reviewer: Zoltan Ivanfi <>

View raw message