hive-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Sahil Takiar (JIRA)" <j...@apache.org>
Subject [jira] [Created] (HIVE-16998) Add config to enable HoS DPP only for map-joins
Date Fri, 30 Jun 2017 01:05:00 GMT
Sahil Takiar created HIVE-16998:
-----------------------------------

             Summary: Add config to enable HoS DPP only for map-joins
                 Key: HIVE-16998
                 URL: https://issues.apache.org/jira/browse/HIVE-16998
             Project: Hive
          Issue Type: Sub-task
          Components: Logical Optimizer, Spark
            Reporter: Sahil Takiar
            Assignee: Sahil Takiar


HoS DPP will split a given operator tree in two under the following conditions: it has detected
that the query can benefit from DPP, and the filter is not a map-join (see SplitOpTreeForDPP).

This can hurt performance if the the non-partitioned side of the join involves a complex operator
tree - e.g. the query {{select count(*) from srcpart where srcpart.ds in (select max(srcpart.ds)
from srcpart union all select min(srcpart.ds) from srcpart)}} will require running the subquery
twice, once in each Spark job.

Queries with map-joins don't get split into two operator trees and thus don't suffer from
this drawback. Thus, it would be nice to have a config key that just enables DPP on HoS for
map-joins.



--
This message was sent by Atlassian JIRA
(v6.4.14#64029)

Mime
View raw message