hive-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Janaki Lahorani (JIRA)" <>
Subject [jira] [Updated] (HIVE-16998) Add config to enable HoS DPP only for map-joins
Date Thu, 27 Jul 2017 17:42:00 GMT


Janaki Lahorani updated HIVE-16998:
    Status: In Progress  (was: Patch Available)

> Add config to enable HoS DPP only for map-joins
> -----------------------------------------------
>                 Key: HIVE-16998
>                 URL:
>             Project: Hive
>          Issue Type: Sub-task
>          Components: Logical Optimizer, Spark
>            Reporter: Sahil Takiar
>            Assignee: Janaki Lahorani
>         Attachments: HIVE16998.1.patch, HIVE16998.2.patch, HIVE16998.3.patch, HIVE16998.4.patch
> HoS DPP will split a given operator tree in two under the following conditions: it has
detected that the query can benefit from DPP, and the filter is not a map-join (see SplitOpTreeForDPP).
> This can hurt performance if the the non-partitioned side of the join involves a complex
operator tree - e.g. the query {{select count(*) from srcpart where srcpart.ds in (select
max(srcpart.ds) from srcpart union all select min(srcpart.ds) from srcpart)}} will require
running the subquery twice, once in each Spark job.
> Queries with map-joins don't get split into two operator trees and thus don't suffer
from this drawback. Thus, it would be nice to have a config key that just enables DPP on HoS
for map-joins.

This message was sent by Atlassian JIRA

View raw message