hive-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Sahil Takiar (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (HIVE-20032) Don't serialize hashCode when groupByShuffle and RDD cacheing is disabled
Date Fri, 20 Jul 2018 16:36:00 GMT

    [ https://issues.apache.org/jira/browse/HIVE-20032?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16550959#comment-16550959
] 

Sahil Takiar commented on HIVE-20032:
-------------------------------------

Attached an updated patch that should fix the unit tests. Since the new approach is using
reflection to create a class in the {{kryo-registrator}} module, we have to make sure that
{{hive-kryo-registrator}} jar file is on the current classpath. Apparently the {{--jar}} option
in {{spark-submit}} does not guarantee that for the Spark Driver. So I'm using {{spark.driver.extraClassPath}}
to make sure its present when running the {{SparkPlanGenerator}}.

When {{spark.master}} is {{local}} things are a bit tricker because there is no separate driver
process. So instead I use {{SparkClientUtilities#addJarToContextLoader}} to add the jar to
the current classpath. The constructor for {{LocalHiveSparkClient}} is already doing this,
but for whatever reason the classloader is getting reset by the time {{SparkPlanGenerator}}
runs. I tried debugging it, but didn't have any luck, and given that its just a test infra
issue it doesn't seem work resolving unless anyone has suggestions on how to fix it.

> Don't serialize hashCode when groupByShuffle and RDD cacheing is disabled
> -------------------------------------------------------------------------
>
>                 Key: HIVE-20032
>                 URL: https://issues.apache.org/jira/browse/HIVE-20032
>             Project: Hive
>          Issue Type: Improvement
>          Components: Spark
>            Reporter: Sahil Takiar
>            Assignee: Sahil Takiar
>            Priority: Major
>         Attachments: HIVE-20032.1.patch, HIVE-20032.2.patch, HIVE-20032.3.patch, HIVE-20032.4.patch,
HIVE-20032.5.patch, HIVE-20032.6.patch
>
>
> Follow up on HIVE-15104, if we don't enable RDD cacheing or groupByShuffles, then we
don't need to serialize the hashCode when shuffling data in HoS.



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

Mime
View raw message