hive-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Brock Noland (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (HIVE-860) Persistent distributed cache
Date Mon, 17 Feb 2014 04:44:19 GMT

    [ https://issues.apache.org/jira/browse/HIVE-860?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=13902950#comment-13902950
] 

Brock Noland commented on HIVE-860:
-----------------------------------

I think we can take a similar approach to https://issues.apache.org/jira/browse/PIG-2672.

> Persistent distributed cache
> ----------------------------
>
>                 Key: HIVE-860
>                 URL: https://issues.apache.org/jira/browse/HIVE-860
>             Project: Hive
>          Issue Type: Improvement
>            Reporter: Zheng Shao
>
> DistributedCache is shared across multiple jobs, if the hdfs file name is the same.
> We need to make sure Hive put the same file into the same location every time and do
not overwrite if the file content is the same.
> We can achieve 2 different results:
> A1. Files added with the same name, timestamp, and md5 in the same session will have
a single copy in distributed cache.
> A2. Filed added with the same name, timestamp, and md5 will have a single copy in distributed
cache.
> A2 has a bigger benefit in sharing but may raise a question on when Hive should clean
it up in hdfs.



--
This message was sent by Atlassian JIRA
(v6.1.5#6160)

Mime
View raw message