hbase-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "churro morales (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (HBASE-15482) Provide an option to skip calculating block locations for SnapshotInputFormat
Date Fri, 18 Mar 2016 20:07:33 GMT

    [ https://issues.apache.org/jira/browse/HBASE-15482?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15202058#comment-15202058

churro morales commented on HBASE-15482:

I think Liyin is referring to only those InputFormats that deal specifically with store files,
if the Input format scans meta, that should still be fine.  We encountered the same issue
when dealing with snapshots because when you have a million store files and you calculate
the block distribution of each one that creates quite a bit of stress on the namenode.  

> Provide an option to skip calculating block locations for SnapshotInputFormat
> -----------------------------------------------------------------------------
>                 Key: HBASE-15482
>                 URL: https://issues.apache.org/jira/browse/HBASE-15482
>             Project: HBase
>          Issue Type: Improvement
>          Components: mapreduce
>            Reporter: Liyin Tang
>            Priority: Minor
> When a MR job is reading from SnapshotInputFormat, it needs to calculate the splits based
on the block locations in order to get best locality. However, this process may take a long
time for large snapshots. 
> In some setup, the computing layer, Spark, Hive or Presto could run out side of HBase
cluster. In these scenarios, the block locality doesn't matter. Therefore, it will be great
to have an option to skip calculating the block locations for every job. That will super useful
for the Hive/Presto/Spark connectors.

This message was sent by Atlassian JIRA

View raw message