hadoop-mapreduce-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "dhruba borthakur (JIRA)" <j...@apache.org>
Subject [jira] Commented: (MAPREDUCE-1597) combinefileinputformat does not work with non-splittable files
Date Mon, 06 Sep 2010 20:12:34 GMT

    [ https://issues.apache.org/jira/browse/MAPREDUCE-1597?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12906579#action_12906579
] 

dhruba borthakur commented on MAPREDUCE-1597:
---------------------------------------------

+1, code looks good. This patch would have to get merged with the recently committed one from
MAPREDUCE-2046

> combinefileinputformat does not work with non-splittable files
> --------------------------------------------------------------
>
>                 Key: MAPREDUCE-1597
>                 URL: https://issues.apache.org/jira/browse/MAPREDUCE-1597
>             Project: Hadoop Map/Reduce
>          Issue Type: Bug
>            Reporter: Namit Jain
>            Assignee: Amareshwari Sriramadasu
>             Fix For: 0.22.0
>
>         Attachments: patch-1597.txt
>
>
> CombineFileInputFormat.getSplits() does not take into account whether a file is splittable.
> This can lead to a problem for compressed text files - for example, getSplits() may return
more
> than 1 split depending on the size of the compressed file, all the splits recordreader
will read the
> complete file.
> I ran into this problem while using Hive on hadoop 20.

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


Mime
View raw message