hadoop-mapreduce-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Amareshwari Sriramadasu (JIRA)" <j...@apache.org>
Subject [jira] Commented: (MAPREDUCE-1597) combinefileinputformat does not work with non-splittable files
Date Tue, 07 Sep 2010 03:26:34 GMT

    [ https://issues.apache.org/jira/browse/MAPREDUCE-1597?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12906651#action_12906651

Amareshwari Sriramadasu commented on MAPREDUCE-1597:

bq. This patch would have to get merged with the recently committed one from MAPREDUCE-2046
Thanks Dhruba for the review. Uploaded patch is already merged with commit of MAPREDUCE-2046.
Will update the test results soon.

> combinefileinputformat does not work with non-splittable files
> --------------------------------------------------------------
>                 Key: MAPREDUCE-1597
>                 URL: https://issues.apache.org/jira/browse/MAPREDUCE-1597
>             Project: Hadoop Map/Reduce
>          Issue Type: Bug
>            Reporter: Namit Jain
>            Assignee: Amareshwari Sriramadasu
>             Fix For: 0.22.0
>         Attachments: patch-1597.txt
> CombineFileInputFormat.getSplits() does not take into account whether a file is splittable.
> This can lead to a problem for compressed text files - for example, getSplits() may return
> than 1 split depending on the size of the compressed file, all the splits recordreader
will read the
> complete file.
> I ran into this problem while using Hive on hadoop 20.

This message is automatically generated by JIRA.
You can reply to this email to add a comment to the issue online.

View raw message