hadoop-mapreduce-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Wilfred Spiegelenburg (JIRA)" <j...@apache.org>
Subject [jira] [Updated] (MAPREDUCE-6558) multibyte delimiters with compressed input files generate duplicate records
Date Tue, 10 May 2016 16:05:13 GMT

     [ https://issues.apache.org/jira/browse/MAPREDUCE-6558?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel

Wilfred Spiegelenburg updated MAPREDUCE-6558:
    Attachment: MAPREDUCE-6558.1.patch

A patch with test input that fails before the fix was made and passes after the fix was made.

I have run all the tests that were there and they all still pass with this change. The tests
have been run a large number of times with different input splits and none of them have failed.

The use cases that passed before the fix was applied have been documented in the code.

> multibyte delimiters with compressed input files generate duplicate records
> ---------------------------------------------------------------------------
>                 Key: MAPREDUCE-6558
>                 URL: https://issues.apache.org/jira/browse/MAPREDUCE-6558
>             Project: Hadoop Map/Reduce
>          Issue Type: Bug
>          Components: mrv1, mrv2
>    Affects Versions: 2.7.2
>            Reporter: Wilfred Spiegelenburg
>            Assignee: Wilfred Spiegelenburg
>         Attachments: MAPREDUCE-6558.1.patch
> This is the follow up for MAPREDUCE-6549. Compressed files cause record duplications
as shown in different junit tests. The number of duplicated records changes with the splitsize:
> Unexpected number of records in split (splitsize = 10)
> Expected: 41051
> Actual: 45062
> Unexpected number of records in split (splitsize = 100000)
> Expected: 41051
> Actual: 41052
> Test passes with splitsize = 147445 which is the compressed file length.The file is a
bzip2 file with 100k blocks and a total of 11 blocks

This message was sent by Atlassian JIRA

To unsubscribe, e-mail: mapreduce-issues-unsubscribe@hadoop.apache.org
For additional commands, e-mail: mapreduce-issues-help@hadoop.apache.org

View raw message