hadoop-mapreduce-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Vinod Kumar Vavilapalli (JIRA)" <j...@apache.org>
Subject [jira] [Updated] (MAPREDUCE-4892) CombineFileInputFormat node input split can be skewed on small clusters
Date Wed, 27 Feb 2013 17:19:13 GMT

     [ https://issues.apache.org/jira/browse/MAPREDUCE-4892?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]

Vinod Kumar Vavilapalli updated MAPREDUCE-4892:
-----------------------------------------------

    Attachment: MAPREDUCE-4892.1.alt.patch

Same patch reattached.
                
> CombineFileInputFormat node input split can be skewed on small clusters
> -----------------------------------------------------------------------
>
>                 Key: MAPREDUCE-4892
>                 URL: https://issues.apache.org/jira/browse/MAPREDUCE-4892
>             Project: Hadoop Map/Reduce
>          Issue Type: Bug
>            Reporter: Bikas Saha
>            Assignee: Bikas Saha
>             Fix For: 3.0.0
>
>         Attachments: MAPREDUCE-4892.1.alt.patch, MAPREDUCE-4892.1.alt.patch, MAPREDUCE-4892.1.patch
>
>
> The CombineFileInputFormat split generation logic tries to group blocks by node in order
to create splits. It iterates through the nodes and creates splits on them until there aren't
enough blocks left on a node that can be grouped into a valid split. If the first few nodes
have a lot of blocks on them then they can end up getting a disproportionately large share
of the total number of splits created. This can result in poor locality of maps. This problem
is likely to happen on small clusters where its easier to create a skew in the distribution
of blocks on nodes.

--
This message is automatically generated by JIRA.
If you think it was sent incorrectly, please contact your JIRA administrators
For more information on JIRA, see: http://www.atlassian.com/software/jira

Mime
View raw message