hadoop-mapreduce-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Hudson (JIRA)" <j...@apache.org>
Subject [jira] Commented: (MAPREDUCE-1248) Redundant memory copying in StreamKeyValUtil
Date Fri, 29 Oct 2010 02:05:34 GMT

    [ https://issues.apache.org/jira/browse/MAPREDUCE-1248?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=12926043#action_12926043

Hudson commented on MAPREDUCE-1248:

Integrated in Hadoop-Mapreduce-trunk-Commit #523 (See [https://hudson.apache.org/hudson/job/Hadoop-Mapreduce-trunk-Commit/523/])

> Redundant memory copying in StreamKeyValUtil
> --------------------------------------------
>                 Key: MAPREDUCE-1248
>                 URL: https://issues.apache.org/jira/browse/MAPREDUCE-1248
>             Project: Hadoop Map/Reduce
>          Issue Type: Improvement
>          Components: contrib/streaming
>            Reporter: Ruibang He
>            Priority: Minor
>             Fix For: 0.22.0
>         Attachments: MAPREDUCE-1248-v1.0.patch
> I found that when MROutputThread collecting the output of  Reducer, it calls StreamKeyValUtil.splitKeyVal()
and two local byte-arrays are allocated there for each line of output. Later these two byte-arrays
are passed to variable key and val. There are twice memory copying here, one is the System.arraycopy()
method, the other is inside key.set() / val.set().
> This causes double times of memory copying for the whole output (may lead to higher CPU
consumption), and frequent temporay object allocation.

This message is automatically generated by JIRA.
You can reply to this email to add a comment to the issue online.

View raw message