hadoop-common-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Robert Chansler (JIRA)" <j...@apache.org>
Subject [jira] Updated: (HADOOP-3019) want input sampler & sorted partitioner
Date Tue, 21 Oct 2008 19:56:44 GMT

     [ https://issues.apache.org/jira/browse/HADOOP-3019?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel
]

Robert Chansler updated HADOOP-3019:
------------------------------------

    Release Note: Added a partitioner that effects a total order of output data, and an input
sampler for generating the partition keyset for TotalOrderPartitioner for when the map's input
keytype and distribution approximates its output.  (was: Adds a partitioner capable of effecting
a total order of output data. Also includes an input sampler for generating the partition
keyset for TotalOrderPartitioner, useful where the map's input keytype and distribution approximates
its output.)

> want input sampler & sorted partitioner
> ---------------------------------------
>
>                 Key: HADOOP-3019
>                 URL: https://issues.apache.org/jira/browse/HADOOP-3019
>             Project: Hadoop Core
>          Issue Type: New Feature
>          Components: mapred
>            Reporter: Doug Cutting
>            Assignee: Chris Douglas
>             Fix For: 0.19.0
>
>         Attachments: 3019-0.patch, 3019-1.patch, 3019-2.patch, 3019-3.patch, 3019-4.patch,
3019-5.patch
>
>
> The input sampler should generate a small, random sample of the input, saved to a file.
> The partitioner should read the sample file and partition keys into relatively even-sized
key-ranges, where the partition numbers correspond to key order.
> Note that when the sampler is used for partitioning, the number of samples required is
proportional to the number of reduce partitions.  10x the intended reducer count should give
good results.

-- 
This message is automatically generated by JIRA.
-
You can reply to this email to add a comment to the issue online.


Mime
View raw message