hadoop-pig-dev mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Daniel Dai (JIRA)" <j...@apache.org>
Subject [jira] Updated: (PIG-1144) set default_parallelism construct does not set the number of reducers correctly
Date Thu, 17 Dec 2009 18:52:18 GMT

     [ https://issues.apache.org/jira/browse/PIG-1144?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel

Daniel Dai updated PIG-1144:

    Attachment: PIG-1144-4.patch

Include logic for local mode. However, this is only for user use "-x local". If user do not
use -x option, and run into local mode because of no hadoop configuration file in CLASSPATH,
Pig do not have a way to detect that, and order by job will fail.

> set default_parallelism construct does not set the number of reducers correctly
> -------------------------------------------------------------------------------
>                 Key: PIG-1144
>                 URL: https://issues.apache.org/jira/browse/PIG-1144
>             Project: Pig
>          Issue Type: Bug
>          Components: impl
>    Affects Versions: 0.6.0
>         Environment: Hadoop 20 cluster with multi-node installation
>            Reporter: Viraj Bhat
>            Assignee: Daniel Dai
>             Fix For: 0.6.0
>         Attachments: brokenparallel.out, genericscript_broken_parallel.pig, PIG-1144-1.patch,
PIG-1144-2.patch, PIG-1144-3.patch, PIG-1144-4.patch
> Hi all,
>  I have a Pig script where I set the parallelism using the following set construct: "set
default_parallel 100" . I modified the "MRPrinter.java" to printout the parallelism
> {code}
> ...
> public void visitMROp(MapReduceOper mr)
> mStream.println("MapReduce node " + mr.getOperatorKey().toString() + " Parallelism "
+ mr.getRequestedParallelism());
> ...
> {code}
> When I run an explain on the script, I see that the last job which does the actual sort,
runs as a single reducer job. This can be corrected, by adding the PARALLEL keyword in front
of the ORDER BY.
> Attaching the script and the explain output
> Viraj

This message is automatically generated by JIRA.
You can reply to this email to add a comment to the issue online.

View raw message