Mailing-List: contact user-help@hive.apache.org; run by ezmlm
Precedence: bulk
Reply-To: user@hive.apache.org
Received-SPF: pass (nike.apache.org: domain of dimad@microsoft.com designates
 207.46.100.23 as permitted sender)
From: Dima Datsenko <dimad@microsoft.com>
To: Bennie Schut <bschut@ebuddy.com>, "user@hive.apache.org"
	<user@hive.apache.org>
Subject: RE: Effecient partitions usage in join
Thread-Topic: Effecient partitions usage in join
Thread-Index: Ac3Itp7vKQ0i7EFESLirkoP8vQengQAAoWLgAAC37MAAAa2F0A==
Date: Thu, 22 Nov 2012 15:06:45 +0000
Message-ID: 
 <2BBBD93B295A4442A5B731B34A80C5622AA72A8E@DB3EX14MBXC316.europe.corp.microsoft.com>
References: 
 <2BBBD93B295A4442A5B731B34A80C5622AA72A09@DB3EX14MBXC316.europe.corp.microsoft.com>
 <FF1DF58D04F11D4291D09795D1A4EF16185935593C@SRV-MAIL>
In-Reply-To: <FF1DF58D04F11D4291D09795D1A4EF16185935593C@SRV-MAIL>
Accept-Language: en-US
Content-Language: en-US
Content-Type: multipart/alternative;
	boundary="_000_2BBBD93B295A4442A5B731B34A80C5622AA72A8EDB3EX14MBXC316e_"
MIME-Version: 1.0

--_000_2BBBD93B295A4442A5B731B34A80C5622AA72A8EDB3EX14MBXC316e_
Content-Type: text/plain; charset="windows-1255"
Content-Transfer-Encoding: quoted-printable

Hi Benny,

The udf solution sounds like a plan. Much better than generating hive query=
 with hardcoded partition out of table B. Can you please provide a sample o=
f what you=92re doing there?

Thanks,
Dima

From: Bennie Schut [mailto:bschut@ebuddy.com]
Sent: =E9=E5=ED =E4 22 =F0=E5=E1=EE=E1=F8 2012 16:28
To: user@hive.apache.org
Cc: Dima Datsenko
Subject: RE: Effecient partitions usage in join

Unfortunately at the moment partition pruning is a bit limited in hive. Whe=
n hive creates the query plan it decides what partitions to use. So if you =
put hardcoded list of partition_id items in the where clause it will know w=
hat to do. In the case of a join (or a subquery) it would have to run the q=
uery before it can know what it can prune.  There are obvious solutions to =
this but they are simply not implemented at the moment.
Generally speaking people try to work around this by not normalizing the da=
ta. So if you plan on doing a clean star schema with a calendar table then =
do yourself a favor and but the actual date in the fact table and not a mea=
ningless key.
It=92s also good to realize you can (in some special cases) work around it =
by using udf=92s. I=92ve used it once by creating a udf which produced the =
current date which I flagged as deterministic (ugly I know). This causes th=
e planner to run the udf during planning and use the result as if it=92s a =
constant and thus partition pruning works again. It=92s currently the only =
way I know to select x days of data with partition pruning working.


From: Dima Datsenko [mailto:dimad@microsoft.com]
Sent: Thursday, November 22, 2012 2:56 PM
To: user@hive.apache.org<mailto:user@hive.apache.org>
Subject: Effecient partitions usage in join

Hi Guys,

I wonder if you could help me.

I have a huge Hive table partitioned by some field. It has thousands of par=
titions.
Now I have another small table containing tens of partitions id. I=92d like=
 to get the data only from those partitions.

However when I run
Select * from A join B on (A.partition_id =3D B.partition_id),
It reads all data from A, then from B and on reduce stage performs join.

I tried /*+ MAPJOIN*/ it ran faster sparing reduce operation, but still rea=
d the whole A table.

Is there a more efficient way to perform the query w/o reading the whole A =
content?


Thanks
Dima

--_000_2BBBD93B295A4442A5B731B34A80C5622AA72A8EDB3EX14MBXC316e_
Content-Type: text/html; charset="windows-1255"
Content-Transfer-Encoding: quoted-printable

<html>
<head>
<meta http-equiv=3D"Content-Type" content=3D"text/html; charset=3Dwindows-1=
255">
<meta name=3D"Generator" content=3D"Microsoft Word 14 (filtered medium)">
<style><!--
/* Font Definitions */
@font-face
	{font-family:Helvetica;
	panose-1:2 11 6 4 2 2 2 2 2 4;}
@font-face
	{font-family:Helvetica;
	panose-1:2 11 6 4 2 2 2 2 2 4;}
@font-face
	{font-family:Calibri;
	panose-1:2 15 5 2 2 2 4 3 2 4;}
@font-face
	{font-family:Tahoma;
	panose-1:2 11 6 4 3 5 4 4 2 4;}
/* Style Definitions */
p.MsoNormal, li.MsoNormal, div.MsoNormal
	{margin:0cm;
	margin-bottom:.0001pt;
	font-size:11.0pt;
	font-family:"Calibri","sans-serif";}
a:link, span.MsoHyperlink
	{mso-style-priority:99;
	color:blue;
	text-decoration:underline;}
a:visited, span.MsoHyperlinkFollowed
	{mso-style-priority:99;
	color:purple;
	text-decoration:underline;}
p.MsoAcetate, li.MsoAcetate, div.MsoAcetate
	{mso-style-priority:99;
	mso-style-link:"Balloon Text Char";
	margin:0cm;
	margin-bottom:.0001pt;
	font-size:8.0pt;
	font-family:"Tahoma","sans-serif";}
span.EmailStyle17
	{mso-style-type:personal;
	font-family:"Calibri","sans-serif";
	color:windowtext;}
span.EmailStyle18
	{mso-style-type:personal;
	font-family:"Calibri","sans-serif";
	color:#1F497D;}
span.EmailStyle19
	{mso-style-type:personal;
	font-family:"Calibri","sans-serif";
	color:#1F497D;}
span.EmailStyle20
	{mso-style-type:personal-reply;
	font-family:"Calibri","sans-serif";
	color:#1F497D;}
span.BalloonTextChar
	{mso-style-name:"Balloon Text Char";
	mso-style-priority:99;
	mso-style-link:"Balloon Text";
	font-family:"Tahoma","sans-serif";}
.MsoChpDefault
	{mso-style-type:export-only;
	font-size:10.0pt;}
@page WordSection1
	{size:612.0pt 792.0pt;
	margin:72.0pt 72.0pt 72.0pt 72.0pt;}
div.WordSection1
	{page:WordSection1;}
--></style><!--[if gte mso 9]><xml>
<o:shapedefaults v:ext=3D"edit" spidmax=3D"1026" />
</xml><![endif]--><!--[if gte mso 9]><xml>
<o:shapelayout v:ext=3D"edit">
<o:idmap v:ext=3D"edit" data=3D"1" />
</o:shapelayout></xml><![endif]-->
</head>
<body lang=3D"EN-US" link=3D"blue" vlink=3D"purple">
<div class=3D"WordSection1">
<p class=3D"MsoNormal"><span style=3D"color:#1F497D">Hi Benny,<o:p></o:p></=
span></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D">The udf solution sound=
s like a plan. Much better than generating hive query with hardcoded partit=
ion out of table B. Can you please provide a sample of what you=92re doing =
there?<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D">Thanks,<o:p></o:p></sp=
an></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D">Dima<o:p></o:p></span>=
</p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<div>
<div style=3D"border:none;border-top:solid #B5C4DF 1.0pt;padding:3.0pt 0cm =
0cm 0cm">
<p class=3D"MsoNormal"><b><span style=3D"font-size:10.0pt;font-family:&quot=
;Tahoma&quot;,&quot;sans-serif&quot;">From:</span></b><span style=3D"font-s=
ize:10.0pt;font-family:&quot;Tahoma&quot;,&quot;sans-serif&quot;"> Bennie S=
chut [mailto:bschut@ebuddy.com]
<br>
<b>Sent:</b> <span lang=3D"HE" dir=3D"RTL">=E9=E5=ED</span><span dir=3D"LTR=
"></span><span dir=3D"LTR"></span>&nbsp;<span lang=3D"HE" dir=3D"RTL">=E4</=
span><span dir=3D"LTR"></span><span dir=3D"LTR"></span> 22
<span lang=3D"HE" dir=3D"RTL">=F0=E5=E1=EE=E1=F8</span><span dir=3D"LTR"></=
span><span dir=3D"LTR"></span> 2012 16:28<br>
<b>To:</b> user@hive.apache.org<br>
<b>Cc:</b> Dima Datsenko<br>
<b>Subject:</b> RE: Effecient partitions usage in join<o:p></o:p></span></p=
>
</div>
</div>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D">Unfortunately at the m=
oment partition pruning is a bit limited in hive. When hive creates the que=
ry plan it decides what partitions to use. So if you put hardcoded list of =
partition_id items in the where clause
 it will know what to do. In the case of a join (or a subquery) it would ha=
ve to run the query before it can know what it can prune. &nbsp;There are o=
bvious solutions to this but they are simply not implemented at the moment.=
<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D">Generally speaking peo=
ple try to work around this by not normalizing the data. So if you plan on =
doing a clean star schema with a calendar table then do yourself a favor an=
d but the actual date in the fact table
 and not a meaningless key.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D">It=92s also good to re=
alize you can (in some special cases) work around it by using udf=92s. I=92=
ve used it once by creating a udf which produced the current date which I f=
lagged as deterministic (ugly I know). This
 causes the planner to run the udf during planning and use the result as if=
 it=92s a constant and thus partition pruning works again. It=92s currently=
 the only way I know to select x days of data with partition pruning workin=
g.<o:p></o:p></span></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<p class=3D"MsoNormal"><span style=3D"color:#1F497D"><o:p>&nbsp;</o:p></spa=
n></p>
<div>
<div style=3D"border:none;border-top:solid #B5C4DF 1.0pt;padding:3.0pt 0cm =
0cm 0cm">
<p class=3D"MsoNormal"><b><span style=3D"font-size:10.0pt;font-family:&quot=
;Tahoma&quot;,&quot;sans-serif&quot;">From:</span></b><span style=3D"font-s=
ize:10.0pt;font-family:&quot;Tahoma&quot;,&quot;sans-serif&quot;"> Dima Dat=
senko [<a href=3D"mailto:dimad@microsoft.com">mailto:dimad@microsoft.com</a=
>]
<br>
<b>Sent:</b> Thursday, November 22, 2012 2:56 PM<br>
<b>To:</b> <a href=3D"mailto:user@hive.apache.org">user@hive.apache.org</a>=
<br>
<b>Subject:</b> Effecient partitions usage in join<o:p></o:p></span></p>
</div>
</div>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Hi Guys,<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">I wonder if you could help me.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">I have a huge Hive table partitioned by some field. =
It has thousands of partitions.<o:p></o:p></p>
<p class=3D"MsoNormal">Now I have another small table containing tens of pa=
rtitions id. I=92d like to get the data only from those partitions.<o:p></o=
:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">However when I run<o:p></o:p></p>
<p class=3D"MsoNormal">Select * from A join B on (A.partition_id =3D B.part=
ition_id),
<o:p></o:p></p>
<p class=3D"MsoNormal">It reads all data from A, then from B and on reduce =
stage performs join.<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">I tried <span style=3D"font-size:10.0pt;font-family:=
&quot;Helvetica&quot;,&quot;sans-serif&quot;">
/*&#43; MAPJOIN*/ it ran faster sparing reduce operation, but still read th=
e whole A table.</span><o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Is there a more efficient way to perform the query w=
/o reading the whole A content?<o:p></o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal"><o:p>&nbsp;</o:p></p>
<p class=3D"MsoNormal">Thanks<o:p></o:p></p>
<p class=3D"MsoNormal">Dima<o:p></o:p></p>
</div>
</body>
</html>

--_000_2BBBD93B295A4442A5B731B34A80C5622AA72A8EDB3EX14MBXC316e_--