Mailing-List: contact dev-help@spark.apache.org; run by ezmlm
Precedence: bulk
MIME-Version: 1.0
Date: Wed, 14 Oct 2015 15:46:08 +0530
Message-ID: 
 <CAFiYKR9KP8kBYxt0FuJtSj_oz45bwk-EUA9Vp8p-5Ktq40r04A@mail.gmail.com>
Subject: Contributing Receiver based Low Level Kafka Consumer from
 Spark-Packages to Apache Spark Project
From: Dibyendu Bhattacharya <dibyendu.bhattachary@gmail.com>
To: "dev@spark.apache.org" <dev@spark.apache.org>
Content-Type: multipart/alternative; boundary=e89a8f83a90de1f45905220dd98c

--e89a8f83a90de1f45905220dd98c
Content-Type: text/plain; charset=UTF-8

Hi,

I have raised a JIRA ( https://issues.apache.org/jira/browse/SPARK-11045)
to track the discussion but also mailing dev group for your opinion. There
are some discussions already happened in Jira and love to hear what others
think. You can directly comment against the Jira if you wish.

This kafka consumer is around for a while in spark-packages (
http://spark-packages.org/package/dibbhatt/kafka-spark-consumer ) and I see
many people started using it , I am now thinking of contributing back to
Apache Spark core project so that it can get better support ,visibility and
adoption.

Few Point about this consumer

*Why this is needed , and how I position this Consumer : *

This Consumer is NOT the replacement for existing DirectStream API.
DirectStream solves the problem around "Exactly Once" semantics and "Global
Ordering" of messages . But to achieve this DirectStream comes with an
overhead. The overhead of maintaining the offset externally
, limited parallelism while processing the RDD ( as the RDD partition is
same as Kafka Partition ), and higher latency while processing RDD ( as
messages are fetched when RDD is processed) . There are many who does not
want "Exact Once" and "Global Ordering" of messages, or ordering are
managed in external store ( say HBase),  and want more parallelism and
lower latency in their Streaming channel . At this point Spark does not
have a better fallback option available in terms of Receiver Based API.
Present Receiver Based API use Kafka High Level API which is low
performance and has serious issue. [For this reason Kafka is coming up with
new High Level Consumer API in 0.9]

The Consumer which I implemented is using the Kafka Low Level API which
gives more performance.  This consumer has built in fault tolerant features
for all failures recovery. This Consumer extended the code from Storm Kafka
Spout which is being around for some time and has matured over the years
and has all built in Kafka fault tolerant capabilities. This same Kafka
consumer for spark is being running in various production scenarios
presently and already being adopted by many in the spark community.

*Why Can't we fix existing Receiver based API in Spark* :

This is not possible unless you move to Kafka Low Level API . Or let wait
for Kafka 0.9 where they are re-writing the HighLevel Consumer API and
built another consumer for Kafka 0.9 customers .
This approach seems to be not good in my opinion. The Kafka Low Level API
which I used in my consumer ( and even DirectStream uses ) will not going
to be deprecated in near future. So if Kafka Consumer for Spark is using
Low Level API for Receiver based mode, that will make sure all Kafka
Customers who are presently in 0.8.x or who will use 0.9 , benefited form
this same API. This will give easier maintenance to manage single API for
any Kafka versions. Also this will make sure both Direct Stream and
Receiver mode utilize same Kafka API.

*Concerns around Low Level API Complexity*

Yes, implementing a reliable consumer using Kafka Low Level consumer API is
complex. But same has been done for Strom -Kafka Spout and has been stable
for quite some time. This consumer for Spark is battle tested in various
production loads and gives much better performance than existing Kafka
Consumers for Spark and has better fault tolerant approach than existing
Receiver based mode. I do not think having a complex code should be a major
concern to deny a stable and high performance consumer for community. I am
okay if anyone interested to benchmark against other Kafka Consumers for
Spark and do various fault testing to make sure what I am saying is correct.

*Why can't this consumer continue to be in Spark-Package ?*

This can be possible. But what I see , many customer who want to fallback
to receiver based mode as they may not need "Exact Once" semantics or
"Global Ordering" , seems to little tentative using a spark-package library
for their critical streaming pipeline. And they are forced to use faulty
and buggy Kafka High Level API based mode. This consumer being part of
Spark project will give much higher adoption and support from community.

*Some Major features around this consumer :*

This consumer is controlling the rate limit by maintaining the constant
Block size where as default rate limiting in other Spark consumers are done
by number of messages. This is an issue when Kafka has messages of
different sizes and there is no deterministic way to know the actual block
sizes and memory utilization if rate control done by number of messages.

This consumer has in-built PID controller which controls the Rate of
consumption again by modifying the block size and consume only that much
amount of messages needed from Kafka . In default Spark consumer , it
fetches chunk of messages and then apply throttle to control the rate.
Which can lead to excess I/O while consuming from Kafka.

There are other features in this Consumer which we can discuss at length
once we are convinced that Kafka Low Level API is way to go.

Regards,
Dibyendu

--e89a8f83a90de1f45905220dd98c
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr">Hi,=C2=A0<div><br></div><div>I have raised a JIRA (=C2=A0<=
a href=3D"https://issues.apache.org/jira/browse/SPARK-11045">https://issues=
.apache.org/jira/browse/SPARK-11045</a>) to track the discussion but also m=
ailing dev group for your opinion. There are some discussions already happe=
ned in Jira and love to hear what others think. You can directly comment ag=
ainst the Jira if you wish.</div><div><br></div><div><span style=3D"color:r=
gb(51,51,51);font-family:&#39;Helvetica Neue&#39;,Helvetica,&#39;Segoe UI&#=
39;,Arial,freesans,sans-serif;font-size:14px;line-height:22.4px">This kafka=
 consumer is around for a while in spark-packages (</span><font color=3D"#3=
33333" face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans, sans-s=
erif"><span style=3D"font-size:14px;line-height:22.4px"><a href=3D"http://s=
park-packages.org/package/dibbhatt/kafka-spark-consumer">http://spark-packa=
ges.org/package/dibbhatt/kafka-spark-consumer</a> )=C2=A0</span></font><spa=
n style=3D"color:rgb(51,51,51);font-family:&#39;Helvetica Neue&#39;,Helveti=
ca,&#39;Segoe UI&#39;,Arial,freesans,sans-serif;font-size:14px;line-height:=
22.4px">and I see many people started using it , I am now thinking of contr=
ibuting back to Apache Spark core project so that it can get better support=
 ,visibility and adoption.</span></div><div><span style=3D"color:rgb(51,51,=
51);font-family:&#39;Helvetica Neue&#39;,Helvetica,&#39;Segoe UI&#39;,Arial=
,freesans,sans-serif;font-size:14px;line-height:22.4px"><br></span></div><d=
iv><span style=3D"color:rgb(51,51,51);font-family:&#39;Helvetica Neue&#39;,=
Helvetica,&#39;Segoe UI&#39;,Arial,freesans,sans-serif;font-size:14px;line-=
height:22.4px">Few Point about this consumer=C2=A0</span></div><div><span s=
tyle=3D"color:rgb(51,51,51);font-family:&#39;Helvetica Neue&#39;,Helvetica,=
&#39;Segoe UI&#39;,Arial,freesans,sans-serif;font-size:14px;line-height:22.=
4px"><br></span></div><div><span style=3D"color:rgb(51,51,51);font-family:&=
#39;Helvetica Neue&#39;,Helvetica,&#39;Segoe UI&#39;,Arial,freesans,sans-se=
rif;font-size:14px;line-height:22.4px"><u>Why this is needed , and how I po=
sition this Consumer :=C2=A0</u></span></div><div><span style=3D"color:rgb(=
51,51,51);font-family:&#39;Helvetica Neue&#39;,Helvetica,&#39;Segoe UI&#39;=
,Arial,freesans,sans-serif;font-size:14px;line-height:22.4px"><br></span></=
div><div><font color=3D"#333333" face=3D"Helvetica Neue, Helvetica, Segoe U=
I, Arial, freesans, sans-serif"><span style=3D"font-size:14px;line-height:2=
2.4px">This Consumer is NOT the replacement for existing DirectStream API. =
DirectStream solves the problem around &quot;Exactly Once&quot; semantics a=
nd &quot;Global Ordering&quot; of messages . But to=C2=A0achieve=C2=A0this =
DirectStream comes with an overhead. The overhead of=C2=A0maintaining=C2=A0=
the offset externally ,=C2=A0limited=C2=A0parallelism=C2=A0while processing=
 the RDD ( as the RDD partition is same as Kafka Partition ), and higher la=
tency while processing RDD ( as messages are fetched when RDD is processed)=
 . There are many who does not want &quot;Exact Once&quot; and &quot;Global=
 Ordering&quot; of messages, or ordering are managed in external store ( sa=
y HBase), =C2=A0and want more=C2=A0parallelism=C2=A0and lower latency in th=
eir Streaming channel . At this point Spark does not have a better fallback=
 option available in terms of Receiver Based API. Present Receiver Based AP=
I use Kafka High Level API which is low performance and has serious issue. =
[For this reason Kafka is coming up with new High Level Consumer API in 0.9=
]=C2=A0</span></font></div><div><font color=3D"#333333" face=3D"Helvetica N=
eue, Helvetica, Segoe UI, Arial, freesans, sans-serif"><span style=3D"font-=
size:14px;line-height:22.4px"><br></span></font></div><div><font color=3D"#=
333333" face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans, sans-=
serif"><span style=3D"font-size:14px;line-height:22.4px">The Consumer which=
 I implemented is using the Kafka Low Level API which gives more performanc=
e.=C2=A0 This consumer has built in fault tolerant features for all failure=
s recovery. This Consumer extended the code from Storm Kafka Spout which is=
 being around for some time and has matured over the years and has all buil=
t in Kafka fault tolerant capabilities. This same Kafka consumer for spark =
is being running in various production scenarios presently and already bein=
g adopted by many in the spark community.</span></font></div><div><font col=
or=3D"#333333" face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans=
, sans-serif"><span style=3D"font-size:14px;line-height:22.4px"><br></span>=
</font></div><div><font color=3D"#333333" face=3D"Helvetica Neue, Helvetica=
, Segoe UI, Arial, freesans, sans-serif"><span style=3D"font-size:14px;line=
-height:22.4px"><u>Why Can&#39;t we fix existing Receiver based API in Spar=
k</u> :</span></font></div><div><font color=3D"#333333" face=3D"Helvetica N=
eue, Helvetica, Segoe UI, Arial, freesans, sans-serif"><span style=3D"font-=
size:14px;line-height:22.4px"><br></span></font></div><div><font color=3D"#=
333333" face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans, sans-=
serif"><span style=3D"font-size:14px;line-height:22.4px">This is not possib=
le unless you move to Kafka Low Level API . Or let wait for Kafka 0.9 where=
 they are re-writing the HighLevel Consumer API and built another consumer =
for Kafka 0.9 customers .</span></font></div><div><font color=3D"#333333" f=
ace=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans, sans-serif"><s=
pan style=3D"font-size:14px;line-height:22.4px">This approach seems to be n=
ot good in my opinion. The Kafka Low Level API which I used in my consumer =
( and even DirectStream uses ) will not going to be deprecated in near futu=
re. So if Kafka Consumer for Spark is using Low Level API for Receiver base=
d mode, that will make sure all Kafka Customers who are presently in 0.8.x =
or who will use 0.9 , benefited=C2=A0form this same API. This will give eas=
ier=C2=A0maintenance=C2=A0to manage single API for any Kafka versions. Also=
 this will make sure both Direct Stream and Receiver mode utilize same Kafk=
a API.</span></font></div><div><font color=3D"#333333" face=3D"Helvetica Ne=
ue, Helvetica, Segoe UI, Arial, freesans, sans-serif"><span style=3D"font-s=
ize:14px;line-height:22.4px"><br></span></font></div><div><font color=3D"#3=
33333" face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans, sans-s=
erif"><span style=3D"font-size:14px;line-height:22.4px"><u>Concerns around =
Low Level API Complexity</u></span></font></div><div><font color=3D"#333333=
" face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans, sans-serif"=
><span style=3D"font-size:14px;line-height:22.4px"><br></span></font></div>=
<div><font color=3D"#333333" face=3D"Helvetica Neue, Helvetica, Segoe UI, A=
rial, freesans, sans-serif"><span style=3D"font-size:14px;line-height:22.4p=
x">Yes, implementing a reliable consumer using Kafka Low Level consumer API=
 is complex. But same has been done for Strom -Kafka Spout and has been sta=
ble for=C2=A0quite=C2=A0some time. This consumer for Spark is battle tested=
 in various production loads and gives much better performance than existin=
g Kafka Consumers for Spark and has better fault tolerant approach than exi=
sting Receiver based mode. I do not think having a complex code should be a=
 major concern to deny a stable and high performance consumer for community=
. I am okay if anyone interested to benchmark against other Kafka Consumers=
 for Spark and do various fault testing to make sure what I am saying is co=
rrect.</span></font></div><div><font color=3D"#333333" face=3D"Helvetica Ne=
ue, Helvetica, Segoe UI, Arial, freesans, sans-serif"><span style=3D"font-s=
ize:14px;line-height:22.4px"><br></span></font></div><div><font color=3D"#3=
33333" face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans, sans-s=
erif"><span style=3D"font-size:14px;line-height:22.4px"><u>Why can&#39;t th=
is consumer continue to be in Spark-Package ?</u></span></font></div><div><=
font color=3D"#333333" face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, =
freesans, sans-serif"><span style=3D"font-size:14px;line-height:22.4px"><br=
></span></font></div><div><font color=3D"#333333" face=3D"Helvetica Neue, H=
elvetica, Segoe UI, Arial, freesans, sans-serif"><span style=3D"font-size:1=
4px;line-height:22.4px">This can be possible. But what I see , many custome=
r who want to fallback to receiver based mode as they may not need &quot;Ex=
act Once&quot; semantics or &quot;Global Ordering&quot; , seems to little t=
entative using a spark-package library for their critical streaming pipelin=
e. And they are forced to use faulty and buggy Kafka High Level API based m=
ode. This consumer being part of Spark project will give much higher adopti=
on and support from community.</span></font></div><div><font color=3D"#3333=
33" face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans, sans-seri=
f"><span style=3D"font-size:14px;line-height:22.4px"><br></span></font></di=
v><div><font color=3D"#333333" face=3D"Helvetica Neue, Helvetica, Segoe UI,=
 Arial, freesans, sans-serif"><span style=3D"font-size:14px;line-height:22.=
4px"><u>Some Major features around this consumer :</u></span></font></div><=
div><br></div><div><div>This consumer is controlling the rate limit by main=
taining the constant Block size where as default rate limiting in other Spa=
rk consumers are done by number of messages. This is an issue when Kafka ha=
s messages of different sizes and there is no deterministic way to know the=
 actual block sizes and memory utilization if rate control done by number o=
f messages.</div><div><br></div><div>This consumer has in-built PID control=
ler which controls the Rate of consumption again by modifying the block siz=
e and consume only that much amount of messages needed from Kafka . In defa=
ult Spark consumer , it fetches chunk of messages and then apply throttle t=
o control the rate. Which can lead to excess I/O while consuming from Kafka=
. =C2=A0</div></div><div><font color=3D"#333333" face=3D"Helvetica Neue, He=
lvetica, Segoe UI, Arial, freesans, sans-serif"><span style=3D"font-size:14=
px;line-height:22.4px"><br></span></font></div><div><font color=3D"#333333"=
 face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans, sans-serif">=
<span style=3D"font-size:14px;line-height:22.4px">There are other features =
in this Consumer which we can discuss at length once we are convinced=C2=A0=
that Kafka Low Level API is way to go.</span></font></div><div><font color=
=3D"#333333" face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans, =
sans-serif"><span style=3D"font-size:14px;line-height:22.4px"><br></span></=
font></div><div><font color=3D"#333333" face=3D"Helvetica Neue, Helvetica, =
Segoe UI, Arial, freesans, sans-serif"><span style=3D"font-size:14px;line-h=
eight:22.4px">Regards,=C2=A0</span></font></div><div><font color=3D"#333333=
" face=3D"Helvetica Neue, Helvetica, Segoe UI, Arial, freesans, sans-serif"=
><span style=3D"font-size:14px;line-height:22.4px">Dibyendu</span></font></=
div><div><font color=3D"#333333" face=3D"Helvetica Neue, Helvetica, Segoe U=
I, Arial, freesans, sans-serif"><span style=3D"font-size:14px;line-height:2=
2.4px"><br></span></font></div><div><span style=3D"color:rgb(51,51,51);font=
-family:&#39;Helvetica Neue&#39;,Helvetica,&#39;Segoe UI&#39;,Arial,freesan=
s,sans-serif;font-size:14px;line-height:22.4px"><br></span></div><div><span=
 style=3D"color:rgb(51,51,51);font-family:&#39;Helvetica Neue&#39;,Helvetic=
a,&#39;Segoe UI&#39;,Arial,freesans,sans-serif;font-size:14px;line-height:2=
2.4px"><br></span></div><div><span style=3D"color:rgb(51,51,51);font-family=
:&#39;Helvetica Neue&#39;,Helvetica,&#39;Segoe UI&#39;,Arial,freesans,sans-=
serif;font-size:14px;line-height:22.4px"><br></span></div><div><span style=
=3D"color:rgb(51,51,51);font-family:&#39;Helvetica Neue&#39;,Helvetica,&#39=
;Segoe UI&#39;,Arial,freesans,sans-serif;font-size:14px;line-height:22.4px"=
><br></span></div></div>

--e89a8f83a90de1f45905220dd98c--