hadoop-yarn-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Pradeep Ambati (JIRA)" <j...@apache.org>
Subject [jira] [Comment Edited] (YARN-8242) YARN NM: OOM error while reading back the state store on recovery
Date Thu, 16 Aug 2018 17:44:00 GMT

    [ https://issues.apache.org/jira/browse/YARN-8242?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16582862#comment-16582862
] 

Pradeep Ambati edited comment on YARN-8242 at 8/16/18 5:43 PM:
---------------------------------------------------------------

In the latest patch, I have addressed all but one of the comments. There seems to be check
style issue that I still have to address. I will address that next patch that I will upload
shortly.

 
{quote}It's a bit trickier to do the full iteration for localized resource state, but it should
be possible. I would be fine with punting that to a followup JIRA since this current work
is still a significant improvement over the old method of loading everything at once.
{quote}
 

I will create a followup Jira for LocalTrackerState iterator. 

 


was (Author: pradeepambati):
In the latest patch, I have addressed all but one of the comments. There seems to be check
style issue that I still have to address. I will address that next patch that I will upload
shortly.

 

I  It's a bit trickier to do the full iteration for localized resource state, but it should
be possible. I would be fine with punting that to a followup JIRA since this current work
is still a significant improvement over the old method of loading everything at once.

 

I will create a followup Jira for LocalTrackerState iterator. 

 

> YARN NM: OOM error while reading back the state store on recovery
> -----------------------------------------------------------------
>
>                 Key: YARN-8242
>                 URL: https://issues.apache.org/jira/browse/YARN-8242
>             Project: Hadoop YARN
>          Issue Type: Improvement
>          Components: yarn
>    Affects Versions: 2.6.0, 2.9.0, 2.6.5, 2.8.3, 3.1.0, 2.7.6, 3.0.2
>            Reporter: Kanwaljeet Sachdev
>            Assignee: Pradeep Ambati
>            Priority: Critical
>         Attachments: YARN-8242.001.patch, YARN-8242.002.patch, YARN-8242.003.patch, YARN-8242.004.patch,
YARN-8242.005.patch, YARN-8242.006.patch
>
>
> On startup the NM reads its state store and builds a list of application in the state
store to process. If the number of applications in the state store is large and have a lot
of "state" connected to it the NM can run OOM and never get to the point that it can start
processing the recovery.
> Since it never starts the recovery there is no way for the NM to ever pass this point.
It will require a change in heap size to get the NM started.
>  
> Following is the stack trace
> {code:java}
> at java.lang.OutOfMemoryError.<init> (OutOfMemoryError.java:48) at com.google.protobuf.ByteString.copyFrom
(ByteString.java:192) at com.google.protobuf.CodedInputStream.readBytes (CodedInputStream.java:324)
at org.apache.hadoop.yarn.proto.YarnProtos$StringStringMapProto.<init> (YarnProtos.java:47069)
at org.apache.hadoop.yarn.proto.YarnProtos$StringStringMapProto.<init> (YarnProtos.java:47014)
at org.apache.hadoop.yarn.proto.YarnProtos$StringStringMapProto$1.parsePartialFrom (YarnProtos.java:47102)
at org.apache.hadoop.yarn.proto.YarnProtos$StringStringMapProto$1.parsePartialFrom (YarnProtos.java:47097)
at com.google.protobuf.CodedInputStream.readMessage (CodedInputStream.java:309) at org.apache.hadoop.yarn.proto.YarnProtos$ContainerLaunchContextProto.<init>
(YarnProtos.java:41016) at org.apache.hadoop.yarn.proto.YarnProtos$ContainerLaunchContextProto.<init>
(YarnProtos.java:40942) at org.apache.hadoop.yarn.proto.YarnProtos$ContainerLaunchContextProto$1.parsePartialFrom
(YarnProtos.java:41080) at org.apache.hadoop.yarn.proto.YarnProtos$ContainerLaunchContextProto$1.parsePartialFrom
(YarnProtos.java:41075) at com.google.protobuf.CodedInputStream.readMessage (CodedInputStream.java:309)
at org.apache.hadoop.yarn.proto.YarnServiceProtos$StartContainerRequestProto.<init>
(YarnServiceProtos.java:24517) at org.apache.hadoop.yarn.proto.YarnServiceProtos$StartContainerRequestProto.<init>
(YarnServiceProtos.java:24464) at org.apache.hadoop.yarn.proto.YarnServiceProtos$StartContainerRequestProto$1.parsePartialFrom
(YarnServiceProtos.java:24568) at org.apache.hadoop.yarn.proto.YarnServiceProtos$StartContainerRequestProto$1.parsePartialFrom
(YarnServiceProtos.java:24563) at com.google.protobuf.AbstractParser.parsePartialFrom (AbstractParser.java:141)
at com.google.protobuf.AbstractParser.parseFrom (AbstractParser.java:176) at com.google.protobuf.AbstractParser.parseFrom
(AbstractParser.java:188) at com.google.protobuf.AbstractParser.parseFrom (AbstractParser.java:193)
at com.google.protobuf.AbstractParser.parseFrom (AbstractParser.java:49) at org.apache.hadoop.yarn.proto.YarnServiceProtos$StartContainerRequestProto.parseFrom
(YarnServiceProtos.java:24739) at org.apache.hadoop.yarn.server.nodemanager.recovery.NMLeveldbStateStoreService.loadContainerState
(NMLeveldbStateStoreService.java:217) at org.apache.hadoop.yarn.server.nodemanager.recovery.NMLeveldbStateStoreService.loadContainersState
(NMLeveldbStateStoreService.java:170) at org.apache.hadoop.yarn.server.nodemanager.containermanager.ContainerManagerImpl.recover
(ContainerManagerImpl.java:253) at org.apache.hadoop.yarn.server.nodemanager.containermanager.ContainerManagerImpl.serviceInit
(ContainerManagerImpl.java:237) at org.apache.hadoop.service.AbstractService.init (AbstractService.java:163)
at org.apache.hadoop.service.CompositeService.serviceInit (CompositeService.java:107) at org.apache.hadoop.yarn.server.nodemanager.NodeManager.serviceInit
(NodeManager.java:255) at org.apache.hadoop.service.AbstractService.init (AbstractService.java:163)
at org.apache.hadoop.yarn.server.nodemanager.NodeManager.initAndStartNodeManager (NodeManager.java:474)
at org.apache.hadoop.yarn.server.nodemanager.NodeManager.main (NodeManager.java:521){code}



--
This message was sent by Atlassian JIRA
(v7.6.3#76005)

---------------------------------------------------------------------
To unsubscribe, e-mail: yarn-issues-unsubscribe@hadoop.apache.org
For additional commands, e-mail: yarn-issues-help@hadoop.apache.org


Mime
View raw message