hadoop-common-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "Peng Zhang (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (HADOOP-10584) ActiveStandbyElector goes down if ZK quorum become unavailable
Date Sun, 04 Jan 2015 09:11:35 GMT

    [ https://issues.apache.org/jira/browse/HADOOP-10584?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14263809#comment-14263809
] 

Peng Zhang commented on HADOOP-10584:
-------------------------------------

I met the similar error for YARN RM that enabled HA automatic-failover. 

{noformat}
2015-01-04,12:42:30,682 FATAL org.apache.hadoop.ha.ActiveStandbyElector: Received stat error
from Zookeeper. code:CONNECTIONLOSS. Not retrying further znode monitoring connection errors.
2015-01-04,12:42:30,886 INFO org.apache.zookeeper.ZooKeeper: Session: 0x2498936f2a8c448 closed
2015-01-04,12:42:30,888 INFO org.apache.zookeeper.ClientCnxn: EventThread shut down
2015-01-04,12:42:30,888 FATAL org.apache.hadoop.yarn.server.resourcemanager.ResourceManager:
Received a org.apache.hadoop.yarn.server.resourcemanager.RMFatalEvent of type EMBEDDED_ELECTOR_FAILED.
Cause:
Received stat error from Zookeeper. code:CONNECTIONLOSS. Not retrying further znode monitoring
connection errors.
2015-01-04,12:42:30,891 INFO org.apache.hadoop.util.ExitUtil: Exiting with status 1
{noformat}

> ActiveStandbyElector goes down if ZK quorum become unavailable
> --------------------------------------------------------------
>
>                 Key: HADOOP-10584
>                 URL: https://issues.apache.org/jira/browse/HADOOP-10584
>             Project: Hadoop Common
>          Issue Type: Bug
>          Components: ha
>    Affects Versions: 2.4.0
>            Reporter: Karthik Kambatla
>            Assignee: Karthik Kambatla
>            Priority: Critical
>         Attachments: hadoop-10584-prelim.patch
>
>
> ActiveStandbyElector retries operations for a few times. If the ZK quorum itself is down,
it goes down and the daemons will have to be brought up again. 
> Instead, it should log the fact that it is unable to talk to ZK, call becomeStandby on
its client, and continue to attempt connecting to ZK.



--
This message was sent by Atlassian JIRA
(v6.3.4#6332)

Mime
View raw message