hadoop-common-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "lvchuanwen (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (HADOOP-10251) Both NameNodes could be in STANDBY State if SNN network is unstable
Date Fri, 12 Jun 2015 01:37:01 GMT

    [ https://issues.apache.org/jira/browse/HADOOP-10251?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=14582800#comment-14582800

lvchuanwen commented on HADOOP-10251:

Use Cases´╝Ü
1.NN1 was Active and NN2 was Standby ,kill NN1 . NN2 transition to active.
2.hadoop-daemon.sh start namenode NN2. NOW.NN1 was Standby and NN2 was Active .
3.kill NN2 ,NN1 transition to active.
Attaching hdfs-nn1-zkfc-host195.log file and hdfs-nn2-zkfc-host196.log file

> Both NameNodes could be in STANDBY State if SNN network is unstable
> -------------------------------------------------------------------
>                 Key: HADOOP-10251
>                 URL: https://issues.apache.org/jira/browse/HADOOP-10251
>             Project: Hadoop Common
>          Issue Type: Bug
>          Components: ha
>    Affects Versions: 2.2.0
>            Reporter: Vinayakumar B
>            Assignee: Vinayakumar B
>            Priority: Critical
>             Fix For: 2.5.0
>         Attachments: HADOOP-10251.patch, HADOOP-10251.patch, HADOOP-10251.patch, HADOOP-10251.patch,
> Following corner scenario happened in one of our cluster.
> 1. NN1 was Active and NN2 was Standby
> 2. NN2 machine's network was slow 
> 3. NN1 got shutdown.
> 4. NN2 ZKFC got the notification and trying to check for old active for fencing. (This
took little more time, again due to slow network)
> 5. In between, NN1 got restarted by our automatic monitoring, and ZKFC made it Active.
> 6. Now NN2 ZKFC got Old Active as NN1 and it did graceful fencing of NN1 to STANBY.
> 7. Before writing ActiveBreadCrumb to ZK, NN2 ZKFC got session timeout and got shutdown
before making NN2 Active.
> *Now cluster having both NameNodes as STANDBY.*
> NN1 ZKFC still thinks that its nameNode is in Active state. 
> NN2 ZKFC waiting for election.

This message was sent by Atlassian JIRA

View raw message