Mailing-List: contact hdfs-issues-help@hadoop.apache.org; run by ezmlm
Precedence: bulk
Reply-To: hdfs-issues@hadoop.apache.org
Date: Mon, 18 Apr 2016 15:41:25 +0000 (UTC)
From: "Daryn Sharp (JIRA)" <jira@apache.org>
To: hdfs-issues@hadoop.apache.org
Message-ID: <JIRA.12959486.1460922644000.256221.1460994085546@Atlassian.JIRA>
In-Reply-To: <JIRA.12959486.1460922644000@Atlassian.JIRA>
References: <JIRA.12959486.1460922644000@Atlassian.JIRA>
 <JIRA.12959486.1460922644843@arcas>
Subject: [jira] [Commented] (HDFS-10301) Blocks removed by thousands due to
 falsely detected zombie storages
MIME-Version: 1.0
Content-Type: text/plain; charset=utf-8
Content-Transfer-Encoding: 7bit


    [ https://issues.apache.org/jira/browse/HDFS-10301?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15245867#comment-15245867 ] 

Daryn Sharp commented on HDFS-10301:
------------------------------------

Enabling HDFS-9198 will fifo process BRs.  It doesn't solve this implementation bug but virtually eliminates it from occurring.

> Blocks removed by thousands due to falsely detected zombie storages
> -------------------------------------------------------------------
>
>                 Key: HDFS-10301
>                 URL: https://issues.apache.org/jira/browse/HDFS-10301
>             Project: Hadoop HDFS
>          Issue Type: Bug
>          Components: namenode
>    Affects Versions: 2.6.1
>            Reporter: Konstantin Shvachko
>            Priority: Critical
>         Attachments: zombieStorageLogs.rtf
>
>
> When NameNode is busy a DataNode can timeout sending a block report. Then it sends the block report again. Then NameNode while process these two reports at the same time can interleave processing storages from different reports. This screws up the blockReportId field, which makes NameNode think that some storages are zombie. Replicas from zombie storages are immediately removed, causing missing blocks.


--
This message was sent by Atlassian JIRA
(v6.3.4#6332)