Return-Path: X-Original-To: archive-asf-public-internal@cust-asf2.ponee.io Delivered-To: archive-asf-public-internal@cust-asf2.ponee.io Received: from cust-asf.ponee.io (cust-asf.ponee.io [163.172.22.183]) by cust-asf2.ponee.io (Postfix) with ESMTP id 6A5FA200B46 for ; Sat, 2 Jul 2016 01:21:12 +0200 (CEST) Received: by cust-asf.ponee.io (Postfix) id 68F5A160A6C; Fri, 1 Jul 2016 23:21:12 +0000 (UTC) Delivered-To: archive-asf-public@cust-asf.ponee.io Received: from mail.apache.org (hermes.apache.org [140.211.11.3]) by cust-asf.ponee.io (Postfix) with SMTP id D8E1F160A61 for ; Sat, 2 Jul 2016 01:21:11 +0200 (CEST) Received: (qmail 84688 invoked by uid 500); 1 Jul 2016 23:21:11 -0000 Mailing-List: contact issues-help@mesos.apache.org; run by ezmlm Precedence: bulk List-Help: List-Unsubscribe: List-Post: List-Id: Reply-To: dev@mesos.apache.org Delivered-To: mailing list issues@mesos.apache.org Received: (qmail 84677 invoked by uid 99); 1 Jul 2016 23:21:11 -0000 Received: from arcas.apache.org (HELO arcas) (140.211.11.28) by apache.org (qpsmtpd/0.29) with ESMTP; Fri, 01 Jul 2016 23:21:11 +0000 Received: from arcas.apache.org (localhost [127.0.0.1]) by arcas (Postfix) with ESMTP id EF5022C027F for ; Fri, 1 Jul 2016 23:21:10 +0000 (UTC) Date: Fri, 1 Jul 2016 23:21:10 +0000 (UTC) From: "Jie Yu (JIRA)" To: issues@mesos.apache.org Message-ID: In-Reply-To: References: Subject: [jira] [Commented] (MESOS-5763) Task stuck in fetching is not cleaned up after --executor_registration_timeout. MIME-Version: 1.0 Content-Type: text/plain; charset=utf-8 Content-Transfer-Encoding: 7bit X-JIRA-FingerPrint: 30527f35849b9dde25b450d4833f0394 archived-at: Fri, 01 Jul 2016 23:21:12 -0000 [ https://issues.apache.org/jira/browse/MESOS-5763?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=15359818#comment-15359818 ] Jie Yu commented on MESOS-5763: ------------------------------- Yep, definitely a bug to me. We'll need to backport it to 0.28.x and 0.27.x. Older releases are no longer supported. > Task stuck in fetching is not cleaned up after --executor_registration_timeout. > ------------------------------------------------------------------------------- > > Key: MESOS-5763 > URL: https://issues.apache.org/jira/browse/MESOS-5763 > Project: Mesos > Issue Type: Bug > Components: containerization > Affects Versions: 0.28.0, 1.0.0, 0.29.0 > Reporter: Yan Xu > Assignee: Yan Xu > Priority: Critical > Fix For: 0.28.3, 1.0.0, 0.27.4 > > > When the fetching process hangs forever due to reasons such as HDFS issues, Mesos containerizer would attempt to destroy the container and kill the executor after {{--executor_registration_timeout}}. However this reliably fails for us: the executor would be killed by the launcher destroy and the container would be destroyed but the agent would never find out that the executor is terminated thus leaving the task in the STAGING state forever. -- This message was sent by Atlassian JIRA (v6.3.4#6332)