Mailing-List: contact dev-help@reef.apache.org; run by ezmlm
Precedence: bulk
Reply-To: dev@reef.apache.org
Date: Sun, 28 Feb 2016 00:50:18 +0000 (UTC)
From: "Julia (JIRA)" <jira@apache.org>
To: dev@reef.apache.org
Message-ID: <JIRA.12945223.1456620583000.166068.1456620618171@Atlassian.JIRA>
In-Reply-To: <JIRA.12945223.1456620583000@Atlassian.JIRA>
References: <JIRA.12945223.1456620583000@Atlassian.JIRA>
 <JIRA.12945223.1456620583147@arcas>
Subject: [jira] [Assigned] (REEF-1223) Fault tolerant - restart failed
 evaluators
MIME-Version: 1.0
Content-Type: text/plain; charset=utf-8
Content-Transfer-Encoding: 7bit


     [ https://issues.apache.org/jira/browse/REEF-1223?page=com.atlassian.jira.plugin.system.issuetabpanels:all-tabpanel ]

Julia reassigned REEF-1223:
---------------------------

    Assignee: Julia

> Fault tolerant - restart failed evaluators
> ------------------------------------------
>
>                 Key: REEF-1223
>                 URL: https://issues.apache.org/jira/browse/REEF-1223
>             Project: REEF
>          Issue Type: New Feature
>            Reporter: Julia
>            Assignee: Julia
>
> Currently in .Net Group Communication and IMRU scenario, if one of the Evaluator failed for whatever reason, all the Evaluators will be killed by the driver. 
> There are multiple levels of fault tolerant. The scenario we would like to support in this JIRA is:
> *  When an evaluator failed, this failed evaluator will be killed and other good Evaluators will stay, but all the tasks running on those Evaluators will be stopped. 
> *  A new Evaluator will be requested and started with the original task. 
> *  Same tasks will be resubmitted to the rest the Evaluators
> *  The topology of those tasks will be kept in the same group communication as before
> *  The data that have been downloaded in those good Evaluators will stay. 


--
This message was sent by Atlassian JIRA
(v6.3.4#6332)