flink-issues mailing list archives

Site index · List index
Message view « Date » · « Thread »
Top « Date » · « Thread »
From "ASF GitHub Bot (JIRA)" <j...@apache.org>
Subject [jira] [Commented] (FLINK-7231) SlotSharingGroups are not always released in time for new restarts
Date Thu, 20 Jul 2017 13:33:00 GMT

    [ https://issues.apache.org/jira/browse/FLINK-7231?page=com.atlassian.jira.plugin.system.issuetabpanels:comment-tabpanel&focusedCommentId=16094688#comment-16094688
] 

ASF GitHub Bot commented on FLINK-7231:
---------------------------------------

Github user StephanEwen commented on a diff in the pull request:

    https://github.com/apache/flink/pull/4370#discussion_r128514948
  
    --- Diff: flink-runtime/src/test/java/org/apache/flink/runtime/executiongraph/ExecutionGraphRestartTest.java
---
    @@ -661,10 +682,117 @@ public void testConcurrentGlobalFailAndRestarts() throws Exception
{
     		}
     	}
     
    +	@Test
    +	public void testRestartWithEagerSchedulingAndSlotSharing() throws Exception {
    --- End diff --
    
    Exactly, it extends test coverage to possible related conditions.
    I realized that we had no test about the "happy path" of eager scheduling and slot sharing
as well.


> SlotSharingGroups are not always released in time for new restarts
> ------------------------------------------------------------------
>
>                 Key: FLINK-7231
>                 URL: https://issues.apache.org/jira/browse/FLINK-7231
>             Project: Flink
>          Issue Type: Bug
>          Components: Distributed Coordination
>    Affects Versions: 1.3.1
>            Reporter: Stephan Ewen
>            Assignee: Stephan Ewen
>             Fix For: 1.4.0, 1.3.2
>
>
> In the case where there are not enough resources to schedule the streaming program, a
race condition can lead to a sequence of the following errors:
> {code}
> java.lang.IllegalStateException: SlotSharingGroup cannot clear task assignment, group
still has allocated resources.
> {code}
> This eventually recovers, but may involve many fast restart attempts before doing so.
> The root cause is that slots are not cleared before the next restart attempt.



--
This message was sent by Atlassian JIRA
(v6.4.14#64029)

Mime
View raw message