Commit
This commit does not belong to any branch on this repository, and may belong to a fork outside of the repository.
sched: Fix lockup by limiting load-balance retries on lock-break
Eric and David reported dead machines and traced it to commit a195f00 ("sched: Fix load-balance lock-breaking"), it turns out there's still a scenario where we can end up re-trying forever. Since there is no strict forward progress guarantee in the load-balance iteration we can get stuck re-retrying the same task-set over and over. Creating a forward progress guarantee with the existing structure is somewhat non-trivial, for now simply terminate the retry loop after a few tries. Reported-by: Eric Dumazet <[email protected]> Tested-by: Eric Dumazet <[email protected]> Reported-by: David Ahern <[email protected]> [ logic cleanup as suggested by Eric ] Signed-off-by: Peter Zijlstra <[email protected]> Cc: Linus Torvalds <[email protected]> Cc: Martin Schwidefsky <[email protected]> Cc: Frederic Weisbecker <[email protected]> Cc: Suresh Siddha <[email protected]> Link: http://lkml.kernel.org/r/1326297936.2442.157.camel@twins Signed-off-by: Ingo Molnar <[email protected]>
- Loading branch information