Fault Tolerance in Grids Using Job Replication
Loading...
Files
Date
Authors
Journal Title
Journal ISSN
Volume Title
Publisher
ТНЕУ
Abstract
As grids consist of a large number of resources, fault tolerance forms an important aspect of the scheduling
process. In this paper, we address the problem of scheduling user jobs in grids so that failures can be avoided in the
presence of resources faults. We employ job replication as an effective mechanism to achieve efficient and fault-tolerant
scheduling system. Most of the existing replication-based algorithms use a fixed number of replications for each job
which consumes more grid resources. We first propose an algorithm to determine adaptively the number of job replicas
according to the grid failure history. Then we propose an algorithm to schedule these replicas. The proposed
algorithms have been evaluated through simulation and have shown better performance in terms of grid load,
throughput and failure tendency.
Description
Keywords
Citation
Amoon, М. Fault Tolerance in Grids Using Job Replication [Text] / Mohammed Amoon // Computing = Комп’ютинг. - 2012. - Vol. 11, is. 2. - P. 115-121.