Sciweavers

8329 search results - page 489 / 1666
» An ATM-based Distributed High Performance Computing System
Sort
View
JSSPP
2004
Springer
16 years 29 days ago
Performance Implications of Failures in Large-Scale Cluster Scheduling
As we continue to evolve into large-scale parallel systems, many of them employing hundreds of computing engines to take on mission-critical roles, it is crucial to design those s...
Yanyong Zhang, Mark S. Squillante, Anand Sivasubra...
ICPPW
2003
IEEE
16 years 27 days ago
Constructing Nondominated Local Coteries for Distributed Resource Allocation
The resource allocation problem is a fundamental problem in distributed systems. In this paper, we focus on constructing nondominated (ND) local coteries to solve the problem. Dis...
Jehn-Ruey Jiang, Cheng-Sheng Chou, Shing-Tsaan Hua...
IPPS
2007
IEEE
16 years 1 months ago
FixD : Fault Detection, Bug Reporting, and Recoverability for Distributed Applications
Model checking, logging, debugging, and checkpointing/recovery are great tools to identify bugs in small sequential programs. The direct application of these techniques to the dom...
Cristian Tapus, David A. Noblet
PODC
2003
ACM
16 years 26 days ago
Routing networks for distributed hash tables
Routing topologies for distributed hashing in peer-to-peer networks are classified into two categories: deterministic and randomized. A general technique for constructing determi...
Gurmeet Singh Manku
PPL
2002
83views more  PPL 2002»
15 years 7 months ago
Trading Replication for Communication in Parallel Distributed-Memory Dense Solvers
We present new communication-efficient parallel dense linear solvers: a solver for triangular linear systems with multiple right-hand sides and an LU factorization algorithm. Thes...
Dror Irony, Sivan Toledo