BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20230124T171522Z
LOCATION:C143-149
DTSTART;TZID=America/Chicago:20221114T160900
DTEND;TZID=America/Chicago:20221114T163400
UID:submissions.supercomputing.org_SC22_sess461_ws_ftxs102@linklings.com
SUMMARY:ReStore: In-Memory REplicated STORagE for Rapid Recovery in Fault-
 Tolerant Algorithms
DESCRIPTION:Workshop\n\nReStore: In-Memory REplicated STORagE for Rapid Re
 covery in Fault-Tolerant Algorithms\n\nHübner, Hespe, Sanders, Stamatakis\
 n\nFault-tolerant applications need to recover data lost after process fai
 lures. It is typically impractical to request replacement resources after 
 a failure. Therefore, applications have to continue with the remaining res
 ources. This requires redistributing the workload. We present an algorithm
 ic framework and its C++ implementation ReStore that enables recovery of d
 ata after process failures. By storing all required data in memory via an 
 appropriate data distribution and replication, recovery is substantially f
 aster than with standard checkpointing schemes that rely on a parallel fil
 e system. As the application developer can specify which data to load, we 
 also support shrinking recovery instead of recovery using spare compute no
 des. Our experiments show loading times of lost input data in the range of
  milliseconds on up to 24,576 processors and a substantial speedup of the 
 recovery time for the fault-tolerant version of a widely used bioinformati
 cs application.\n\nSession Format: Recorded\n\nRegistration Category: Work
 shop Reg Pass
END:VEVENT
END:VCALENDAR
