BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20230124T171522Z
LOCATION:C143-149
DTSTART;TZID=America/Chicago:20221114T155400
DTEND;TZID=America/Chicago:20221114T160900
UID:submissions.supercomputing.org_SC22_sess461_ws_ftxs104@linklings.com
SUMMARY:Toward Precision-Aware Fault Tolerance Approaches for Mixed-Precis
 ion Applications
DESCRIPTION:Workshop\n\nToward Precision-Aware Fault Tolerance Approaches 
 for Mixed-Precision Applications\n\nFang, Hari, Tsai, Li, Gopalakrishnan..
 .\n\nGraphics Processing Units (GPUs), the dominantly adopted accelerators
  in HPC systems, are susceptible to a transient hardware fault. A new gene
 ration of GPUs features mixed-precision architectures such as NVIDIA Tenso
 r Cores to accelerate matrix multiplications. While widely adapted, how th
 ey would behave under transient hardware faults remain unclear. In this st
 udy, we conduct large-scale fault injection experiments on GEMM kernels im
 plemented with different floating-point data types on the V100 and A100 Te
 nsor Cores and show distinct error resilience characteristics for the GEMM
 S with different formats. We plan to explore this space in the future by b
 uilding precision-aware floating-point fault tolerance techniques for appl
 ications such as DNNs that exercise low-precision computations.\n\nSession
  Format: Recorded\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
