BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20230124T171526Z
LOCATION:D222
DTSTART;TZID=America/Chicago:20221113T143000
DTEND;TZID=America/Chicago:20221113T150000
UID:submissions.supercomputing.org_SC22_sess436_ws_mchpc101@linklings.com
SUMMARY:Maximizing Performance Through Memory Hierarchy-Driven Data Layout
  Transformations
DESCRIPTION:Workshop\n\nMaximizing Performance Through Memory Hierarchy-Dr
 iven Data Layout Transformations\n\nSepanski, Zhao, Johansen, Williams\n\n
 Computations on structured grids using standard multidimensional array lay
 outs can incur substantial data movement costs through the memory hierarch
 y. This presentation explores the benefits of using a framework (Bricks) t
 o separate the complexity of data layout and optimized communication from 
 the functional representation. To that end, we provide three novel contrib
 utions and evaluate them on several kernels taken from GENE, a phase-space
  fusion tokamak simulation code. We extend Bricks to support 6-dimensional
  arrays and kernels that operate on complex data types, and integrate Bric
 ks with cuFFT. We demonstrate how to optimize Bricks for data reuse, spati
 al locality, and GPU hardware utilization achieving up to a 2.67× speedup 
 on a single A100 GPU. We conclude with insights on how to rearchitect memo
 ry subsystems.\n\nSession Format: Recorded\n\nRegistration Category: Works
 hop Reg Pass
END:VEVENT
END:VCALENDAR
