BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20230124T170804Z
LOCATION:C1-2-3
DTSTART;TZID=America/Chicago:20221116T083000
DTEND;TZID=America/Chicago:20221116T170000
UID:submissions.supercomputing.org_SC22_sess274_rpost134@linklings.com
SUMMARY:Case Study for Performance-Portability of Lattice Boltzmann Kernel
 s
DESCRIPTION:Posters, Research Posters\n\nCase Study for Performance-Portab
 ility of Lattice Boltzmann Kernels\n\nLiu, Randles, Insley, Patel, Rizzi..
 .\n\nIn this work, we study the performance-portability of offloaded latti
 ce Boltzmann kernels and the trade-off between portability and efficiency.
  The study is based on a proxy application for the lattice Boltzmann metho
 d (LBM). The performance portability programming framework of Kokkos (with
  CUDA or SYCL backend) is used and compared with programming models of nat
 ive CUDA and native SYCL. The Kokkos library supports the mainstream GPU p
 roducts in the market. The performance of the code can vary with accelerat
 ing models, number of GPUs, scale of the problem, propagation patterns and
  architectures. Both Kokkos library and CUDA toolkit are studied on the su
 percomputer of ThetaGPU (Argonne Leadership Computing Facility). It is fou
 nd that Kokkos (CUDA) has almost the same performance as native CUDA. The 
 automatic data and kernel management in Kokkos may sacrifice the efficienc
 y, but the parallelization parameters can also be tuned by Kokkos to optim
 ize the performances.\n\nRegistration Category: Tech Program Reg Pass, Exh
 ibits Reg Pass
END:VEVENT
END:VCALENDAR
