BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20230124T171519Z
LOCATION:C148
DTSTART;TZID=America/Chicago:20221114T145000
DTEND;TZID=America/Chicago:20221114T145500
UID:submissions.supercomputing.org_SC22_sess460_ws_pdswwip104@linklings.co
 m
SUMMARY:Revisit Data Partitioning in Data-Intensive Workflows
DESCRIPTION:Workshop\n\nRevisit Data Partitioning in Data-Intensive Workfl
 ows\n\nLiem, Ibrahim\n\nIn this work in progress, we will showcase a compr
 ehensive analysis of the current state-of-the-art solutions for data skew 
 mitigation in several environments. Our experiments and evaluation compris
 e several data-intensive workflows running on Spark using the Grid’5000 te
 stbed. The data-intensive workflows vary from a highly optimized WordCount
  application, an iterative application like PageRank, to an SQL-based deci
 sion support system benchmark, TPC-H with various sizes and configurations
 . Going forward,  we will discuss our current efforts toward heterogeneity
 -aware multi-stages data partitioning.\n\nSession Format: Recorded\n\nRegi
 stration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
