BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Chicago
X-LIC-LOCATION:America/Chicago
BEGIN:DAYLIGHT
TZOFFSETFROM:-0600
TZOFFSETTO:-0500
TZNAME:CDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0500
TZOFFSETTO:-0600
TZNAME:CST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20230124T171520Z
LOCATION:C141
DTSTART;TZID=America/Chicago:20221113T160000
DTEND;TZID=America/Chicago:20221113T163000
UID:submissions.supercomputing.org_SC22_sess421_ws_drbsd110@linklings.com
SUMMARY:Dynamic Clustering-Based Sharding in Distributed Deduplication Sys
 tems
DESCRIPTION:Workshop\n\nDynamic Clustering-Based Sharding in Distributed D
 eduplication Systems\n\nZhou, Xia, Zou\n\nWith a massive upsurge in data, 
 combining deduplication with distributed storage continuously suffer from 
 a low deduplication ratio when providing the corresponding throughput.  It
  is because distributed storage requires sharding data on different nodes,
  while global deduplication needs eliminating redundancies in a unified vi
 ew.In this paper, we present clustering-based sharding method, D-Shard, in
  distributed deduplication storage systems that leads to a comparable dedu
 plication efficiency on a single system while supporting a high throughput
 . First, using Dynamic K-Means approach to cluster super-blocks, then extr
 acting every cluster center feature as the anchor point for sharding; Seco
 nd, Construct a secondary deduplication index based on the Compact Hamming
  Index.  Currently, preliminary results show that super-block clustering i
 s convergent, and routing strategy based on anchor points can achieve a hi
 gher deduplication ratio compared to the state-of-the-art approach and the
  throughput of system has been greatly improved.\n\nSession Format: Record
 ed\n\nRegistration Category: Workshop Reg Pass
END:VEVENT
END:VCALENDAR
