Skip to main navigation Skip to search Skip to main content

Accelerating erasure coding by exploiting multiple repair paths in distributed storage systems

  • Chanki Kim
  • , Kang Wook Chon*
  • *Corresponding author for this work
    • Korea University of Technology and Education

    Research output: Contribution to journalJournal articlepeer-review

    Abstract

    High reliability must be ensured in distributed storage systems (DSSs) to maintain the stability of warehouse-scale computing and high-performance computing (HPC) systems. For system-level reliability, a repair operation using redundant storage nodes can be used in conjunction with erasure coding (EC), which can also affect the system performance. The existing EC design mainly focused on minimizing the required bandwidth for the repair and storage overheads. However, the computing performance for EC should be considered to achieve high bandwidth in order to exploit back-end network link capacity with heterogeneous and high-speed interconnects over 10 Gbps Ethernet. In this study, a new computing acceleration method for repair operation in EC is proposed using multiple repair paths and modifying the computation kernel on the graphics processing unit (GPU) device. For the Cauchy Reed–Solomon (CRS) codes, the proposed scheme is observed to achieve sufficient repair bandwidth compared to the theoretical bound or exceed the current maximum Ethernet link bandwidth.

    Original languageEnglish
    Pages (from-to)8621-8635
    Number of pages15
    JournalCluster Computing
    Volume27
    Issue number6
    DOIs
    StatePublished - 2024.09

    Keywords

    • Cauchy Reed–Solomon (CRS) codes
    • Computing acceleration
    • Erasure coding
    • Fault-tolerant system
    • Graphics processing units (GPUs)
    • High-performance computing (HPC)

    Quacquarelli Symonds(QS) Subject Topics

    • Computer Science & Information Systems

    Fingerprint

    Dive into the research topics of 'Accelerating erasure coding by exploiting multiple repair paths in distributed storage systems'. Together they form a unique fingerprint.

    Cite this