SupremeRAID™ GPU-Accelerated RAID Performance Evaluation

Benchmark comparison of SupremeRAID™, Linux MD RAID and a single NVMe-oF over a GP-SPARK direct-attach connection

Key Finding: SupremeRAID™'s Advantage in Parity RAID

On RAID 5 / RAID 6 write performance, SupremeRAID™ (GPU offload) substantially outperforms MDRAID (CPU software RAID).

RAID 5 Sequential Write
Intel Machine 4.4×
DGX Spark 6.6×
RAID 6 Sequential Write
Intel Machine 4.5×
DGX Spark 5.7×
RAID 5 Random Write
Intel Machine 8.4×
DGX Spark 6.4×
RAID 6 Random Write
Intel Machine 5.7×
DGX Spark 5.1×

SupremeRAID™ throughput expressed as a multiple of MDRAID throughput.

GPU-offloaded parity computation substantially mitigates the write performance degradation MDRAID experiences on RAID 5/6, where the CPU becomes the bottleneck.

Sequential Write Bandwidth Comparison

SupremeRAID™ vs. MDRAID at the same RAID level. MB/s; bs=2M, numjobs=40, iodepth=32, 60s per pattern.

SupremeRAID™ MDRAID
RAID 5 · Intel Machine
2926
660
RAID 5 · DGX Spark
2930
446
RAID 6 · Intel Machine
2389
536
RAID 6 · DGX Spark
1964
343

The chart covers the parity levels only (RAID 5 / RAID 6); RAID 0 and the single-NVMe references are in the table below.

Performance Summary (Sequential, MB/s)

Configuration Intel Read Intel Write DGX Read DGX Write
SupremeRAID™ RAID 0 8080 4731 7805 4804
SupremeRAID™ RAID 5 7906 2926 7725 2930
SupremeRAID™ RAID 6 8129 2389 7687 1964
MDRAID RAID 0 7990 4811 7401 4821
MDRAID RAID 5 7202 660 7399 446
MDRAID RAID 6 7280 536 7467 343
Single NVMe-oF 2060 1448 1923 1438
Single NVMe (Local) 6302 2402 11700 12700

RAID 0 shows near-parity across both approaches (no parity computation); the SupremeRAID™ advantage becomes pronounced on RAID 5/6.

dd Throughput Cross-Validation

100 GiB file write × 5 runs, average of the steady-state runs 2-5 (GB/s), cross-validating the fio trend.

Configuration SupremeRAID™ MDRAID
RAID 5 ≈ 2.95 – 3.00 GB/s ≈ 0.85 – 1.60 GB/s
RAID 6 ≈ 2.00 – 2.13 GB/s ≈ 0.76 – 1.35 GB/s

Values are approximate readings from the source report's chart, spanning both test environments; indicative of the trend only.

GP-SPARK - Test Environment

Intel Machine Environment
CPU
Intel Xeon Platinum 8380 ×2 (40 cores/socket, 2.30 GHz)
Memory
256 GB
OS Drive
1.92 TB NVMe SSD
GPU
NVIDIA A40 (Driver 610.43.02)
Connection
Direct-attach to GP-SPARK
OS
Ubuntu 24.04.4
DGX Spark Environment
CPU
NVIDIA Cortex-X925 3.9 GHz (10 cores)
Memory
128 GB (Unified Memory)
OS Drive
4 TB NVMe SSD
GPU
NVIDIA GB10 (Driver 580.159.03)
Connection
Direct-attach to GP-SPARK
OS
Ubuntu 24.04.4

GP-SPARK - Test Methodology

fio Benchmark
I/O Engine
libaio (direct=1)
Block Size
2 MiB
File Size
20 GiB / job
Parallel Jobs
40
I/O Depth
32
Runtime
60s per pattern (30s cooldown between)
Patterns Tested
seq_write / rand_write / seq_read / rand_read
dd Throughput Test
Command
dd if=/dev/zero of=/data/data-N.dat bs=1M count=100K
Transfer Size
100 GiB / run
Repetitions
5 consecutive runs
Aggregation
Average of steady-state runs (2nd-5th)

Configurations tested (both environments)

  • SupremeRAID™ RAID 0 / 5 / 6 (GPU-accelerated RAID)
  • Linux MD RAID 0 / 5 / 6 (software RAID)
  • Single NVMe-oF (comparison baseline)
  • Single NVMe (local disk, reference only)

Notes

  • The figures above are taken from the HPCTECH Corporation report "SupremeRAID™ GPU-Accelerated RAID Performance Evaluation".
  • Actual performance varies with host platform, SSD model, network topology and OS tuning.
  • For a PoC in your own environment or for the full test report, please get in touch with us.