×

AMD × Dayhoff Health

Accelerating Single-Cell Tumor Classification with AMD ROCm™ and HIP

Download the joint AMD × Dayhoff Health white paper on GPU-accelerating the 10x Cell Ranger Flex pipeline and the CopyKAT tumor-cell classifier - measured on a real archival breast tumor of 1.03 billion reads across four multiplexed FFPE regions.

  • 6.59× end to end - 76.1 minutes on CPU to 11.6 on two GPUs
  • 6.41× faster Cell Ranger, FASTQ to count matrix
  • Deterministic tumor calls byte-identical to CPU
Single-cell data accelerated by AMD GPUs into tumor and normal cell calls

Get Your Free White Paper

We respect your privacy. Your information is safe with us.

Validated Speed, Not Just Faster Numbers

Classifying malignant cells takes two compute-heavy steps in sequence: Cell Ranger turns raw Flex reads into a gene-by-cell count matrix, then CopyKAT calls each cell aneuploid or diploid from its inferred copy-number profile. Both are CPU-bound jobs repeating the same arithmetic over millions of reads or thousands of cells. This paper measures what moving them onto AMD Radeon™ GPUs actually delivers - holding to one rule throughout: validate every kernel against the exact CPU function it replaces before reporting any speedup.

Workload CPU One GPU Two GPUs
Cell Ranger
complete output
3900.548 s 608.550 s6.41× 528.965 s7.37×
CopyKAT
all-four standard batch
645.382 s 279.608 s2.31× 148.355 s4.35×
Complete workflow
FASTQ to tumor calls
76.1 min 15.0 min5.09× 11.6 min6.59×

The CPU reference and all one- and two-GPU repetitions are continuously timed complete-workflow runs. GPU figures are three-run medians.

Where the Gains Come From

Removing the count path's disk round-trip, mapping millions of cell-pair distances onto the GPU, and running independent regions concurrently across two devices - plus the stages that saw no repeatable gain.

Correctness Before Speed

Per-gene correlation of 1.00000 with roughly 99% element-level parity, deterministic CopyKAT predictions byte-identical across CPU and GPU, and the precision, sensitivity and abstention figures behind the tumor calls.

Open Infrastructure

The complete workflow runs on one or two AMD Radeon AI PRO R9700S devices (gfx1201, RDNA4) on the open ROCm™ 7.x stack, with HIP kernels compiled for the target architecture.