AMD × Dayhoff Health
Download the joint AMD × Dayhoff Health white paper on GPU-accelerating the 10x Cell Ranger Flex pipeline and the CopyKAT tumor-cell classifier - measured on a real archival breast tumor of 1.03 billion reads across four multiplexed FFPE regions.
We respect your privacy. Your information is safe with us.
Classifying malignant cells takes two compute-heavy steps in sequence: Cell Ranger turns raw Flex reads into a gene-by-cell count matrix, then CopyKAT calls each cell aneuploid or diploid from its inferred copy-number profile. Both are CPU-bound jobs repeating the same arithmetic over millions of reads or thousands of cells. This paper measures what moving them onto AMD Radeon™ GPUs actually delivers - holding to one rule throughout: validate every kernel against the exact CPU function it replaces before reporting any speedup.
| Workload | CPU | One GPU | Two GPUs |
|---|---|---|---|
| Cell Ranger complete output | 3900.548 s | 608.550 s6.41× | 528.965 s7.37× |
| CopyKAT all-four standard batch | 645.382 s | 279.608 s2.31× | 148.355 s4.35× |
| Complete workflow FASTQ to tumor calls | 76.1 min | 15.0 min5.09× | 11.6 min6.59× |
The CPU reference and all one- and two-GPU repetitions are continuously timed complete-workflow runs. GPU figures are three-run medians.
Removing the count path's disk round-trip, mapping millions of cell-pair distances onto the GPU, and running independent regions concurrently across two devices - plus the stages that saw no repeatable gain.
Per-gene correlation of 1.00000 with roughly 99% element-level parity, deterministic CopyKAT predictions byte-identical across CPU and GPU, and the precision, sensitivity and abstention figures behind the tumor calls.
The complete workflow runs on one or two AMD Radeon AI PRO R9700S devices (gfx1201, RDNA4) on the open ROCm™ 7.x stack, with HIP kernels compiled for the target architecture.