Y. Lan, L. Tang, N. Zhang and Franz Franchetti (Proc. High Performance Extreme Computing (HPEC), 2026)
Double-Precision Floating-Point on AWS Trainium Using Integer Units: A First Look
Comment: Paper with poster
Preprint (2 MB)
Bibtex

Modern accelerators increasingly emphasize low- precision arithmetic for performance and energy efficiency, while many scientific computing workloads still require double precision. This creates a need to understand whether software- emulated double-precision arithmetic can be efficiently mapped to such architectures. In this paper, we present a first-look study of double-precision floating-point (FP64) arithmetic using integer units on AWS Trainium through the custom C++ operator flow targeting its GPSIMD engine. This approach provides FP64- level accuracy on hardware that primarily targets lower-precision computation. Beyond accuracy, our results show that the batched radix-4 software-emulated floating-point kernel achieves a 1.88 times speedup over MKL-based CPU baselines for the equivalent 1D 4-point FFT computation on the Trainium host CPU. These results provide early evidence that our software-based approach can extend Trainium to FP64 kernels, while also revealing custom-operator limitations for larger FFT pipelines.

Keywords:
Fast Fourier Transform, Floating-point emulation, Extended precision, AWS Trainium