Copyrights to these papers may be held by the publishers. The download files are preprints. It is understood that all persons copying this information will adhere to the terms and constraints invoked by each author's copyright. These works may not be reposted without the explicit permission of the copyright holder.
Y. Lan, L. Tang, N. Zhang and Franz Franchetti (Proc. High Performance Extreme Computing (HPEC), 2026)
Double-Precision Floating-Point on AWS Trainium Using Integer Units: A First Look
Comment: Paper with poster
Preprint (2 MB)
Bibtex
Modern accelerators increasingly emphasize low- precision arithmetic for performance and energy efficiency, while many scientific computing workloads still require double precision. This creates a need to understand whether software- emulated double-precision arithmetic can be efficiently mapped to such architectures. In this paper, we present a first-look study of double-precision floating-point (FP64) arithmetic using integer units on AWS Trainium through the custom C++ operator flow targeting its GPSIMD engine. This approach provides FP64- level accuracy on hardware that primarily targets lower-precision computation. Beyond accuracy, our results show that the batched radix-4 software-emulated floating-point kernel achieves a 1.88 times speedup over MKL-based CPU baselines for the equivalent 1D 4-point FFT computation on the Trainium host CPU. These results provide early evidence that our software-based approach can extend Trainium to FP64 kernels, while also revealing custom-operator limitations for larger FFT pipelines.
Keywords: Fast Fourier Transform, Floating-point emulation, Extended precision, AWS Trainium