ComputingSummer Publishing Program, August 18, 2026

I Made the World’s First Exascale Supercomputer Run Faster From My Living Room

Frontier copies GPU data to its CPUs just to add numbers up. An undergraduate on a small allocation made that step run 50 times faster across 8,192 nodes, on her own time.

Published
August 18, 2026
Series
Summer Publishing Program
Licence
CC BY 4.0

Every argument about the future of computing is an argument about hardware. More chips, more datacenters, more power plants. The four biggest cloud companies are projected to spend roughly $725 billion in capital, most of it for AI infrastructure, and the federal government has proposed another $1.2 billion for new AI supercomputers.[1,2] Almost nobody is arguing about the software. That’s where the cheapest performance in computing is currently sitting, and nobody is paid to go get it.

I know because I spent two summers going and getting some of it

A supercomputer is not just an extra-large computer. Frontier, the Department of Energy system at Oak Ridge National Laboratory that was the first in the world to break the exascale barrier of a quintillion operations per second, is really 9,408 separate computers that have to talk to each other constantly.[3,4] The language they speak is called Message-Passing Interface (MPI), the de-facto communication standard for high performance computing.[5] One of the most important operations is Allreduce, which takes a number from every processor, combines them, and sends the result back. It’s the most heavily used collective operation in production science.[6] Every step of training a large model, from climate simulations to AI training, includes thousands of graphics processing units (GPUs) stopping to agree on a single set of numbers.

Here’s the strange part. Frontier’s power comes from its GPUs, clustered in eights (also called a node) with tens of thousands of nodes in total.[4] When MPI performs an Allreduce on GPU data, it copies that data back to the machine’s conventional central processing units (CPUs), does the arithmetic much slower there, and copies the result out again.[7] It’s like owning a sports car and towing it everywhere with a truck.

Last summer, as an intern in the Mathematics and Computer Science division at Argonne National Laboratory, I helped fix that by integrating a GPU-based library into MPICH, Argonne’s open-source MPI implementation. With my design, small messages still run on the CPU, where the latency is lower, but larger ones leverage the GPU for faster performance. The code has since been merged into MPICH and is accessible to the public.[8]

However, I had only validated the results on two nodes. So this summer, on a small research allocation and entirely on my own time, I ran experiments sweeping from one node up to 8,192, about 83% of Frontier. The advantage still held. At a thousand nodes and beyond, the GPU path ran the large messages that dominate modern workloads roughly 50× faster than the default. On a 1.4 GB gradient, corresponding to a full-precision BERT-Large model, my tuned path outpaced even Frontier’s own vendor library by 2-3×.

AI training at tech companies now runs on vendor libraries, and teams of engineers employed to keep them fast.[9] Public research has no equivalent. Millions of lines of physics, climate, and biology code, written over decades, are all built to speak MPI,[5] so its maintenance is key. My entire investigation cost the equivalent of five hours of the full machine’s time. I ran it as an undergraduate for free, but that’s not a maintainable funding model.

Meanwhile the money is moving the other way. The 2027 budget request would cut the DOE Office of Science by 15% to $7.14 billion.[10] Argonne and Fermilab offered staff buyouts last fall after their budgets dropped.[11] The entire Office of Science, essentially all of American public physics, chemistry, and computing, costs about 1% of what the cloud companies will spend on hardware this year.

The performance we’re building power plants for isn’t all locked in unbuilt silicon. Some of it is sitting in code we already own, on computers we already built, waiting for someone whose job it is to look. A supercomputer without software people is just a very expensive way to heat a building.

References

  1. Financial Times analysis, reported in Tom’s Hardware, “Google, Microsoft, Meta, and Amazon capex spending to hit $725 billion in 2026, up 77% from last year,” April 2026, https://www.tomshardware.com/tech-industry/big-tech/big-techs-ai-spending-plans-reach-725-billion.
  2. American Institute of Physics, FYI, “DOE Proposes Boosts for Supercomputers, Cuts to Research,” May 2026, https://www.aip.org/fyi/doe-proposes-boosts-for-supercomputers-cuts-to-research.
  3. Oak Ridge Leadership Computing Facility, “Frontier,” https://www.olcf.ornl.gov/olcf-resources/compute-systems/frontier/.
  4. Oak Ridge Leadership Computing Facility, Frontier User Guide, “System Overview,” https://docs.olcf.ornl.gov/systems/frontier_user_guide.html.
  5. I. Laguna, R. Marshall, K. Mohror, M. Ruefenacht, A. Skjellum, and N. Sultana, “A Large-Scale Study of MPI Usage in Open-Source HPC Applications,” SC ‘19, https://www.osti.gov/servlets/purl/1575876.
  6. S. Chunduri, S. Parker, P. Balaji, K. Harms, and K. Kumaran, “Characterization of MPI Usage on a Production Supercomputer,” SC ‘18, https://pavanbalaji.github.io/pubs/2018/sc/sc18.mpiusage.pdf.
  7. S. Singh, K. Pradeep, M. Singh, C. Wei, and A. Bhatele, “The Big Send-off: Scalable and Performant Collectives for Deep Learning,” §III.B, https://arxiv.org/abs/2504.18658.
  8. MPICH pull request #7493, “RCCL Allreduce Backend Implementation,” pmodels/mpich, https://github.com/pmodels/mpich/pull/7493.
  9. PyTorch, “Distributed Communication Package,” https://docs.pytorch.org/docs/stable/distributed.html.
  10. Computing Research Association, Government Affairs, “Department of Energy FY 2027 Request: Office of Science Faces Substantial Cuts; Slight Reduction for ASCR,” May 2026, https://cra.org/govaffairs/blog/2026/05/doe-sc-fy2027-pbr/.
  11. WTTW News, “Staff Shakeup at Fermilab and Argonne as Buyouts Follow Budgeted Funding Drop, Federal Research Shift,” September 9, 2025, https://news.wttw.com/2025/09/09/staff-shakeup-fermilab-and-argonne-buyouts-follow-budgeted-funding-drop-federal-research.

How to cite this article

Li, G. (2026). I Made the World’s First Exascale Supercomputer Run Faster From My Living Room. Columbia Scientist, Summer Publishing Program. https://columbiascientist.org/articles/exascale-supercomputer-living-room

© 2026 Grace Li. This is an open-access article distributed under the terms of the Creative Commons Attribution 4.0 International licence, which permits use, distribution, and reproduction in any medium, provided the original author and source are credited.

More from the Summer Publishing Program

All SPP pieces →