Date of Award

8-2026

Document Type

Dissertation

Degree Name

Doctor of Philosophy (PhD)

Department

Computer Engineering

Committee Chair/Advisor

Jon C. Calhoun

Committee Member

Melissa C. Smith

Committee Member

Rong Ge

Committee Member

Tao Wei

Committee Member

Bogdan Nicolae

Abstract

Modern high-performance computing (HPC) and artificial intelligence (AI) workloads increasingly generate data at rates that outpace the storage systems responsible for persisting it. While computational throughput has scaled dramatically through parallel processors, GPUs, and specialized accelerators, storage systems have improved far more slowly. As such, application performance is increasingly limited by the ability to move data efficiently through deep, heterogeneous storage hierarchies rather than by raw computational capability. This dissertation investigates how aggregation and data movement strategies must be redesigned to close this gap, addressing four central challenges: resource contention between foreground computation and background I/O, the fine-grained and irregular access patterns of modern workloads, the scalability limits of coordinating I/O across many concurrent processes, and the difficulty of moving data efficiently across increasingly heterogeneous memory and storage tiers. We address these challenges across three representative domains: large-scale HPC checkpointing, LLM training, and LLM inference. We first show that collective I/O techniques designed for synchronous checkpointing are poorly suited to asynchronous execution due to their reliance on globally coordinated synchronization. Through an extensive characterization of aggregation strategies, spanning thread concurrency, buffer management, data contiguity, and file organization, we design a hierarchical aggregation strategy for VeloC that partitions processes into independent I/O groups, avoiding global synchronization while achieving up to 2$\times$ higher checkpoint throughput than existing systems such as GenericIO and ADIOS2. We then extend these ideas to LLM training, where checkpoints are fragmented into hundreds or thousands of small, heterogeneous shards dictated by tensor, pipeline, and data parallelism. Characterizing how these workloads interact with modern kernel-level I/O mechanisms, including io\_uring, buffered I/O, and direct I/O, we uncover a counterintuitive result: checkpoint restoration can take nearly twice as long as checkpoint generation. Profiling production checkpointing frameworks traces this overhead not to storage bandwidth but to software inefficiencies, including per-tensor synchronization and repeated dynamic memory allocation. Guided by this analysis, we redesign the DataStates-LLM checkpoint engine and DeepSpeed's restore pipeline around batched, asynchronous I/O submission and reusable buffer pools, reducing end-to-end restoration time by more than a factor of two. Finally, we examine checkpoint loading for LLM inference, where model initialization has become an increasingly significant systems bottleneck. We characterize how checkpoint layout, filesystem organization, operating-system buffering, tensor-parallel partitioning, and GPU data movement jointly determine inference startup latency, and show that globally coordinated loading strategies expose greater storage concurrency and reduce redundant accesses compared to independently issued reads. Collectively, these contributions demonstrate that scalable I/O is best understood as a coordination and scheduling problem rather than a storage bandwidth problem, that effective aggregation strategies must adapt to workload-specific data structures rather than apply a single fixed policy, and that efficient data movement requires coordination across application frameworks, operating systems, and storage systems rather than treating each layer in isolation. These findings inform the design of future storage systems capable of supporting the continued growth of large-scale HPC and AI applications.

Author ORCID Identifier

https://orcid.org/0009-0002-1668-0250

Share

COinS
 
 

To view the content in your browser, please download Adobe Reader or, alternately,
you may Download the file to your hard drive.

NOTE: The latest versions of Adobe Reader do not support viewing PDF files within Firefox on Mac OS and if you are using a modern (Intel) Mac, there is no official plugin for viewing PDF files within the browser window.