Date of Award
8-2026
Document Type
Dissertation
Degree Name
Doctor of Philosophy (PhD)
Department
Computer Engineering
Committee Chair/Advisor
Jon C. Calhoun
Committee Member
Melissa C. Smith
Committee Member
Rong Ge
Committee Member
Tao Wei
Committee Member
Bogdan Nicolae
Abstract
Modern high-performance computing (HPC) and artificial intelligence (AI) workloads increasingly generate data at rates that outpace the storage systems responsible for persisting it. While computational throughput has scaled dramatically through parallel processors, GPUs, and specialized accelerators, storage systems have improved far more slowly. As such, application performance is increasingly limited by the ability to move data efficiently through deep, heterogeneous storage hierarchies rather than by raw computational capability. This dissertation investigates how aggregation and data movement strategies must be redesigned to close this gap, addressing four central challenges: resource contention between foreground computation and background I/O, the fine-grained and irregular access patterns of modern workloads, the scalability limits of coordinating I/O across many concurrent processes, and the difficulty of moving data efficiently across increasingly heterogeneous memory and storage tiers. We address these challenges across three representative domains: large-scale HPC checkpointing, LLM training, and LLM inference. We first show that collective I/O techniques designed for synchronous checkpointing are poorly suited to asynchronous execution due to their reliance on globally coordinated synchronization. Through an extensive characterization of aggregation strategies, spanning thread concurrency, buffer management, data contiguity, and file organization, we design a hierarchical aggregation strategy for VeloC that partitions processes into independent I/O groups, avoiding global synchronization while achieving up to 2$\times$ higher checkpoint throughput than existing systems such as GenericIO and ADIOS2. We then extend these ideas to LLM training, where checkpoints are fragmented into hundreds or thousands of small, heterogeneous shards dictated by tensor, pipeline, and data parallelism. Characterizing how these workloads interact with modern kernel-level I/O mechanisms, including io\_uring, buffered I/O, and direct I/O, we uncover a counterintuitive result: checkpoint restoration can take nearly twice as long as checkpoint generation. Profiling production checkpointing frameworks traces this overhead not to storage bandwidth but to software inefficiencies, including per-tensor synchronization and repeated dynamic memory allocation. Guided by this analysis, we redesign the DataStates-LLM checkpoint engine and DeepSpeed's restore pipeline around batched, asynchronous I/O submission and reusable buffer pools, reducing end-to-end restoration time by more than a factor of two. Finally, we examine checkpoint loading for LLM inference, where model initialization has become an increasingly significant systems bottleneck. We characterize how checkpoint layout, filesystem organization, operating-system buffering, tensor-parallel partitioning, and GPU data movement jointly determine inference startup latency, and show that globally coordinated loading strategies expose greater storage concurrency and reduce redundant accesses compared to independently issued reads. Collectively, these contributions demonstrate that scalable I/O is best understood as a coordination and scheduling problem rather than a storage bandwidth problem, that effective aggregation strategies must adapt to workload-specific data structures rather than apply a single fixed policy, and that efficient data movement requires coordination across application frameworks, operating systems, and storage systems rather than treating each layer in isolation. These findings inform the design of future storage systems capable of supporting the continued growth of large-scale HPC and AI applications.
Recommended Citation
Gossman, Mikaila J., "Improving Collective Aggregation Strategies for HPC and AI Workloads" (2026). All Dissertations. 4298.
https://open.clemson.edu/all_dissertations/4298
Author ORCID Identifier
https://orcid.org/0009-0002-1668-0250