Computer Science > Hardware Architecture Title:Fleet: Hierarchical Task-based Abstraction for Megakernels on Multi-Die GPUs View PDF HTML (experimental)Abstract:Modern GPUs adopt chiplet-based designs with multiple private cache hierarchies, but current programming models (CUDA/HIP) expose a flat execution hierarchy that cannot express chiplet-level locality or synchronization. This mismatch leads to redundant memory traffic and poor cache utilization in memory-bound workloads such as LLM...

Fleet: Hierarchical Task-Based Abstraction for Megakernels on Multi-Die GPUs
Chowdhary; Sangeeta; Swann; Ryan; Siddens; Sean; Osama; Muhammad; Neuendorffer; Stephen; Dutu; Alexandru; Sangaiah; Karthik; Bhuyan; Sandeepa; Bayliss; Samuel; Dasika; Ganesh

