Splitting a large computation into blocks small enough to fit in a GPU's fast on-chip memory, cutting slow trips to main memory.
Continue to AI University →