One scheduled unit of GPU work; a model typically runs as a sequence of them, each paying its own startup and memory-traffic cost.
Continue to AI University →