Capping how many requests, tokens, or dollars a single user or account can consume from an API within a given time window, to control cost and prevent overload.
Continue to AI University →