Apple Proposes a 'Length Value Model' to Predict How Long an LLM's Answer Will Be, Token by Token

Apple researchers introduce LenVM, a model that estimates remaining response length at each decoding step to help manage LLM inference cost.