AWS describes a RAG cost-cutting pattern where a small model filters retrieved chunks against the query before the main model answers, trimming input tokens.
Continue to AI University →