AWS Shows a Query-Aware Compression Pattern to Cut RAG Costs on Bedrock

AWS describes a RAG cost-cutting pattern where a small model filters retrieved chunks against the query before the main model answers, trimming input tokens.