Semantic Caching for LLMs: How It Works and the Tools That Do It
TL;DR
* Semantic caching for LLMs serves a stored response when a new request's embedding falls within a similarity threshold of a cached one, so paraphrased repeats skip the model call.
* Bifrost runs an exact hash lookup first and a semantic lookup only on a miss, so identical