How to cache Semantic Kernel LLM calls with Crowkis
Add a semantic cache to Semantic Kernel so repeated and reworded questions are served for free, no rewrite, self-hosted.
Your users ask the same things all day, phrased a hundred different ways. If you build with Semantic Kernel, most of that repetition is invisible in your code but very visible on your bill. A semantic cache in front of your model calls fixes it.
The lowest-friction path is the OpenAI-compatible gateway: point Semantic Kernel's base URL at Crowkis and every model call flows through a semantic cache. Repeated and reworded prompts are served from cache with no upstream call; new ones pass through and get cached.
# point Semantic Kernel at the Crowkis gateway
base_url = "http://127.0.0.1:6380/v1" # semantic cache in front of your providerIn plain words. You don't restructure your Semantic Kernel app. You change where the calls go, and repeats stop costing money.
On repetitive workloads this cuts LLM costs up to 60-70% on repetitive workloads, and every hit comes back with a confidence score so reuse stays safe. Runs self-hosted with zero egress, nothing leaves your machine.