Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSOlder articles
Page 14 of 18-
CKEYLIMIT: per-tenant rate limits that stop the runaway before it starts
A runaway agent or a noisy tenant can torch a budget in minutes. CKEYLIMIT sets per-tenant requests-per-minute and tokens-per-minute ceilings, enforced locally before the spend happens.
-
The ROI timeline: hour one, week one, quarter one
Caching ROI isn't a hockey stick, it's a staircase that starts the first hour. Here's the honest schedule of when each saving shows up.
-
How to cache the Anthropic SDK LLM calls with Crowkis
Add a semantic cache to the Anthropic SDK so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Crowkis vs OpenAI prompt caching: a discount is not a cache
Provider prompt caching discounts your repeated prefixes. You still call the model, still wait, and still pay, just slightly less. There's a bigger idea available.
-
Give the Anthropic SDK agents long-term memory with Crowkis
Durable, per-user memory for the Anthropic SDK agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
Three levels, one strategy: compaction without the tuning PhD
LSM compaction is where storage engines breed complexity. Crowkis ships exactly one strategy across three levels, chosen for cache workloads, closed for configuration.
-
Consumer chat at scale: when every millisecond and every token multiply
At consumer scale, traffic converges on shared intents while costs and latency multiply by millions. The cache becomes load-bearing infrastructure.
-
Cache poisoning is the whole problem
Semantic caching has an obvious failure mode nobody likes to talk about: one bad write, served forever to everyone nearby. This is how Crowkis decides what to trust.
-
How to cache the Gemini SDK LLM calls with Crowkis
Add a semantic cache to the Gemini SDK so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
CTHINK and CREUSE: banking a chain of thought and replaying it
The reasoning is the expensive part of a hard answer. CTHINK stores a chain-of-thought trace as a reusable step graph; CREUSE fetches the matching plan for a new query at a fraction of the original token cost.
-
Give the Gemini SDK agents long-term memory with Crowkis
Durable, per-user memory for the Gemini SDK agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
Crowkis vs Anthropic prompt caching: cache writes that bill you are telling you something
Anthropic's prompt caching is excellent at its actual job, cheap long contexts. It was never designed to be your response cache, and the pricing says so.
-
How to cache the Mistral SDK LLM calls with Crowkis
Add a semantic cache to the Mistral SDK so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Streaming cache hits: instant answers that still feel like typing
Users expect LLM answers to arrive as a typing stream. CGETSTREAM serves cached answers chunk by chunk, so a sub-millisecond hit doesn't break the interface's rhythm.
-
Voice assistants: caching as a conversational necessity
Voice gives you about a second before silence feels broken. Model round-trips don't fit. Cache hits do, with room to spare for the speech stack.
-
Give the Mistral SDK agents long-term memory with Crowkis
Durable, per-user memory for the Mistral SDK agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
How to cache LangChain.js LLM calls with Crowkis
Add a semantic cache to LangChain.js so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Crowkis vs Gemini context caching: renting memory by the hour
Google bills cached context per token per hour, a parking meter for your own prompts. Compare that with a cache you simply own.
-
Give LangChain.js agents long-term memory with Crowkis
Durable, per-user memory for LangChain.js agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
347 tests and a murder weapon: how the suite is organized
Bottom-heavy by design: the layers that hold your data get the most hostile coverage, and the smoke suite's signature move is killing the process to prove a point.
-
Translation pipelines: the same strings, the same languages, every release
Product copy, help docs, and templates get re-translated continuously as releases churn. Most of the content didn't change. Stop paying as if it did.
-
How to cache Spring AI LLM calls with Crowkis
Add a semantic cache to Spring AI so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Give Spring AI agents long-term memory with Crowkis
Durable, per-user memory for Spring AI agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
Crowkis vs vLLM prefix caching: different layers, different physics
vLLM's prefix caching saves GPU work inside one inference server. Crowkis saves the inference itself. You probably want both, but only one cuts the bill to zero on a hit.