Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSOlder articles
Page 13 of 18-
The CFO pitch: explaining the cache to the person who signs things
Three sentences, one dashboard number, and a flat price. The rare infrastructure purchase that finance understands faster than engineering does.
-
Give the Vercel AI SDK agents long-term memory with Crowkis
Durable, per-user memory for the Vercel AI SDK agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
Closed-source as a security posture, argued honestly
'Many eyes' assumes the eyes show up. For your hot path, a signed single binary with zero dependencies is a smaller attack surface than a thousand auditable packages nobody audits.
-
Crowkis vs ElastiCache: managed Redis is still Redis
AWS will happily run an exact-match cache for you at any scale. It will miss your LLM traffic at any scale, too.
-
How to cache LiteLLM LLM calls with Crowkis
Add a semantic cache to LiteLLM so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Why we kept the Redis protocol instead of inventing an API
Every new API is a tax on adoption: clients, docs, muscle memory, tooling. RESP3 meant inheriting twenty years of all four on day one.
-
Multi-tenant SaaS: one cache, many customers, zero leaks
Caching across customers multiplies savings and multiplies risk. Tenant isolation has to be architecture, not a WHERE clause.
-
Give LiteLLM agents long-term memory with Crowkis
Durable, per-user memory for LiteLLM agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
CPII: scrubbing personal data and honouring the right to be forgotten
A cache of LLM traffic is a cache of whatever users typed, including PII. CPII reports what personal data is present and executes right-to-erasure, so compliance is a command, not a project.
-
Provider arbitrage: paying frontier prices only for frontier questions
Model prices vary 50x for overlapping quality on easy queries. The arbitrage router exploits the spread automatically, with a quality bar you set per intent.
-
How to cache Ollama LLM calls with Crowkis
Add a semantic cache to Ollama so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Crowkis vs Memcached: a beautiful fossil meets a new workload
Memcached is the purest cache ever written, and purity is exactly the problem when your keys are sentences.
-
Give Ollama agents long-term memory with Crowkis
Durable, per-user memory for Ollama agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
One actor, no locks across await: the concurrency design
Crowkis serves thousands of connections through async IO, then funnels every cache decision through a single deterministic actor. Here's why that's a feature.
-
Startups: your LLM bill is eating runway you'll want back
Seed-stage AI products routinely spend salary-sized sums recomputing known answers. Free Community edition exists precisely for this moment of your company.
-
How to cache the OpenAI Python SDK LLM calls with Crowkis
Add a semantic cache to the OpenAI Python SDK so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
CINFO and the dashboard: a cache you can actually watch work
Infrastructure you can't observe is infrastructure you don't trust. CINFO and the built-in dashboard expose hit rate, saved spend, safety blocks, memory pressure, and license state in real time.
-
The hidden invoice of a cold cache: what model migrations really cost
Swap models with a normal cache and you re-purchase your entire corpus at the new model's prices. Migration leasing is the line item that prevents the line item.
-
Give the OpenAI Python SDK agents long-term memory with Crowkis
Durable, per-user memory for the OpenAI Python SDK agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
Crowkis vs Dragonfly, Valkey, and KeyDB: faster exact-matching is still exact-matching
The new Redis-compatibles race each other on throughput. On LLM traffic they all hit the same wall at full speed: the keys never repeat.
-
How to cache the OpenAI Node SDK LLM calls with Crowkis
Add a semantic cache to the OpenAI Node SDK so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Designing the MCP server: a cache as a tool the model can hold
MCP turns Crowkis into something an AI assistant can use deliberately, check the cache, store the answer, over plain stdio, with the banner silenced so JSON-RPC stays clean.
-
Platform teams: make caching a paved road, not a per-team adventure
Every product team is duct-taping its own LLM cache right now. Platform engineering exists to end exactly this kind of duplication.
-
Give the OpenAI Node SDK agents long-term memory with Crowkis
Durable, per-user memory for the OpenAI Node SDK agents that survives restarts and consolidates contradictions, self-hosted, zero egress.