Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSOlder articles
Page 10 of 18-
How to use CDEDUP: collapse near-duplicate answers
CDEDUP finds entries that mean the same thing and collapses them, reporting clusters and memory reclaimed, best run as off-peak maintenance.
-
How to cache LangChain LLM calls with Crowkis
Add a semantic cache to LangChain so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Tenant isolation as physics, not policy
A WHERE clause is a promise; a namespace is a wall. How Crowkis makes cross-tenant leakage structurally impossible rather than procedurally unlikely.
-
Agent fleets are token furnaces. Crowkis is the heat exchanger.
Agents re-ask, re-plan, and re-fetch with industrial enthusiasm. Multiply by a fleet and you get the most cacheable traffic in existence, if the cache understands agents.
-
How to use CINFO: the Crowkis-flavoured INFO
CINFO returns server, cache, savings, security, db, and license sections in one call, the fastest read on what the cache is doing right now.
-
Does the vector index go cold under churn? We tried to break it
A v0.2.1 bug let the HNSW index go cold under heavy write-and-flush churn, semantic search silently stopped finding neighbours. Here's the soak test that reproduces it and proves v0.2.2 fixed it.
-
Give LangChain agents long-term memory with Crowkis
Durable, per-user memory for LangChain agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
CSTALE: serve the slightly-old answer now, refresh it behind the scenes
A hard TTL turns a one-second-expired answer into a full model call. CSTALE serves the cached answer past its TTL with a stale flag, so you choose freshness versus latency per request.
-
Crowkis on Kubernetes: a well-behaved citizen
One container, a PVC, real health probes, hard memory bounds, graceful shutdown. Everything your cluster expects from a tenant that's read the manual.
-
Replay: the demo that uses your data instead of our slides
Every cache vendor promises a hit rate. Crowkis Replay computes yours, on your real queries, before you spend anything. The pitch is a number with your name on it.
-
How to cache LangGraph LLM calls with Crowkis
Add a semantic cache to LangGraph so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
The write-ahead log: how the cache survives a kill -9
Durability isn't a checkbox, it's a sequence of writes in the right order with checksums at every step. Here's the boring machinery that makes restarts uneventful.
-
Give LangGraph agents long-term memory with Crowkis
Durable, per-user memory for LangGraph agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
CBUDGET: per-tenant spend you can see before the invoice does
Token spend is usually a month-end surprise. CBUDGET tracks per-tenant token and dollar consumption in real time and surfaces alerts, so a runaway tenant is a notification, not a billing shock.
-
PII in a cache: scrub, isolate, erase, prove
Users put personal data in prompts whether you like it or not. The cache's job is a full lifecycle: keep it out of shared entries, find it on demand, erase it provably.
-
How to cache LlamaIndex LLM calls with Crowkis
Add a semantic cache to LlamaIndex so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Bloom filters: how the engine knows what it doesn't know
The fastest disk read is the one that never happens. A few bits per key let Crowkis skip files that can't contain your answer, at a 1% false-positive cost we chose on purpose.
-
AI coding assistants: the cache your team didn't know it was sharing
Every developer on your team asks the assistant the same questions about the same codebase. With Crowkis behind MCP, the second ask is free for everyone.
-
Give LlamaIndex agents long-term memory with Crowkis
Durable, per-user memory for LlamaIndex agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
Running a model canary: the operator's walkthrough
Slice the traffic, compare against cached baselines, promote or retreat, model upgrades as a controlled experiment with the cache as your measuring instrument.
-
Budgets with teeth: why your LLM spend needs a circuit breaker
Every team has a runaway-loop story that ends with a shocking invoice. Per-key budgets with hard TPM and dollar walls end the genre.
-
A million vectors on a laptop: the honest vector-search numbers
Crowkis is a cache with a vector index, not a vector database, but it should still hold up at scale. We indexed 100K and 1M vectors and measured build time, search latency, and recall. Including where dedicated vector DBs still win.
-
How to cache CrewAI LLM calls with Crowkis
Add a semantic cache to CrewAI so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
The AI Gateway: a semantic cache in front of any OpenAI-compatible API
Point your existing OpenAI client at Crowkis and change nothing else. The gateway proxies /v1/chat/completions, serves semantic hits without an upstream call, and adds retries, routing, and rate limits.