Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSOlder articles
Page 5 of 18-
Bi-temporal (time-travel) memory: how it works and when to use it
Bi-temporal (time-travel) memory, answers what the agent believed at a past instant using validity windows, so you can reconstruct 'what did we know then?'. Here's how Crowkis does it and why it matters for cost and safety.
-
Restart-safe by design: a cache that survives a crash without losing a thing
A cache that forgets everything on restart isn't much of a memory. Crowkis writes durably first, so a crash, a deploy, or a reboot never costs you a single learned answer.
-
Why Rust for AI infrastructure
Why Rust for AI infrastructure. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Auto fact extraction from conversations: how it works and when to use it
Auto fact extraction from conversations, pulls durable facts out of a transcript deterministically, dropping questions, greetings, and filler, with no model call. Here's how Crowkis does it and why it matters for cost and safety.
-
The five gates every write passes before it can poison your cache
In any shared cache, one crafted answer could get served to thousands. So every write runs a five-stage gauntlet before it's ever eligible to be reused.
-
Sub-millisecond retrieval: why in-process beats a network hop
Sub-millisecond retrieval: why in-process beats a network hop. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Input guardrails (CGUARD): how it works and when to use it
Input guardrails (CGUARD), scans prompts for injection and jailbreaks after normalizing leetspeak, whitespace, and zero-width evasion, model-free and stateless. Here's how Crowkis does it and why it matters for cost and safety.
-
Why 'who wrote Hamlet' and 'today's stock price' can't share a TTL
A single time-to-live for every cached answer is a bug in disguise. Some facts are true for a decade; some are stale in minutes. Freshness has to know the difference.
-
FinOps for LLMs: attributing AI spend
FinOps for LLMs: attributing AI spend. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Output guardrails (COUTCHECK): how it works and when to use it
Output guardrails (COUTCHECK), scans responses for PII, toxicity, and JSON validity before they ship, so the model's output is checked at the trust boundary. Here's how Crowkis does it and why it matters for cost and safety.
-
A cache hit is not the same as a correct answer
Most caches treat every hit as equally trustworthy, a binary yes. But LLM answers are probabilistic and time-sensitive. Crowkis scores its confidence before it serves.
-
Evals without an LLM judge
Evals without an LLM judge. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Model-free online evals (CEVAL): how it works and when to use it
Model-free online evals (CEVAL), grades output with deterministic evaluators (toxicity, PII, relevance, JSON validity and more) and tracks the results over time, no LLM-judge. Here's how Crowkis does it and why it matters for cost and safety.
-
LRU throws away your most expensive knowledge
Least-recently-used eviction is semantically blind. It happily discards the rare, costly answer you'll pay dearly to regenerate, and keeps the cheap trivia everyone asks. There's a smarter way.
-
Prompt versioning and A/B testing at the data layer
Prompt versioning and A/B testing at the data layer. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Prompt versioning and A/B testing (CPROMPT): how it works and when to use it
Prompt versioning and A/B testing (CPROMPT), versions prompts on every write, renders variables, and runs sticky per-user A/B splits, so you roll back or test without a code deploy. Here's how Crowkis does it and why it matters for cost and safety.
-
Reasoning reuse: stop paying for the same chain of thought twice
Chain-of-thought costs several times more tokens than the answer it produces, and it's usually thrown away. For structurally similar problems, most of that reasoning can be reused.
-
Streaming responses from cache without breaking UX
Streaming responses from cache without breaking UX. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Self-hosted RAG (CDOC): how it works and when to use it
Self-hosted RAG (CDOC), adds documents with auto-chunking and metadata, then runs filtered ANN search with optional reranking, no separate vector database. Here's how Crowkis does it and why it matters for cost and safety.
-
One similarity cutoff is always wrong
'What's 2+2?' needs a near-exact match to reuse safely. 'Give me creative ideas for X' can tolerate a loose one. A single global threshold guarantees you're too strict somewhere and too loose somewhere else.
-
The case for a Redis-compatible AI cache
The case for a Redis-compatible AI cache. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Local offline embeddings (CEMBED): how it works and when to use it
Local offline embeddings (CEMBED), turns text into vectors with the bundled local model, no API key and no egress, with a micro-cache so repeats are free. Here's how Crowkis does it and why it matters for cost and safety.
-
Tenant isolation is a feature you test, not a checkbox you claim
Every multi-tenant product says its tenants are isolated. The ones you can trust are the ones that try to break it on purpose. Here's how we prove no tenant can read another's data.
-
Cache poisoning is the whole problem with semantic caches
Cache poisoning is the whole problem with semantic caches. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.