Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSOlder articles
Page 7 of 18-
Per-key budgets and circuit breakers (CBUDGET): how it works and when to use it
Per-key budgets and circuit breakers (CBUDGET), meters spend per tenant and can hard-block upstream calls past a circuit-breaker threshold, so a runaway loop hits a wall before the invoice. Here's how Crowkis does it and why it matters for cost and safety.
-
How to use CVECCOUNT: see how many vectors are live
CVECCOUNT returns the live entry count in the vector index, the quickest health signal for a semantic cache.
-
A drop-in CachedOpenAI for Node, in one wrapper
The Node SDK ships a typed client and a CachedOpenAI wrapper, keep your OpenAI calls exactly as they are, and a semantic cache slips in underneath.
-
Memcached walked so semantic caches could think
Classic caches taught the industry a durable lesson: never compute the same thing twice. LLMs just changed what 'the same thing' means, from identical bytes to identical meaning.
-
Per-tenant rate limits (CKEYLIMIT): how it works and when to use it
Per-tenant rate limits (CKEYLIMIT), caps requests and tokens per minute per tenant, enforced at the cache before the spend happens. Here's how Crowkis does it and why it matters for cost and safety.
-
CEVAL: nine evaluators that grade your LLM output without a second LLM
LLM-as-judge is expensive and leaks your data. CEVAL ships nine deterministic evaluators, toxicity, PII, injection-safety, relevance, JSON validity and more, that score input/output pairs locally and track the results over time.
-
How to use CFLUSH: clear the semantic cache, by tenant
CFLUSH empties the semantic cache, globally, or scoped to a single tenant so one customer's reset doesn't touch another's.
-
Golden answer pinning (CPIN): how it works and when to use it
Golden answer pinning (CPIN), serves a human-approved answer verbatim for any phrasing of a question, with an audit trail of who approved it. Here's how Crowkis does it and why it matters for cost and safety.
-
How to use CTHINK and CREUSE: bank a chain of thought
CTHINK stores a reasoning trace as a reusable step graph; CREUSE fetches the matching plan for a new query at a fraction of the token cost.
-
Giving an agent memory from Python: CMEM in practice
The memory commands from application code, store facts, recall them semantically, and watch consolidation retire the stale ones. A worked example in Python.
-
Negative / anti-hallucination cache (CFLAG): how it works and when to use it
Negative / anti-hallucination cache (CFLAG), records known-bad answers so every paraphrase of the question that would reproduce a hallucination is caught. Here's how Crowkis does it and why it matters for cost and safety.
-
How to use CSTALE: serve slightly-old, refresh behind it
CSTALE returns a cached answer even past its TTL, flagged as stale, so an expired entry is a snappy answer plus a refresh signal, not a cold miss.
-
Where the milliseconds go: an honest latency profile
A semantic cache hit isn't free, it has to embed your query first. We measured every operation's percentiles so you know exactly what you're paying for, and where the cache engine itself is microsecond-fast.
-
Answer lineage and cascade purge (CSOURCE): how it works and when to use it
Answer lineage and cascade purge (CSOURCE), ties answers to their source so that when a document changes, every answer built on it can be purged in one move. Here's how Crowkis does it and why it matters for cost and safety.
-
CPROMPT: version your prompts and A/B test them like code
Prompts are production logic edited like sticky notes. CPROMPT gives them named templates, automatic versioning, rollback, variable rendering, and sticky weighted A/B splits, all surviving restart.
-
How to use CINVALIDATE: purge by meaning, preview first
CINVALIDATE clears entries whose meaning matches a natural-language instruction, and previews exactly what it would remove until you add COMMIT.
-
Natural-language cache invalidation (CINVALIDATE): how it works and when to use it
Natural-language cache invalidation (CINVALIDATE), purges entries whose meaning matches a plain-English instruction, previewing by default and only acting on COMMIT. Here's how Crowkis does it and why it matters for cost and safety.
-
How to use CWHYEVICT: ask why an entry would be dropped
CWHYEVICT explains the retention maths for an entry, recency, frequency, isolation, and cost, so eviction is auditable, not mysterious.
-
Guardrails in your request path: CGUARD and COUTCHECK in code
Two commands wrap your model call in an input and an output gate, prompt-injection scanning before, PII and toxicity scanning after. No second model, no egress.
-
Stale-while-revalidate (CSTALE): how it works and when to use it
Stale-while-revalidate (CSTALE), returns a cached answer past its TTL with a stale flag, so expiry is a snappy answer plus a refresh signal, not a cold miss. Here's how Crowkis does it and why it matters for cost and safety.
-
CDOC: a self-hosted RAG store that chunks, filters, and reranks
You don't always need a separate vector database to do retrieval. CDOC is a mini RAG store inside Crowkis, auto-chunking, metadata filtering, and optional cross-encoder reranking, sharing the cache's embedder.
-
How to use CFLAG and CCHECKBAD: a memory for wrong answers
CFLAG records a known-bad answer in the negative cache; CCHECKBAD catches every paraphrase of the question that would reproduce it.
-
Semantic dedup (CDEDUP): how it works and when to use it
Semantic dedup (CDEDUP), folds near-duplicate answers into clusters and reports the memory reclaimed, best run as scheduled off-peak maintenance. Here's how Crowkis does it and why it matters for cost and safety.
-
How to use CPIN: serve a human-approved answer verbatim
CPIN pins a golden answer that's served word-for-word for matching questions, with an audit trail of who approved it.