Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSOlder articles
Page 8 of 18-
The throughput ceiling we won't hide, and the fix
On v0.2.2, throwing 16 threads at Crowkis got the same throughput as one. That's a real ceiling, we found it in our own harness, and here's both why it happened and how the embedding-deferral work lifts it.
-
PII scrubbing and right-to-erasure (CPII): how it works and when to use it
PII scrubbing and right-to-erasure (CPII), reports what personal data is cached and executes right-to-erasure on request, so compliance is a command. Here's how Crowkis does it and why it matters for cost and safety.
-
CSESSION: conversation buffers with semantic recall built in
Chat history is more than the last N turns. CSESSION stores a multi-turn buffer per session, bounded and TTL'd, with both recent-window reads and semantic search across the whole conversation.
-
How to use CSOURCE: tie answers to their source and cascade-purge
CSOURCE links cache entries to the source they derived from, so when the source changes you can purge everything built on it in one move.
-
Multi-tenant isolation: how it works and when to use it
Multi-tenant isolation, namespaces keys per tenant and tags every entry, so one customer's answer can never be served to another. Here's how Crowkis does it and why it matters for cost and safety.
-
How to use CTOOLSET and CTOOLGET: cache a tool call
CTOOLSET caches a tool result keyed by tool plus exact arguments; CTOOLGET returns it, so a deterministic call runs once and serves many.
-
Self-hosted RAG in twenty lines with CDOC
Add documents, auto-chunk them, search with metadata filters and reranking, a working retrieval pipeline without a separate vector database.
-
Observability dashboard + Prometheus: how it works and when to use it
Observability dashboard + Prometheus, shows hit rate, saved spend, safety blocks, and memory pressure live, and exposes Prometheus /metrics, all in the box. Here's how Crowkis does it and why it matters for cost and safety.
-
CPIN: golden answers that are served verbatim, with an audit trail
For the questions where 'close enough' is unacceptable, pricing, legal, brand lines, CPIN serves a human-approved answer verbatim, records who approved it, and never lets the model improvise.
-
Why Crowkis is Rust all the way down
A cache lives in the hot path of every request. The language choice isn't aesthetic, it's the difference between predictable microseconds and mystery pauses.
-
How to use CDOC: a RAG store in the CLI
CDOC adds documents with auto-chunking and metadata, then searches them with filters and optional reranking, no separate vector database.
-
Redis-compatible (RESP3): how it works and when to use it
Redis-compatible (RESP3), speaks RESP3 so redis-py, ioredis, and Lettuce connect unmodified across 40+ commands, adoption is a port change. Here's how Crowkis does it and why it matters for cost and safety.
-
The five-minute deploy, timed honestly
Pull, run, first hit in the dashboard, with no config file, no signup, and no environment variables you're required to set. We timed it. It holds.
-
Crowkis vs Redis: same protocol, different century
Redis is magnificent infrastructure for exact-match workloads. LLM traffic isn't one. Here's why speaking the same protocol doesn't mean solving the same problem.
-
How to use CSESSION: a conversation buffer with recall
CSESSION stores a multi-turn conversation, reads the recent window, and semantically searches the whole thing, so 'as I mentioned earlier' works.
-
84 of 84: correctness and isolation under a hostile harness
Empty strings, 100 KB values, null bytes, emoji, 16 threads hammering across tenants. The stress harness throws 84 nasty checks at Crowkis and counts the cross-tenant leaks. The leak count is zero.
-
Pure-Rust LSM storage engine: how it works and when to use it
Pure-Rust LSM storage engine, is a from-scratch WAL + MemTable + SSTable + bloom-filter + compaction engine in Rust, no RocksDB and no FFI. Here's how Crowkis does it and why it matters for cost and safety.
-
CFLAG and CCHECKBAD: a memory for the answers that were wrong
Most caches only remember good answers. Crowkis also remembers bad ones, flag a hallucinated or harmful response once, and CCHECKBAD catches every paraphrase of the question that would have reproduced it.
-
Support bots are the single best caching workload in software
Nowhere else do thousands of people ask the same fifty questions, all day, in every phrasing imaginable. Crowkis was practically designed in a support queue.
-
How to use the CMEM commands: long-term agent memory
CMEMSET stores a fact scoped to (agent, user); CMEMGET recalls by meaning, recency-blended; consolidation retires contradictions automatically.
-
In-process HNSW vector index: how it works and when to use it
In-process HNSW vector index, keeps a custom, persistent HNSW graph in the same process as the store, so a lookup and its scoring never cross a network hop, sub-millisecond search. Here's how Crowkis does it and why it matters for cost and safety.
-
Prompt injection meets your cache: the attack nobody threat-modeled
Injected instructions in one response become served truth for every similar query, unless the cache can smell an answer that doesn't answer.
-
The token math of repetition: what your duplicate questions actually cost
Take your daily query volume, multiply by the repeat fraction, multiply by your blended price per call. That number, twelve times a year, is the cache argument.
-
How to use CGUARD: scan input for prompt injection
CGUARD checks a prompt for jailbreaks and injections, normalizing leetspeak and zero-width tricks first, and returns a verdict, category, and match.