Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSOlder articles
Page 6 of 18-
Bring your own embedder / reranker: how it works and when to use it
Bring your own embedder / reranker, swaps in any sentence-transformers MiniLM or GTE export via ONNX, so you're never locked to the default model. Here's how Crowkis does it and why it matters for cost and safety.
-
One container, zero dependencies: what's deliberately absent from our image
The most secure dependency is the one that isn't there. Crowkis ships as a single stripped binary with the model baked in, no Python, no package manager, nothing to poison at runtime.
-
How to warm an LLM cache during a model migration
How to warm an LLM cache during a model migration. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Multi-turn session memory (CSESSION): how it works and when to use it
Multi-turn session memory (CSESSION), keeps a bounded conversation buffer with both recent-window reads and semantic search across the whole chat. Here's how Crowkis does it and why it matters for cost and safety.
-
The agentic era needs a memory layer. Here's what it looks like.
We gave agents tools, planning, and the ability to act. We forgot to give them a place to remember. That gap is why your agents feel brilliant and amnesiac at the same time.
-
Reasoning reuse: the deepest LLM saving nobody talks about
Reasoning reuse: the deepest LLM saving nobody talks about. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Tool-result caching (CTOOLSET): how it works and when to use it
Tool-result caching (CTOOLSET), caches a deterministic tool call keyed by tool plus exact args, so a swarm's duplicate lookups become one call. Here's how Crowkis does it and why it matters for cost and safety.
-
How good is Crowkis agent memory, really? The LoCoMo and LongMemEval numbers
We ran Crowkis memory against two public, hostile retrieval benchmarks, SNAP's LoCoMo and LongMemEval, on a laptop with no cloud calls. Here are the recall numbers, by question type, with the reranker on and off.
-
Five agents asking one question should cost one answer
Multi-agent systems fan out, and they ask overlapping questions constantly. Without a shared cache, that overlap is pure waste, the same answer, bought once per agent.
-
Multimodal caching (image + text): how it works and when to use it
Multimodal caching (image + text), caches image-plus-text lookups, so a repeated vision question is a hit instead of an expensive re-run. Here's how Crowkis does it and why it matters for cost and safety.
-
The crowkis CLI: every subcommand, with the flags that matter
The binary is the whole product, server, REPL, doctor, bench, and the inspect tools. A tour of the crowkis command line, from cold start to debugging a missed hit.
-
What is an embedding, really? A plain-English guide
Embeddings sound like math you need a PhD for. The core idea is simpler and more useful than that, and it's the reason a cache can tell that two different sentences mean the same thing.
-
Streaming response caching: how it works and when to use it
Streaming response caching, serves cached answers chunk by chunk, so a hit feels like live typing and the seam between hit and miss disappears. Here's how Crowkis does it and why it matters for cost and safety.
-
CGUARD: an input guardrail that survives leetspeak and zero-width tricks
Prompt injection rarely arrives in plain English. CGUARD normalizes the evasion first, whitespace, leetspeak, zero-width characters, then scans for jailbreaks, overrides, and system-prompt exfiltration.
-
How to use CSET: store an answer the safe way
CSET writes an answer into the semantic cache, running the five-stage anti-poisoning pipeline before it accepts anything.
-
HNSW explained: finding the needle in a million haystacks
Once meaning is a point in space, the hard part is finding the nearest point out of millions, fast. HNSW is the elegant trick that makes it feel instant, here's the intuition.
-
OpenAI-compatible AI gateway: how it works and when to use it
OpenAI-compatible AI gateway, proxies /v1/chat/completions with a semantic cache in front, so you point your client's base URL at Crowkis and change nothing else. Here's how Crowkis does it and why it matters for cost and safety.
-
How to use CGET: a lookup that matches meaning
CGET finds a cached answer by meaning, not exact bytes, and can return the confidence behind the hit so you decide whether to trust it.
-
Cache an LLM call in three lines: the Python SDK
The Python SDK wraps the semantic cache in an ergonomic client, get-or-compute, streaming, tenants, models. Here's the three-line version and the production version.
-
The hidden cost of RAG nobody puts on the slide
Retrieval-augmented generation stuffs context into every prompt to make answers better. It also makes every repeated question dramatically more expensive, and repeated questions are most of them.
-
Multi-provider routing and fallback: how it works and when to use it
Multi-provider routing and fallback, load-balances and fails over across providers on error class, with retries using exponential backoff and jitter. Here's how Crowkis does it and why it matters for cost and safety.
-
COUTCHECK: catching the PII leak and the toxic line before your user does
The model's output is the other trust boundary. COUTCHECK scans responses for PII leakage and toxicity, and optionally validates JSON, returning a structured verdict you can act on.
-
How to use CSIM: score how close two queries are
CSIM returns the semantic similarity between two strings, the primitive behind every hit decision, exposed so you can calibrate thresholds.
-
Give your coding agent a memory
Coding agents re-read the same schema, re-derive the same conventions, and re-ask the same architecture questions on every run. A shared memory turns that repeated context into a one-time cost.