Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSOlder articles
Page 9 of 18-
A/B testing prompts in production with CPROMPT
Version a prompt, split traffic across versions with sticky per-user bucketing, render variables, and roll back, without a deploy or a feature-flag service.
-
Group-commit WAL: how it works and when to use it
Group-commit WAL, batches fsyncs on a timer instead of per write, for materially higher write throughput when you enable it. Here's how Crowkis does it and why it matters for cost and safety.
-
CSOURCE: answer lineage and cascade-purge when the source changes
Cached answers derive from sources, a doc, a config, an API. CSOURCE ties answers to their origin so that when the source changes, every answer built on it can be purged in one move.
-
Upgrades as non-events: the binary-swap contract
docker pull, restart, done, no schema migrations, no export/import, no upgrade runbook. The on-disk format is a stability promise, not an implementation detail.
-
How to use COUTCHECK: scan output for leaks before you send it
COUTCHECK scans a response for PII and toxicity, optionally validating JSON, and returns the entities found so you can redact, block, or regenerate.
-
i8 vector quantization: how it works and when to use it
i8 vector quantization, stores vectors in int8 for about 4x less memory, so more of your working set fits in RAM. Here's how Crowkis does it and why it matters for cost and safety.
-
Internal copilots: your whole company asks the same questions
HR policy, expense rules, deploy commands, VPN setup, every employee rediscovers them through your copilot, billed per discovery. Give the company one memory.
-
How to use CEVAL: grade output without a second model
CEVAL runs deterministic evaluators, toxicity, PII, relevance, JSON validity and more, over an input/output pair, and tracks the results on /metrics.
-
The 150-second stall we found in our own benchmark
CDEDUP works, and at 1,340 vectors it froze the whole server for 150 seconds in our harness. Here's the honest finding, why it happens, and what it means for how you should run dedup.
-
FinOps chargeback and savings receipts: how it works and when to use it
FinOps chargeback and savings receipts, meters spend across team, project, env, and cost center, and anchors savings receipts in an audit chain. Here's how Crowkis does it and why it matters for cost and safety.
-
CTOOLSET: cache the tool call so the agent stops paying for it twice
Agents call the same tools with the same arguments constantly. CTOOLSET and CTOOLGET cache tool results keyed by tool plus exact arguments, so a deterministic call runs once and serves many.
-
The supply-chain argument, made carefully
After the 2026 gateway compromise, 'how many packages are in your hot path?' became a real procurement question. Our answer is a number: zero.
-
RAG apps: cache the synthesis, not just the retrieval
Your vector store finds the chunks fast. Then the model re-synthesizes the same answer from the same chunks, thousands of times. That second step is the bill.
-
How to use CPROMPT: version and A/B test prompts
CPROMPT stores named prompt templates with versioning, variable rendering, sticky A/B splits, and rollback, all from the CLI, all surviving restart.
-
Model migration without a cold cache: how it works and when to use it
Model migration without a cold cache, canary and migration workflows carry cache value across a model upgrade, so a new model doesn't cold-start your hit rate. Here's how Crowkis does it and why it matters for cost and safety.
-
Three windows into one cache: dashboard, Prometheus, logs
The built-in dashboard for humans, /metrics for your Grafana, one JSON line per event for your pipeline, same truth, three consumers, zero adapters.
-
Why Crowkis refuses to meter you
A cache exists to make costs predictable. Metering the cache would be self-defeating. So Community is free and Enterprise is flat per cluster, priced on a call, not a meter.
-
How to use CBUDGET: read per-tenant spend and alerts
CBUDGET reports token and dollar consumption per tenant in real time, and surfaces the tenants approaching or crossing their thresholds.
-
Let Claude Code use the cache: Crowkis over MCP
One config block turns Crowkis into a tool an AI assistant can hold, check the cache, store the answer, over MCP, with the same trust pipeline as every other write.
-
How Crowkis earned the right to sit in your critical path
347 integration tests, a smoke suite that kills the process on purpose, and a Docker image hardened before anyone asked. The receipts behind 'production-ready.'
-
MCP server for AI apps: how it works and when to use it
MCP server for AI apps, doubles as an MCP server over stdio, so Claude and any MCP-capable agent can check the cache and store what they compute. Here's how Crowkis does it and why it matters for cost and safety.
-
CINVALIDATE: purge the cache by meaning, with a preview before you commit
Sometimes you need to clear 'everything about the old pricing', a fuzzy, semantic set. CINVALIDATE takes a natural-language instruction, previews what it would purge, and only acts on COMMIT.
-
HNSW without the network hop: why the vector index lives inside the engine
Most semantic caches call out to a vector database. Crowkis embeds the HNSW graph in-process, and that placement decision is worth more than any algorithm tweak.
-
Crowkis vs LiteLLM-style gateways: caching is not a checkbox
Python gateways treat caching as one feature among forty. Crowkis treats it as the product, and ships it without a Python supply chain attached.