One-line integration get_or_compute
Wrap your model call. A safe hit skips it; a miss runs it once.

Smarter caching, smaller bills.
A cache for AI apps. When users ask the same thing in different words, Crowkis answers from cache in under a millisecond. You pay your LLM once.
Users ask the same thing in different words. Your LLM bills every one as new.
Fifty asks. One question.Your model bills every one.
Why it keeps costing you.
Reworded means a new, full-price call.
Change one word and a key-value cache misses.
“Cancel my plan” is not “pause my plan”.
Cache one hallucination, serve it to everyone.
Switch models and most caches go cold.
Follow one refund question from your app to a cached answer.
Your app sends a question over RESP3, gRPC, REST or MCP.
Already answered
How long do refunds take?
Refunds take 5-7 business days.
New question
What’s the refund timeline?
overRESP3gRPCRESTMCP
Crowkis reads what the question means and how it is built, not just its words.
CachedHow long do refunds take?
NewWhat’s the refund timeline?
Five checks decide if a cached answer is safe to reuse.
Five checks, with example values
5 of 5 passed. Reuse the cached answer.
A safe match returns in about 0.4 ms. A miss calls your model once.
Served from cache in
0.4 ms
“Refunds take 5-7 business days.”
$0, no model call
On a miss: your LLM answers, Crowkis stores it after its trust checks.
Store an answer once. Ask it again in different words, and it still comes back from cache.
CSET "how do refunds work?"
"Refunds take 5-7 business days."
Stored onceCGET "what's the refund timeline?"
# a paraphrase, still a hitDifferent words · still a hitCSIM "how do refunds work?"
"refund process?"
# -> similarity scoreHow close two questions areYour app asks Crowkis first. A safe match comes back in about 0.4 ms. Anything else goes to your model once, and the answer is kept for next time.
The same fifty questions, all day.
One answer, reused by the whole team.
Cache the finished answer, not the chunks.
Memory per user, and reasoning reuse.
Replies start before the caller notices.
Each caller gets their own details. Your model never hears the repeat.
Crowkis keeps the shape of the answer. The caller’s details are never stored.
Same shape, this caller’s details. No model call.
On a hit, the model’s respond event is skipped. Works with any realtime voice API.
No inference billedNo synthesis billed
Nothing caller‑specific ever enters the cache.
If an answer can’t be de‑personalised, it isn’t cached.
Measured in our tests on the live voice path, single-threaded.
It matches what a question means, then checks the answer is safe to reuse.
1model call per question
Pay for an answer once, however it is asked.
0.4 msper cache hit
Under a millisecond, not a multi-second model call.
5gates, every hit, every time
Five checks before any answer is reused.
1Docker image, every feature in
docker pull crowkis/crowkis:latest
Your existing clients connect unmodified.
FreeCommunity edition, self-hosted
no licence, no sign-up+ SSO · audit log · budgets
Self-hosted. Community is free.
get_or_computeWrap your model call. A safe hit skips it; a miss runs it once.
Hits stream back like live model output.
An image and its text match as one entry.
Reuses the reasoning, not just the words.
Each intent tunes its own bar. Riskier means stricter.
MCPClaude Code checks the cache before spending tokens.
Every hit, miss and block, live. Prometheus and OpenTelemetry built in.
Canary a new model, then migrate. Your cache stays warm.
It decides: reuse or recompute.
Keep your RAG store. Crowkis fronts the model.
Runs fully air-gapped.
A miss passes straight through.
One Docker image for the server, one package for your app.
One container, three local ports.
docker pull crowkis/crowkis:latestdocker run -d --name crowkis \
-p 127.0.0.1:6379:6379 \
-p 127.0.0.1:6380:6380 \
-p 127.0.0.1:6381:6381 \
-v crowkis-data:/data \
crowkis/crowkis:latestcurl 127.0.0.1:6380/healthWrap your model call. Rephrasings hit the cache.
pip install crowkisfrom crowkis import Crowkis
cache = Crowkis(host="127.0.0.1", port=6379, tenant="my-app")
@cache.cached(ttl=3600)
def answer(prompt: str) -> str:
return my_model(prompt)
answer("How do refunds work?") # miss → model runs, cached
answer("What's the refund process?") # semantic HIT → served from cacheOptional: connect Claude Code or any agent over MCP.
claude mcp add crowkis -- crowkis mcp