Crowkis

Crowkis logo: a white geometric crow on red

Smarter caching, smaller bills.

A cache for AI apps. When users ask the same thing in different words, Crowkis answers from cache in under a millisecond. You pay your LLM once.

per cache hit
0.4 ms
safety checks per hit
5
protocols: RESP3, gRPC, REST, MCP
4
Community edition
Free

Your LLM bill is mostly déjà vu.

Users ask the same thing in different words. Your LLM bills every one as new.

  • 09:02 · new callWhen do you close?
  • 09:47 · new callWhat time do you shut?
  • 10:15 · new callclosing time?
  • 10:58 · new callAre you open late tonight?
  • 11:40 · new callTill what time are you open?
  • 12:26 · new callhours today?
  • 13:08 · new callHow late are you open?
  • 13:55 · new callWhat’s your closing time?
  • 14:31 · new callu open till when
  • 15:12 · new callWhen are you closing tonight?
  • 16:04 · new callStill open at 8?
  • 16:48 · new callwhat time do you close today
  • 17:20 · new callWhen do you shut tonight?
  • 18:02 · new callopen late?

Fifty asks. One question.Your model bills every one.

Model calls
×50
Cache hits
0

Why it keeps costing you.

  1. You pay for reruns

    Reworded means a new, full-price call.

  2. Exact caches miss

    Change one word and a key-value cache misses.

  3. Similarity alone is unsafe

    “Cancel my plan” is not “pause my plan”.

  4. Bad answers spread

    Cache one hallucination, serve it to everyone.

  5. Upgrades reset you

    Switch models and most caches go cold.

How Crowkis answers a repeat question

Follow one refund question from your app to a cached answer.

  1. Ask

    Your app sends a question over RESP3, gRPC, REST or MCP.

    Already answered

    How long do refunds take?

    Refunds take 5-7 business days.

    New question

    What’s the refund timeline?

    overRESP3gRPCRESTMCP

  2. Understand

    Crowkis reads what the question means and how it is built, not just its words.

    CachedHow long do refunds take?

    NewWhat’s the refund timeline?

    Intent
    factual
    Template
    refund · {timeline}
  3. Check

    Five checks decide if a cached answer is safe to reuse.

    Five checks, with example values

    1. similarity0.94
    2. templaterefund · {timeline}
    3. confidence0.91 ≥ 0.88
    4. trust
    5. freshness

    5 of 5 passed. Reuse the cached answer.

  4. Answer

    A safe match returns in about 0.4 ms. A miss calls your model once.

    Served from cache in

    0.4 ms

    “Refunds take 5-7 business days.”

    $0, no model call

    On a miss: your LLM answers, Crowkis stores it after its trust checks.

Three commands. That’s the idea.

Store an answer once. Ask it again in different words, and it still comes back from cache.

three commands, any RESP3 client
  1. CSET "how do refunds work?" "Refunds take 5-7 business days."Stored once
  2. CGET "what's the refund timeline?" # a paraphrase, still a hitDifferent words · still a hit
  3. CSIM "how do refunds work?" "refund process?" # -> similarity scoreHow close two questions are

How Crowkis fits in your stack.

Your app asks Crowkis first. A safe match comes back in about 0.4 ms. Anything else goes to your model once, and the answer is kept for next time.

How Crowkis fits in your stack Your app asks Crowkis over RESP3, gRPC, REST or MCP. Crowkis reads the meaning and structure of the question and searches its cached answers. The nearest answer must pass five checks: similarity, template, confidence, trust and freshness. A hit goes back to your app in about 0.4 ms. On a miss your app calls your LLM once and stores the answer, which passes the write checks before Crowkis keeps it. Your app RESP3, gRPC, REST or MCP Your LLM called only on a miss Understand meaning + structure Cached answers vector index, per tenant Five checks similarity, template, confidence, trust, freshness Write checks anti-poisoning, PII scrub ask search nearest hit: answer in about 0.4 ms on a miss store answer kept Your servers: Crowkis, built in Rust

For apps that hear the same questions.

  • Support bots

    The same fifty questions, all day.

  • Copilots & dev tools

    One answer, reused by the whole team.

  • RAG & docs search

    Cache the finished answer, not the chunks.

  • AI agents

    Memory per user, and reasoning reuse.

  • Voice assistants

    Replies start before the caller notices.

Voice agents that answer instantly.

Each caller gets their own details. Your model never hears the repeat.

  1. 01

    The first caller asks. Your model answers once.

    Crowkis keeps the shape of the answer. The caller’s details are never stored.

    Caller 1“Where’s my order?”
    Your model
    Answer“Your order #1182 arrives Tuesday.”
    Crowkis keeps the shape Your order {order_id} arrives {date}.
    #1182Tuesdaynot stored
  2. 02

    The next caller gets their own answer.

    Same shape, this caller’s details. No model call.

    Caller 2“When will my order get here?”
    Your modelnever called
    Same shape · Caller 2’s own values Your order #4471 arrives Thursday.
    0.34 ms
    Agent says“Your order #4471 arrives Thursday.”
  3. 03

    The reply lands before the model would speak.

    On a hit, the model’s respond event is skipped. Works with any realtime voice API.

    One spoken turnbudget ≈ 1,000 ms
    respond event · skipped 0.34 ms p50 · reply lands 63.96 ms p99
    1. Transcript free
    2. Ask Crowkis
    3. Hit: inject reply
    4. Respond event skipped

    No inference billedNo synthesis billed

Nothing caller‑specific ever enters the cache.

If an answer can’t be de‑personalised, it isn’t cached.

  • 0.34 msp50 per spoken turn
  • 63.96 msp99 of a ~1,000 ms turn

Measured in our tests on the live voice path, single-threaded.

Meaning, structure, confidence, trust.

It matches what a question means, then checks the answer is safe to reuse.

CacheHow it matchesResult
Exact-match cacheIdentical text onlyMisses every rephrasing
Vector-only cacheAnything similarServes unsafe near-misses
CrowkisMeaning and structureReuses only when safe
  • Drop-in
  • RESP3
  • gRPC
  • REST
  • MCP
  • Built in Rust
  • One Docker image
  • Zero external API calls
  • Free community edition

Less spend. Faster answers. Safer AI.

  • 1model call per question

    without Crowkis×50

    with Crowkis×1

    Cut your LLM bill

    Pay for an answer once, however it is asked.

  • 0.4 msper cache hit

    model round-tripseconds

    cache hit0.4 ms

    Answer instantly

    Under a millisecond, not a multi-second model call.

  • 5gates, every hit, every time

    1. similarity
    2. template
    3. confidence
    4. trust
    5. freshness

    Keep answers safe

    Five checks before any answer is reused.

  • 1Docker image, every feature in

    docker pull crowkis/crowkis:latest

    Adopt it in minutes

    Your existing clients connect unmodified.

  • FreeCommunity edition, self-hosted

    no licence, no sign-up+ SSO · audit log · budgets

    Run it your way

    Self-hosted. Community is free.

A drop-in cache, written in Rust, that understands questions and reuses an answer only when it is safe.

Eight features, explained.

  • One-line integration get_or_compute

    Wrap your model call. A safe hit skips it; a miss runs it once.

  • Streaming

    Hits stream back like live model output.

  • Images too

    An image and its text match as one entry.

  • Reasoning reuse

    Reuses the reasoning, not just the words.

  • Adaptive thresholds

    Each intent tunes its own bar. Riskier means stricter.

  • Works with coding assistants MCP

    Claude Code checks the cache before spending tokens.

  • See everything

    Every hit, miss and block, live. Prometheus and OpenTelemetry built in.

  • Survives model upgrades

    Canary a new model, then migrate. Your cache stays warm.

What it doesn’t do.

  • Not a replacement for your LLM

    It decides: reuse or recompute.

  • Not a vector database for RAG

    Keep your RAG store. Crowkis fronts the model.

  • Never phones home

    Runs fully air-gapped.

  • No surprises on a miss

    A miss passes straight through.

Up and running in one minute.

One Docker image for the server, one package for your app.

  1. 01

    Run the server

    One container, three local ports.

    6379
    RESP3 clients and CLI
    6380
    HTTP: dashboard, REST, health
    6381
    gRPC
    Terminal
    # pull the image
    docker pull crowkis/crowkis:latest
    # run it
    docker run -d --name crowkis \
      -p 127.0.0.1:6379:6379 \
      -p 127.0.0.1:6380:6380 \
      -p 127.0.0.1:6381:6381 \
      -v crowkis-data:/data \
      crowkis/crowkis:latest
    # check it is up
    curl 127.0.0.1:6380/health
  2. 02

    Add the SDK

    Wrap your model call. Rephrasings hit the cache.

    Terminal
    pip install crowkis
    app.py
    from crowkis import Crowkis
    
    cache = Crowkis(host="127.0.0.1", port=6379, tenant="my-app")
    
    @cache.cached(ttl=3600)
    def answer(prompt: str) -> str:
        return my_model(prompt)
    
    answer("How do refunds work?")        # miss → model runs, cached
    answer("What's the refund process?")  # semantic HIT → served from cache
    Terminal
    npm install @crowkis/client
    app.mjs
    import { Crowkis } from "@crowkis/client";
    
    const cache = new Crowkis({ host: "127.0.0.1", port: 6379, tenant: "my-app" });
    
    const answer = await cache.ask(
      "How do refunds work?",
      async (prompt) => myModel(prompt),
      { ttl: 3600 },
    );
  3. 03

    Done. Ask twice, pay once.

    Optional: connect Claude Code or any agent over MCP.

    MCP
    claude mcp add crowkis -- crowkis mcp

Stop paying twice for the same answer.