Features
What Crowkis does under the hood, one capability at a time. 75 articles.
Subscribe with RSSOlder articles
Page 3 of 3-
Stale-while-revalidate (CSTALE): how it works and when to use it
Stale-while-revalidate (CSTALE), returns a cached answer past its TTL with a stale flag, so expiry is a snappy answer plus a refresh signal, not a cold miss. Here's how Crowkis does it and why it matters for cost and safety.
-
CDOC: a self-hosted RAG store that chunks, filters, and reranks
You don't always need a separate vector database to do retrieval. CDOC is a mini RAG store inside Crowkis, auto-chunking, metadata filtering, and optional cross-encoder reranking, sharing the cache's embedder.
-
CSESSION: conversation buffers with semantic recall built in
Chat history is more than the last N turns. CSESSION stores a multi-turn buffer per session, bounded and TTL'd, with both recent-window reads and semantic search across the whole conversation.
-
CPIN: golden answers that are served verbatim, with an audit trail
For the questions where 'close enough' is unacceptable, pricing, legal, brand lines, CPIN serves a human-approved answer verbatim, records who approved it, and never lets the model improvise.
-
CFLAG and CCHECKBAD: a memory for the answers that were wrong
Most caches only remember good answers. Crowkis also remembers bad ones, flag a hallucinated or harmful response once, and CCHECKBAD catches every paraphrase of the question that would have reproduced it.
-
CSOURCE: answer lineage and cascade-purge when the source changes
Cached answers derive from sources, a doc, a config, an API. CSOURCE ties answers to their origin so that when the source changes, every answer built on it can be purged in one move.
-
CTOOLSET: cache the tool call so the agent stops paying for it twice
Agents call the same tools with the same arguments constantly. CTOOLSET and CTOOLGET cache tool results keyed by tool plus exact arguments, so a deterministic call runs once and serves many.
-
MCP server for AI apps: how it works and when to use it
MCP server for AI apps, doubles as an MCP server over stdio, so Claude and any MCP-capable agent can check the cache and store what they compute. Here's how Crowkis does it and why it matters for cost and safety.
-
CINVALIDATE: purge the cache by meaning, with a preview before you commit
Sometimes you need to clear 'everything about the old pricing', a fuzzy, semantic set. CINVALIDATE takes a natural-language instruction, previews what it would purge, and only acts on COMMIT.
-
CSTALE: serve the slightly-old answer now, refresh it behind the scenes
A hard TTL turns a one-second-expired answer into a full model call. CSTALE serves the cached answer past its TTL with a stale flag, so you choose freshness versus latency per request.
-
CBUDGET: per-tenant spend you can see before the invoice does
Token spend is usually a month-end surprise. CBUDGET tracks per-tenant token and dollar consumption in real time and surfaces alerts, so a runaway tenant is a notification, not a billing shock.
-
The AI Gateway: a semantic cache in front of any OpenAI-compatible API
Point your existing OpenAI client at Crowkis and change nothing else. The gateway proxies /v1/chat/completions, serves semantic hits without an upstream call, and adds retries, routing, and rate limits.
-
CEMBED: free local embeddings, cached, with no API key
Embeddings usually mean an API key and a per-token bill. CEMBED turns text into vectors using the bundled ONNX model, locally, for free, and caches repeats so the second call is instant.
-
Caching what the model saw: multimodal image-plus-text lookups
Vision queries are expensive and repetitive, the same product photo, the same screenshot, asked about again and again. Crowkis caches image-plus-text lookups so a repeated visual question is a hit.
-
Confidence scoring: every hit arrives with a number you can gate on
A cache that only says 'hit' or 'miss' makes you trust it blindly. Crowkis returns a confidence score per hit, a geometric mean of five signals, so you decide the bar reuse must clear.
-
Adaptive thresholds: the cache tunes its own reuse bar over time
A fixed similarity threshold is wrong the day after you set it. Crowkis uses a three-tier scheme, per-intent base, complexity adjustment, and an EMA feedback loop, that learns the right bar and persists it.
-
CDEDUP: collapsing the answers that mean the same thing
A semantic cache slowly accumulates near-duplicate answers. CDEDUP finds the clusters that mean the same thing and collapses them, reclaiming memory, and Crowkis is honest about its cost.
-
CPII: scrubbing personal data and honouring the right to be forgotten
A cache of LLM traffic is a cache of whatever users typed, including PII. CPII reports what personal data is present and executes right-to-erasure, so compliance is a command, not a project.
-
CINFO and the dashboard: a cache you can actually watch work
Infrastructure you can't observe is infrastructure you don't trust. CINFO and the built-in dashboard expose hit rate, saved spend, safety blocks, memory pressure, and license state in real time.
-
CKEYLIMIT: per-tenant rate limits that stop the runaway before it starts
A runaway agent or a noisy tenant can torch a budget in minutes. CKEYLIMIT sets per-tenant requests-per-minute and tokens-per-minute ceilings, enforced locally before the spend happens.
-
CTHINK and CREUSE: banking a chain of thought and replaying it
The reasoning is the expensive part of a hard answer. CTHINK stores a chain-of-thought trace as a reusable step graph; CREUSE fetches the matching plan for a new query at a fraction of the original token cost.
-
What is agent memory, and why your agents need it
What is agent memory, and why your agents need it. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
RAG in production: the parts nobody warns you about
RAG in production: the parts nobody warns you about. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Semantic caching, explained for engineers
Semantic caching, explained for engineers. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.