Features
What Crowkis does under the hood, one capability at a time. 75 articles.
Subscribe with RSSOlder articles
Page 2 of 3-
Model-free online evals (CEVAL): how it works and when to use it
Model-free online evals (CEVAL), grades output with deterministic evaluators (toxicity, PII, relevance, JSON validity and more) and tracks the results over time, no LLM-judge. Here's how Crowkis does it and why it matters for cost and safety.
-
Prompt versioning and A/B testing at the data layer
Prompt versioning and A/B testing at the data layer. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Prompt versioning and A/B testing (CPROMPT): how it works and when to use it
Prompt versioning and A/B testing (CPROMPT), versions prompts on every write, renders variables, and runs sticky per-user A/B splits, so you roll back or test without a code deploy. Here's how Crowkis does it and why it matters for cost and safety.
-
Reasoning reuse: stop paying for the same chain of thought twice
Chain-of-thought costs several times more tokens than the answer it produces, and it's usually thrown away. For structurally similar problems, most of that reasoning can be reused.
-
Streaming responses from cache without breaking UX
Streaming responses from cache without breaking UX. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Self-hosted RAG (CDOC): how it works and when to use it
Self-hosted RAG (CDOC), adds documents with auto-chunking and metadata, then runs filtered ANN search with optional reranking, no separate vector database. Here's how Crowkis does it and why it matters for cost and safety.
-
One similarity cutoff is always wrong
'What's 2+2?' needs a near-exact match to reuse safely. 'Give me creative ideas for X' can tolerate a loose one. A single global threshold guarantees you're too strict somewhere and too loose somewhere else.
-
Local offline embeddings (CEMBED): how it works and when to use it
Local offline embeddings (CEMBED), turns text into vectors with the bundled local model, no API key and no egress, with a micro-cache so repeats are free. Here's how Crowkis does it and why it matters for cost and safety.
-
Multi-turn session memory (CSESSION): how it works and when to use it
Multi-turn session memory (CSESSION), keeps a bounded conversation buffer with both recent-window reads and semantic search across the whole chat. Here's how Crowkis does it and why it matters for cost and safety.
-
The agentic era needs a memory layer. Here's what it looks like.
We gave agents tools, planning, and the ability to act. We forgot to give them a place to remember. That gap is why your agents feel brilliant and amnesiac at the same time.
-
Tool-result caching (CTOOLSET): how it works and when to use it
Tool-result caching (CTOOLSET), caches a deterministic tool call keyed by tool plus exact args, so a swarm's duplicate lookups become one call. Here's how Crowkis does it and why it matters for cost and safety.
-
Multimodal caching (image + text): how it works and when to use it
Multimodal caching (image + text), caches image-plus-text lookups, so a repeated vision question is a hit instead of an expensive re-run. Here's how Crowkis does it and why it matters for cost and safety.
-
What is an embedding, really? A plain-English guide
Embeddings sound like math you need a PhD for. The core idea is simpler and more useful than that, and it's the reason a cache can tell that two different sentences mean the same thing.
-
Streaming response caching: how it works and when to use it
Streaming response caching, serves cached answers chunk by chunk, so a hit feels like live typing and the seam between hit and miss disappears. Here's how Crowkis does it and why it matters for cost and safety.
-
CGUARD: an input guardrail that survives leetspeak and zero-width tricks
Prompt injection rarely arrives in plain English. CGUARD normalizes the evasion first, whitespace, leetspeak, zero-width characters, then scans for jailbreaks, overrides, and system-prompt exfiltration.
-
OpenAI-compatible AI gateway: how it works and when to use it
OpenAI-compatible AI gateway, proxies /v1/chat/completions with a semantic cache in front, so you point your client's base URL at Crowkis and change nothing else. Here's how Crowkis does it and why it matters for cost and safety.
-
Multi-provider routing and fallback: how it works and when to use it
Multi-provider routing and fallback, load-balances and fails over across providers on error class, with retries using exponential backoff and jitter. Here's how Crowkis does it and why it matters for cost and safety.
-
COUTCHECK: catching the PII leak and the toxic line before your user does
The model's output is the other trust boundary. COUTCHECK scans responses for PII leakage and toxicity, and optionally validates JSON, returning a structured verdict you can act on.
-
Give your coding agent a memory
Coding agents re-read the same schema, re-derive the same conventions, and re-ask the same architecture questions on every run. A shared memory turns that repeated context into a one-time cost.
-
CEVAL: nine evaluators that grade your LLM output without a second LLM
LLM-as-judge is expensive and leaks your data. CEVAL ships nine deterministic evaluators, toxicity, PII, injection-safety, relevance, JSON validity and more, that score input/output pairs locally and track the results over time.
-
Golden answer pinning (CPIN): how it works and when to use it
Golden answer pinning (CPIN), serves a human-approved answer verbatim for any phrasing of a question, with an audit trail of who approved it. Here's how Crowkis does it and why it matters for cost and safety.
-
Answer lineage and cascade purge (CSOURCE): how it works and when to use it
Answer lineage and cascade purge (CSOURCE), ties answers to their source so that when a document changes, every answer built on it can be purged in one move. Here's how Crowkis does it and why it matters for cost and safety.
-
CPROMPT: version your prompts and A/B test them like code
Prompts are production logic edited like sticky notes. CPROMPT gives them named templates, automatic versioning, rollback, variable rendering, and sticky weighted A/B splits, all surviving restart.
-
Natural-language cache invalidation (CINVALIDATE): how it works and when to use it
Natural-language cache invalidation (CINVALIDATE), purges entries whose meaning matches a plain-English instruction, previewing by default and only acting on COMMIT. Here's how Crowkis does it and why it matters for cost and safety.