Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSOlder articles
Page 3 of 18-
A self-hosted Jev alternative for typed LLM decisions
Looking for a self-hosted Jev alternative? How Curva compares on features, measured accuracy, speed and cost, and how to switch by changing two variables.
-
Semantic cache vs vector database: they solve different problems
A vector database is built for large-scale retrieval. A semantic cache is built for safe answer reuse. Using one for the other's job is where teams get burned.
-
What is a semantic cache for LLMs? (and why exact-match caching fails)
A plain key-value cache misses the moment a prompt is reworded, and a raw vector cache can serve the wrong answer. A semantic cache understands meaning and structure, and only reuses when it's safe.
-
Reasoning reuse: cache the chain of thought, not just the answer
The expensive part of a hard answer is the thinking. Crowkis stores reasoning as a reusable step graph and replays it for the next question that shares its shape, at a fraction of the token cost.
-
How to cut LLM API costs with semantic caching
The cheapest token is the one you never spend twice. Here's the simple math behind semantic caching, and where the savings actually come from.
-
Self-hosted RAG with CDOC: chunking, metadata filters, reranking
If your corpus fits a cache, you don't need a separate vector database to do retrieval. CDOC adds documents with auto-chunking, filtered search, and reranking, all local.
-
Prompt-injection and jailbreak detection at the cache layer
Attackers disguise injections with odd spacing and character swaps. CGUARD normalizes the disguise first, then scans, so the trick that beats a naive filter doesn't beat this.
-
We put our tiny embedding model up against OpenAI and NVIDIA. It didn't blink.
crowsight is a small, offline embedding model that ships inside Crowkis. We didn't trust it on faith, so we made it compete with the biggest embedding APIs on the one job a semantic cache actually needs. Here's what happened.
-
Give Claude and your agents a cache via MCP
The Crowkis binary doubles as an MCP server, so Claude Desktop, Claude Code, and any MCP-capable agent can check the cache before calling the model and store what they compute.
-
A drop-in OpenAI-compatible AI gateway with a semantic cache in front
Point your existing OpenAI client at Crowkis and change nothing else. Repeated questions are served from cache with no upstream call, and you get retries and routing for free.
-
Stop paying twice for the same answer
Most LLM bills are quietly full of duplicates, the same question, reworded, billed at full price every time. Semantic caching is how you stop paying for an answer you already have.
-
Local, offline embeddings with CEMBED (no external API)
Embeddings usually mean an API key and a per-token bill. CEMBED turns text into vectors using the bundled local model, for free, with nothing leaving your machine.
-
A Redis drop-in for AI: RESP3 compatibility, semantic brain
Crowkis speaks RESP3, so redis-py, ioredis, and Lettuce connect unmodified. Adoption is a port change, not a rewrite, and the semantic commands sit right beside the familiar ones.
-
Redis is the fastest cache alive. It also has no idea what your users are asking.
Redis is a masterpiece, for exact-match lookups. But nobody asks your app exact-match questions. Here's why we kept its wire protocol and taught the cache to understand meaning.
-
Long-term memory for AI agents, explained
Most agents forget the moment a session ends. Real memory consolidates contradictions, blends relevance with recency, and can even tell you what it believed at a past point in time.
-
Looking for a GPTCache alternative? What to compare
If you're evaluating semantic caches, similarity is the easy part. The differences that matter in production are safety, isolation, confidence, and cost control.
-
A million vectors, still instant: the search engine we wrote in Rust
Finding the nearest meaning among a million cached answers, in under a millisecond, without a single external dependency. A look at the pure-Rust HNSW engine underneath Crowkis.
-
MCP explained: giving models tools they can hold
MCP explained: giving models tools they can hold. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Semantic + structural cache matching: how it works and when to use it
Semantic + structural cache matching, matches questions by meaning (HNSW vectors) and by structure (12 intent-class templates), so paraphrases hit but a changed number, entity, or negation does not. Here's how Crowkis does it and why it matters for cost and safety.
-
Your AI agent has amnesia. We gave it memory that stays in its lane.
Most AI agents forget everything the moment a session ends, then re-pay to relearn it. Crowkis gives agents durable, semantic memory, and keeps every tenant's memory strictly walled off from the next.
-
Vector databases vs semantic caches: pick the right tool
Vector databases vs semantic caches: pick the right tool. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Adaptive confidence thresholds: how it works and when to use it
Adaptive confidence thresholds, learns the right reuse bar per intent class with a feedback loop, and persists it across restarts, so it stops both over-serving and missing safe hits. Here's how Crowkis does it and why it matters for cost and safety.
-
crowjudge: the bouncer that keeps your cache honest
A cache that serves a wrong answer is worse than no cache. Meet crowjudge, the second model that re-reads borderline matches and vetoes the ones that don't hold up.
-
Token cost management for AI products
Token cost management for AI products. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.