Engineering
How Crowkis is built: the Rust internals, the search engine and the decisions behind them. 37 articles.
Subscribe with RSSA Redis drop-in for AI: RESP3 compatibility, semantic brain
Crowkis speaks RESP3, so redis-py, ioredis, and Lettuce connect unmodified. Adoption is a port change, not a rewrite, and the semantic commands sit right beside the familiar ones.
Also new
Redis is the fastest cache alive. It also has no idea what your users are asking.
A million vectors, still instant: the search engine we wrote in Rust
You might not need a vector database
Prompt caching vs semantic caching: what the provider feature doesn't cover
GPTCache proved the idea. We went and rebuilt the engine underneath.
More articles
Page 1 of 2-
Restart-safe by design: a cache that survives a crash without losing a thing
A cache that forgets everything on restart isn't much of a memory. Crowkis writes durably first, so a crash, a deploy, or a reboot never costs you a single learned answer.
-
Why Rust for AI infrastructure
Why Rust for AI infrastructure. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Sub-millisecond retrieval: why in-process beats a network hop
Sub-millisecond retrieval: why in-process beats a network hop. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
LRU throws away your most expensive knowledge
Least-recently-used eviction is semantically blind. It happily discards the rare, costly answer you'll pay dearly to regenerate, and keeps the cheap trivia everyone asks. There's a smarter way.
-
The case for a Redis-compatible AI cache
The case for a Redis-compatible AI cache. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
-
Bring your own embedder / reranker: how it works and when to use it
Bring your own embedder / reranker, swaps in any sentence-transformers MiniLM or GTE export via ONNX, so you're never locked to the default model. Here's how Crowkis does it and why it matters for cost and safety.
-
One container, zero dependencies: what's deliberately absent from our image
The most secure dependency is the one that isn't there. Crowkis ships as a single stripped binary with the model baked in, no Python, no package manager, nothing to poison at runtime.
-
HNSW explained: finding the needle in a million haystacks
Once meaning is a point in space, the hard part is finding the nearest point out of millions, fast. HNSW is the elegant trick that makes it feel instant, here's the intuition.
-
Memcached walked so semantic caches could think
Classic caches taught the industry a durable lesson: never compute the same thing twice. LLMs just changed what 'the same thing' means, from identical bytes to identical meaning.
-
Why Crowkis is Rust all the way down
A cache lives in the hot path of every request. The language choice isn't aesthetic, it's the difference between predictable microseconds and mystery pauses.
-
Redis-compatible (RESP3): how it works and when to use it
Redis-compatible (RESP3), speaks RESP3 so redis-py, ioredis, and Lettuce connect unmodified across 40+ commands, adoption is a port change. Here's how Crowkis does it and why it matters for cost and safety.
-
Pure-Rust LSM storage engine: how it works and when to use it
Pure-Rust LSM storage engine, is a from-scratch WAL + MemTable + SSTable + bloom-filter + compaction engine in Rust, no RocksDB and no FFI. Here's how Crowkis does it and why it matters for cost and safety.
-
In-process HNSW vector index: how it works and when to use it
In-process HNSW vector index, keeps a custom, persistent HNSW graph in the same process as the store, so a lookup and its scoring never cross a network hop, sub-millisecond search. Here's how Crowkis does it and why it matters for cost and safety.
-
Group-commit WAL: how it works and when to use it
Group-commit WAL, batches fsyncs on a timer instead of per write, for materially higher write throughput when you enable it. Here's how Crowkis does it and why it matters for cost and safety.
-
i8 vector quantization: how it works and when to use it
i8 vector quantization, stores vectors in int8 for about 4x less memory, so more of your working set fits in RAM. Here's how Crowkis does it and why it matters for cost and safety.
-
How Crowkis earned the right to sit in your critical path
347 integration tests, a smoke suite that kills the process on purpose, and a Docker image hardened before anyone asked. The receipts behind 'production-ready.'
-
HNSW without the network hop: why the vector index lives inside the engine
Most semantic caches call out to a vector database. Crowkis embeds the HNSW graph in-process, and that placement decision is worth more than any algorithm tweak.
-
The write-ahead log: how the cache survives a kill -9
Durability isn't a checkbox, it's a sequence of writes in the right order with checksums at every step. Here's the boring machinery that makes restarts uneventful.
-
Bloom filters: how the engine knows what it doesn't know
The fastest disk read is the one that never happens. A few bits per key let Crowkis skip files that can't contain your answer, at a 1% false-positive cost we chose on purpose.
-
Twelve intents: why the cache treats a poem differently from a fact
One similarity threshold for all traffic is how caches embarrass themselves. Crowkis classifies every query into one of twelve intents, each with its own rules of reuse.
-
Structural templates: the matching layer vectors can't see
Embeddings blur exactly where caches need precision, numbers, dates, entities. Template abstraction catches what cosine similarity structurally cannot.