Engineering
How Crowkis is built: the Rust internals, the search engine and the decisions behind them. 37 articles.
Subscribe with RSSOlder articles
Page 2 of 2-
Reasoning reuse: caching how the model thinks, not just what it says
Chain-of-thought tokens are the most expensive ones you buy. Crowkis extracts the thought's skeleton, abstracts the specifics, and recomposes it for the next input that shares its shape.
-
Why we wrote our own LSM tree instead of bolting onto RocksDB
Every sane checklist says don't write your own storage engine. We did it anyway. Here's the actual reasoning, the architecture, and the parts that were painful.
-
Eviction with a ledger: why LRU is the wrong instinct for an LLM cache
LRU evicts by recency and nothing else. But cache entries have wildly different replacement costs, and forgetting a $0.40 answer to keep a $0.0004 one is just bad accounting.
-
Five TTL policies: engineering the shelf life of truth
Answers age at different speeds, prices in days, math never. A single TTL knob can't express that, so Crowkis ships five policies plus version pinning and webhooks.
-
Why we kept the Redis protocol instead of inventing an API
Every new API is a tax on adoption: clients, docs, muscle memory, tooling. RESP3 meant inheriting twenty years of all four on day one.
-
One actor, no locks across await: the concurrency design
Crowkis serves thousands of connections through async IO, then funnels every cache decision through a single deterministic actor. Here's why that's a feature.
-
Designing the MCP server: a cache as a tool the model can hold
MCP turns Crowkis into something an AI assistant can use deliberately, check the cache, store the answer, over plain stdio, with the banner silenced so JSON-RPC stays clean.
-
Three levels, one strategy: compaction without the tuning PhD
LSM compaction is where storage engines breed complexity. Crowkis ships exactly one strategy across three levels, chosen for cache workloads, closed for configuration.
-
Streaming cache hits: instant answers that still feel like typing
Users expect LLM answers to arrive as a typing stream. CGETSTREAM serves cached answers chunk by chunk, so a sub-millisecond hit doesn't break the interface's rhythm.
-
347 tests and a murder weapon: how the suite is organized
Bottom-heavy by design: the layers that hold your data get the most hostile coverage, and the smoke suite's signature move is killing the process to prove a point.