Practical reads to help you spend less on AI.
Guides, benchmarks and deep dives on semantic caching, agent memory, LLM cost and LLM classification, from the team building Crowkis and Curva.
Subscribe with RSSOlder articles
Page 12 of 18-
Why we wrote our own LSM tree instead of bolting onto RocksDB
Every sane checklist says don't write your own storage engine. We did it anyway. Here's the actual reasoning, the architecture, and the parts that were painful.
-
How to cache DSPy LLM calls with Crowkis
Add a semantic cache to DSPy so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Confidence scoring: every hit arrives with a number you can gate on
A cache that only says 'hit' or 'miss' makes you trust it blindly. Crowkis returns a confidence score per hit, a geometric mean of five signals, so you decide the bar reuse must clear.
-
A tour of the dashboard: six panels, zero mysteries
Live verdicts, hit-type economics, top misses, safety blocks, tenant accounting, system pressure, what each panel answers and who keeps it open.
-
Agent unit economics: making the per-task math survive contact with reality
Agents multiply model calls per user action by 10-50x. Without aggressive reuse, the unit economics of agentic products simply don't close.
-
Give DSPy agents long-term memory with Crowkis
Durable, per-user memory for DSPy agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
Compliance modes: HIPAA, SOC2, GDPR-EU, FedRAMP as configuration
Each regime wants specific retention, audit, and erasure behavior. Enterprise compliance modes preset the whole posture, so the auditor's checklist maps to a flag.
-
Crowkis vs pgvector: your database deserves better than your cache traffic
pgvector is a lovely extension for storing embeddings next to your data. Routing every LLM query through Postgres is how lovely things die.
-
How to cache Instructor LLM calls with Crowkis
Add a semantic cache to Instructor so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Eviction with a ledger: why LRU is the wrong instinct for an LLM cache
LRU evicts by recency and nothing else. But cache entries have wildly different replacement costs, and forgetting a $0.40 answer to keep a $0.0004 one is just bad accounting.
-
Fintech assistants: fast answers, frozen correctness
Money questions repeat endlessly and tolerate zero staleness. Fintech is where freshness control stops being a feature and becomes the product.
-
Give Instructor agents long-term memory with Crowkis
Durable, per-user memory for Instructor agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
Adaptive thresholds: the cache tunes its own reuse bar over time
A fixed similarity threshold is wrong the day after you set it. Crowkis uses a three-tier scheme, per-intent base, complexity adjustment, and an EMA feedback loop, that learns the right bar and persists it.
-
The world's shortest cache runbook
Fail-open design means most 'incidents' are the absence of savings, not the presence of errors. Here's the whole decision tree, which fits on an index card.
-
Why Community is actually free: the honest economics of our free tier
Full engine, production use, no license, no meter, no time bomb. Here's why giving the small end away is the rational structure, not a teaser.
-
How to cache Pydantic AI LLM calls with Crowkis
Add a semantic cache to Pydantic AI so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
Four doors, four locks: the authentication architecture
RESP, gRPC, REST, and the dashboard each get auth that fits their use, constant-time tokens for the data plane, RBAC for the control plane, mandatory locks past loopback.
-
Crowkis vs Momento: your cache shouldn't bill like the thing it's saving you from
Serverless caches meter every operation. A cache that charges per request in front of an API that charges per request is a strange kind of savings.
-
Give Pydantic AI agents long-term memory with Crowkis
Durable, per-user memory for Pydantic AI agents that survives restarts and consolidates contradictions, self-hosted, zero egress.
-
Five TTL policies: engineering the shelf life of truth
Answers age at different speeds, prices in days, math never. A single TTL knob can't express that, so Crowkis ships five policies plus version pinning and webhooks.
-
Government and defense: the cache that works where the internet doesn't
Air-gapped networks, FedRAMP postures, and zero phone-home tolerance rule out most AI infrastructure on page one. Crowkis was designed to pass that page.
-
How to cache the Vercel AI SDK LLM calls with Crowkis
Add a semantic cache to the Vercel AI SDK so repeated and reworded questions are served for free, no rewrite, self-hosted.
-
CDEDUP: collapsing the answers that mean the same thing
A semantic cache slowly accumulates near-duplicate answers. CDEDUP finds the clusters that mean the same thing and collapses them, reclaiming memory, and Crowkis is honest about its cost.
-
Boring on purpose: the operational philosophy
Exciting infrastructure is a contradiction in terms. Every Crowkis design decision optimizes for the same review: 'it just runs.'