Operations
Running Crowkis in production: cache warming, rate limits, dedup and dashboards. 16 articles.
Subscribe with RSSHow to warm an LLM cache during a model migration
How to warm an LLM cache during a model migration. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.
More articles
-
Model migration without a cold cache: how it works and when to use it
Model migration without a cold cache, canary and migration workflows carry cache value across a model upgrade, so a new model doesn't cold-start your hit rate. Here's how Crowkis does it and why it matters for cost and safety.
-
Three windows into one cache: dashboard, Prometheus, logs
The built-in dashboard for humans, /metrics for your Grafana, one JSON line per event for your pipeline, same truth, three consumers, zero adapters.
-
Crowkis on Kubernetes: a well-behaved citizen
One container, a PVC, real health probes, hard memory bounds, graceful shutdown. Everything your cluster expects from a tenant that's read the manual.
-
Running a model canary: the operator's walkthrough
Slice the traffic, compare against cached baselines, promote or retreat, model upgrades as a controlled experiment with the cache as your measuring instrument.
-
Fallback routing: surviving your provider's bad day
Providers have incidents; your product doesn't have to. Health-aware backend routing plus a warm cache turns upstream outages into degraded modes users barely notice.
-
Memory governance: a cache that respects its container
CROWKIS_MEMORY_LIMIT means what it says, no GC mood swings, no mystery RSS, eviction that engages before the kernel has opinions.
-
A tour of the dashboard: six panels, zero mysteries
Live verdicts, hit-type economics, top misses, safety blocks, tenant accounting, system pressure, what each panel answers and who keeps it open.
-
The world's shortest cache runbook
Fail-open design means most 'incidents' are the absence of savings, not the presence of errors. Here's the whole decision tree, which fits on an index card.
-
Boring on purpose: the operational philosophy
Exciting infrastructure is a contradiction in terms. Every Crowkis design decision optimizes for the same review: 'it just runs.'
-
LLM observability: what to actually measure
LLM observability: what to actually measure. A practical, Crowkis-grounded take, no hype, just what actually moves cost, latency, and safety.