Crowkis vs the dedup script: the cron job that thinks it's a cache
Somewhere in your repo is a script that hashes prompts and skips duplicates. It's doing its best. Here's everything it can't see.
The dedup script is folk engineering at its most charming: normalize the prompt, hash it, skip the call if the hash repeats. It catches the easiest tenth of the waste, verbatim repeats inside one process's window, and it does so with a confidence its design cannot justify. Lowercasing and whitespace-stripping is not semantics; 'refund timeline?' and 'how long do refunds take?' hash to different planets.
In plain words. Hashing catches questions spelled the same. Real users never spell anything the same. You need matching that survives contact with humans.
The script also has no opinion about safety, because hashing has no opinions at all. Whatever response got stored gets replayed: the hallucination, the answer computed for a different tenant, the instruction that was true before Tuesday's pricing change. No confidence floor, no trust history, no freshness, no audit trail when someone asks why.
flowchart TD Q["incoming query"] --> I["intent classifier"] I --> T["template match"] T --> V["HNSW neighbours"] V --> C["confidence gate"] C --> TR["trust + freshness"] TR -- pass --> A["answer · <1ms"] I -. veto .-> M["(nil) → your model"] T -. veto .-> M V -. veto .-> M C -. veto .-> M TR -. veto .-> M style A fill:#fbe9e8,stroke:#d62221,stroke-width:2.5px style M fill:#f3eee5
Crowkis is what the script wishes it were when it grows up: normalization plus templates plus embeddings plus intent classes on the matching side; confidence, trust, tenancy, and TTL policies on the safety side; durability, dashboards, and four protocols on the infrastructure side. The integration is the same size as the script's, one wrapper call.
The bottom line
Retire the cron job with honors. It identified the right problem; it was just never going to be the answer. The answer needed an engine.