CGUARD: an input guardrail that survives leetspeak and zero-width tricks
Prompt injection rarely arrives in plain English. CGUARD normalizes the evasion first, whitespace, leetspeak, zero-width characters, then scans for jailbreaks, overrides, and system-prompt exfiltration.
The naive prompt-injection filter loses to a five-minute workaround: insert a zero-width space, swap an 'o' for a '0', pad with newlines, and the regex that caught 'ignore previous instructions' sails right past 'ignore prev1ous 1nstructions'. CGUARD assumes the attacker knows this, so it normalizes before it matches, collapsing whitespace, folding leetspeak, stripping zero-width characters, and only then runs the scan.
In plain words. Attackers disguise prompt injections with odd spacing and character swaps. CGUARD undoes the disguise first, then checks, so the trick that beats a simple filter doesn't beat this one.
What it scans for is the real catalogue: DAN-style jailbreaks, developer-mode prompts, instruction overrides detected by verb-plus-noun co-occurrence rather than fixed strings, and attempts to exfiltrate the system prompt. The return is structured, a verdict, a category, and the matched span, so your app can log the category, show the user a clean refusal, and route the rest onward.
flowchart TD
W["candidate write"] --> S1["coherence · 0.30"]
S1 --> S2["content · 0.10"]
S2 --> S3["source trust · 0.30"]
S3 --> S4["isolation · 0.15"]
S4 --> S5["neighbourhood · 0.15"]
S5 --> G{"composite ≥ 0.75?"}
G -- yes --> OK["accepted"]
G -- no --> NO["refused + ledger entry"]
style OK fill:#fbe9e8,stroke:#d62221,stroke-width:2.5px
style NO fill:#f3eee5It runs as a command, CGUARD, which means it composes anywhere: in front of a cache write, in front of a model call, or as a standalone gate in a pipeline that isn't even using Crowkis for caching yet. No model call, no egress, the detection is deterministic and local, so it costs microseconds and leaks nothing.
The bottom line
A guardrail is only worth shipping if it assumes an adversary, not a typo. CGUARD is built for the person actively trying to break your agent, which is exactly the person a string-match filter was never going to stop.