The Local Prompt Injection Firewall

Intercept, inspect and block LLM requests on your own machine. Nothing leaves it, and clean traffic is not slowed down.

Free for noncommercial use under PolyForm Noncommercial 1.0.0 · commercial licence for company use

localhost:7731
llm-fw dashboard showing a live feed of intercepted requests with detection stage and score

Covers every major provider out of the box — OpenAI, Anthropic, Google Gemini and Vertex, Azure OpenAI, AWS Bedrock, Mistral, Groq, DeepSeek, Cohere, xAI, Perplexity, OpenRouter, Together, Fireworks, Anyscale and Hugging Face.

100% Attacks detected
110/110, heuristic + embedding stages
0% False positives
0/78 benign prompts
95% Recall on unseen attacks
1,071 real InjecAgent cases, never tuned on

Full methodology and held-out numbers: Detection Scorecard · Generalization Benchmark

Watch what it blocks

Every intercepted request lands in a dashboard on localhost:7731, with the stage that caught it, the score, and the decoded payload.

localhost:7731
llm-fw dashboard screenshot

Events feed: intercepted requests appear as they happen, tagged with the stage that caught them — heuristic, embedding, dos, rag or dlp — and their score.

A proxy your tools already know how to use

Point HTTPS_PROXY at 127.0.0.1:8080 and every SDK, CLI and agent that respects it is covered. No SDK wrapper, no code change.

  • Local certificate authority: setup generates a root CA and registers it in your OS trust store, so TLS can be terminated and inspected locally.
  • OS-level sinkhole: for Node.js and native binaries that ignore proxy variables, a hosts entry plus a :443 port redirect brings them in anyway.
  • No cloud dependency: no API keys, no accounts, no telemetry. Detection models run on your CPU and the event log stays on disk.
Diagram of llm-fw sitting between local tools and LLM provider APIs
Three-Stage Defense

Cheap checks first, expensive ones only when needed

A prompt falls through three stages. Most traffic never reaches stage two.

01

Heuristic scoring

< 1ms

Weighted phrase matching over a normalisation pass built to survive evasion: spacing, case, homoglyphs and leetspeak, plus decoding of base64, base32, hex, binary, morse, ROT13, URL-encoding and reversed text back to plaintext.

02

Cross-lingual embeddings

< 20ms warm

Cosine similarity against canonical injection-intent anchors using a local multilingual-e5-small ONNX model. Because the encoder aligns around 100 languages, an injection written in Urdu or Thai lands near the English anchors with no rule written for it.

03

Local LLM judge

Opt-in

A local Ollama model — qwen2.5:3b by default — reasons about intent on the prompts the cheap stages are unsure about. An opt-in ONNX classifier can take that role instead, with no Ollama required.

Injection is not the only thing worth stopping

Outbound secret scanning

52 named credential patterns — AWS, GitHub, GitLab, Stripe, Slack, npm, PyPI, Vault, database URIs, private keys and the API keys of every LLM provider — plus Luhn-validated card numbers and US social security numbers. Matches are redacted before the request leaves.

AWS_SECRET_ACCESS_KEY = [REDACTED_AWS_SECRET_KEY]

Cost and loop circuit breaker

Behavioural rate limiting on tokens and requests, and a breaker for the failure mode that empties a budget overnight: an agent stuck in a loop, resending near-identical payloads until someone notices the bill.

Untrusted content, judged separately

Instructions hidden in retrieved documents, tool results and code fences are scanned as data, not as user intent — and you can set a tighter threshold there than on the prompt surface. Invisible-Unicode smuggling and text rendered into pasted screenshots are decoded and scanned too.

Licensing

llm-fw is not open source. It is free for noncommercial use, and paid for commercial use.

Noncommercial

Free

No key, no signup, no telemetry.

  • Personal projects, hobby work and private study
  • Research and teaching
  • Charities, schools, public research bodies and government
  • Every feature — nothing is held back behind the paid tier
Install it

Which one applies to you

The PolyForm Noncommercial License 1.0.0 does not grant commercial use. If llm-fw runs anywhere in a for-profit organisation's work you need a commercial licence — including on a single developer's laptop, in CI, and on a shared standalone server. Embedding or redistributing llm-fw inside a product you sell needs an OEM agreement rather than a standard licence.

Commercial tiers

The developer count is the number of developers whose traffic llm-fw inspects. Billed yearly.

DevelopersPrice per year
1 to 9 €450
10 to 49 €1,500
50 to 199 €4,000
200+ €9,000
OEM / redistribution from €15,000, negotiated Email

Checkout is handled by Paddle, our merchant of record, which acts as the seller of record and handles VAT and sales tax. PayPal and card are both accepted at checkout. Prices exclude tax; it is added at checkout, and shown in your local currency where Paddle supports it. Your licence key is sent by email after purchase — it is issued by hand, so please allow one business day.

Prefer an invoice or a purchase order, or want to talk first? Email us. Read the terms and the refund policy.

Quick installation

Three commands. setup generates the local CA and registers it in your trust store; the sinkhole step needs administrator rights, and --proxy-only skips it.

Node.js 22 or newer. Windows, macOS and Linux.

BASH
npm install -g llm-fw
llm-fw setup
llm-fw start

Questions

No. Heuristics, the ONNX embedding model, the optional judge and the event database all run locally. llm-fw makes no cloud calls of its own and sends no telemetry. The only traffic that leaves is the traffic you were already sending to your LLM provider.

The streaming heuristic pass completes in under a millisecond, and clean traffic is forwarded as it streams rather than buffered and re-sent. The embedding stage adds under 20ms once warm, and only runs when the cheap stage is unsure. The optional LLM judge is the slow one, which is why it is opt-in and reserved for the ambiguous middle.

llm-fw setup generates a root certificate authority on your machine and registers it in the OS trust store, which lets the proxy terminate TLS locally, inspect the request, and re-establish it upstream. The CA private key never leaves the machine. llm-fw uninstall removes it from the trust store again.

The blocked request appears in the dashboard with the rule that matched and the score it got. Marking it a false positive records it; on a blocked prompt you can also suppress future matches, which downgrades an identical prompt to a warning instead of a block. Thresholds are configurable per surface, so you can run tighter on tool output than on typed prompts.

Not if the work is commercial. PolyForm Noncommercial covers personal projects, study, research, teaching, charities, schools, public research bodies and government institutions. Running llm-fw as part of a for-profit company's work is commercial use — even on one laptop, even in CI — and needs a commercial licence. Evaluating it before you buy is fine; email us if you want longer than a couple of weeks.

No. It is the same software from the same npm package, with every feature available either way. What you buy is the right to use it commercially, plus priority support. Nothing is crippled to sell an upgrade, and there is no phone-home check.