A PII Filter in Front of Cloud Models: Why Self-Hosting LLMs Isn't Necessary
Published: 2026-10-06 · Author: AI Release · @ai_release1
⚡ The gist in 5 seconds - The point: the author gave up on fully self-hosting an LLM and instead deployed their own PII filter in front of cloud models. - Where it's available: the self-hosted guardrails-llm-filter from Cloud.ru (Apache-2.0) is used as a proxy between the client and the cloud model. - Limitation: in enforce mode the filter replaces personal data and secrets with placeholders, while in detect mode it only logs them; if masking fails, a request may go through unprocessed — this is now tracked as an incident. ### 🔍 What was found Over the past two months, the author tested the RTX PRO 6000 96GB, RTX4090 48GB, and even the H200 for self-hosting, bought an RTX5070ti and a V100 32GB, and ran a large number of benchmarks. After hardware prices rose and a string of personal data leaks at LLM providers, they focused on filtering. The first filter went live on September 17, and on September 22 the system went down: due to a missing home directory for the system user, go build failed to produce the binary, Ansible didn't rebuild it, and systemd unsuccessfully tried to start the process more than 2,600 times over about four hours. The cause turned out to be detect mode accidentally left in the settings; after the fix, enforce was set both as the startup default and in the applied settings. Before this, the author used a hook inside the client, but it couldn't restore the original values in the response — model replies contained placeholders like [EMAIL_01] and [EMAIL_02]. After switching to guardrails-llm-filter, processing happens on both the request and the response. LiteLLM was rejected after testing: the intermediate proxy breaks the shared prefix, causing prompt cache savings to drop by roughly 90% — and the author has almost 1.5B cached tokens over three weeks. ### 💡 Why it matters The practical benefit is that to control sensitive data you don't necessarily need to run an LLM on your own hardware: it's enough to sanitize requests and responses on your side. The author keeps using top-tier cloud models while reducing the leak risk, and has left local Gemma4 and Qwen3.8 on standby on a laptop and an LXC with an RTX5070ti for other tasks. It's a working compromise between the performance of cloud models and privacy — without giving up the cloud. ### 🧩 Context Self-hosting was initially considered the primary option, but tests with the RTX PRO 6000 96GB, RTX4090 48GB, and H200, plus the purchase of an RTX5070ti and V100 32GB, showed that local models still fall short of top cloud offerings. At the same time, the author saw a crazy rise in hardware prices and a string of leaks at LLM providers, which accelerated the search for a solution. The filtering system now sits between the cloud and all traffic, and two monitors in OneUptime track heartbeat and fail-open metrics: the metrics contain only counters, data types, and latencies — no request texts. After today's restart with a GPU swap, the filter came up in enforce mode with six data types and fully loaded settings.
⚡ The gist in 5 seconds - The point: the author gave up on fully self-hosting an LLM and instead deployed their own PII filter in front of cloud models.
- Where it's available: the self-hosted guardrails-llm-filter from Cloud.ru (Apache-2.0) is used as a proxy between the client and the cloud model.
- Limitation: in enforce mode the filter replaces personal data and secrets with placeholders, while in detect mode it only logs them; if masking fails, a request may go through unprocessed — this is now tracked as an incident.
🔍 What was found Over the past two months, the author tested the RTX PRO 6000 96GB, RTX4090 48GB, and even the H200 for self-hosting, bought an RTX5070ti and a V100 32GB, and ran a large number of benchmarks.
After hardware prices rose and a string of personal data leaks at LLM providers, they focused on filtering.
The first filter went live on September 17, and on September 22 the system went down: due to a missing home directory for the system user, go build failed to produce the binary, Ansible didn't rebuild it, and systemd unsuccessfully tried to start the process more than 2,600 times over about four hours.
The cause turned out to be detect mode accidentally left in the settings; after the fix, enforce was set both as the startup default and in the applied settings.
Before this, the author used a hook inside the client, but it couldn't restore the original values in the response — model replies contained placeholders like [EMAIL_01] and [EMAIL_02].