A trust controller for LLM agent memory: the MyCity case study
Published: 2026-10-04 · Author: AI Release · @ai_release1
⚡ The gist in 5 seconds - The gist: an MTS blog post on Habr examines a trust controller for LLM agent memory — a Memory Decision Layer (MDL) designed to reduce hallucinations in RAG systems. - Where to find it: the piece was published October 3 in the MTS company blog; reading time is 5 minutes. - Limitation: the controller doesn't guarantee accuracy — 8% of decisions it approves result in hallucinations, and with high risk and relevance, accuracy drops to 55.6%. ### 🔍 What we found The trigger for the analysis was MyCity, a chatbot launched in 2024 by the New York City mayor's office for small business owners. It was supposed to be a convenient reference for city rules and laws, but it dispensed advice that could get entrepreneurs fined or sued: for example, not paying tips to employees until the end, or firing someone for reporting harassment. Its database had correct information sitting right next to outdated material, and the bot couldn't choose between them. The authors explain the problem with classic RAG. The retrieval module finds documents by semantic similarity and feeds them to the language model as context, but the model doesn't know whether to trust those documents. A team from Tianjin University of Technology measured the scale: on TruthfulQA, when correct facts and popular misconceptions are mixed into memory, vanilla RAG produces 53% hallucinations, while a model without memory errs 23% of the time. MDL adds a verification step: instead of a single relevance score, it looks at relevance (cosine similarity), reliability (searching for negations, antonyms, contradictory polarities), and task risk (medicine, law, finance, conspiracy theories — a coefficient starting at 0.70). The scores are laid out across independent orthogonal subspaces via QR decomposition, so relevance can't mask contradictions. At high risk, the coefficient is subtracted from one, shortening the trust vector; if it falls below a threshold, the controller refuses to answer substantively and reports that the question is too sensitive. All of this takes about 40 microseconds. ### 💡 Why it matters In conflict scenarios, the controller significantly reduces hallucinations — from 53% with vanilla RAG to 23.3%, and on highly sensitive queries it makes no errors at all. But the effect depends heavily on the base model. On the relatively weak gemma-4-E4B-it, MDL delivers a tangible gain, while on the stronger deepseek-v4-flash, vanilla RAG already holds hallucinations at 4.3%, and with the controller it's 4.7% — within statistical noise; the benefit remains only in high-risk topics. It's also important that even a green light from the controller doesn't guarantee a clean result: in 8% of cases where it lets a document through, the answer turns out to be a hallucination. And in the zone of simultaneous high risk and high relevance, decision accuracy drops to 55.6% — barely better than a coin flip. MDL makes decisions by geometric rules and doesn't understand the meaning of the situation, so it can't be relied upon as the sole line of defense. ### 🧩 Context The MDL idea was born from a neuroscientists' experiment with the prefrontal cortex of rhesus macaques. When a monkey decides whether to trust
⚡ The gist in 5 seconds - The gist: an MTS blog post on Habr examines a trust controller for LLM agent memory — a Memory Decision Layer (MDL) designed to reduce hallucinations in RAG systems.
- Where to find it: the piece was published October 3 in the MTS company blog; reading time is 5 minutes.
- Limitation: the controller doesn't guarantee accuracy — 8% of decisions it approves result in hallucinations, and with high risk and relevance, accuracy drops to 55.6%.
🔍 What we found The trigger for the analysis was MyCity, a chatbot launched in 2024 by the New York City mayor's office for small business owners.
It was supposed to be a convenient reference for city rules and laws, but it dispensed advice that could get entrepreneurs fined or sued: for example, not paying tips to employees until the end, or firing someone for reporting harassment.
Its database had correct information sitting right next to outdated material, and the bot couldn't choose between them.
The authors explain the problem with classic RAG.
The retrieval module finds documents by semantic similarity and feeds them to the language model as context, but the model doesn't know whether to trust those documents.