Store Assistant on a Local 9B Model: A Case Study Breakdown
Published: 2026-09-30 · Author: AI Release · @ai_release1
⚡ The Gist in 5 Seconds - Author nmrcs built a gift shop assistant on a local qwen3.5 9B model: it clarifies the customer's request and suggests products, while code handles the totals, discounts, and cart. - The source code is open in the nmrcs/cart-agent repository on npm workspaces, alongside a mock store you can download and run locally. - Streaming the response text is nearly pointless: the model takes 2 seconds out of a 9–15 second total wait to write the answer, so the UI shows processing stages instead. ### 🔍 What Was Found The case study on Habr describes a dialogue: a customer writes "a gift for my brother, he's 30, loves board games, under $70." The assistant asks when the birthday is and learns there are 5 business days left until Friday. Per the store's rules, such an order gets a 20% discount, so three gifts totaling $83.99 fit into the $70 budget, coming out to $67.19 after the discount. The dialogue is handled by the qwen3.5 9B model, running in LM Studio in 4-bit format (about 6 GB) on a MacBook M3 Max. Instead of LM Studio, you can plug in any OpenAI-compatible API — the address, model, and key are set in.env. The architecture is a monorepo: apps/backend on NestJS 11, Prisma 7, and PostgreSQL; apps/harness — the assistant itself with no database of its own; apps/frontend on React 19, Vite, and HeroUI; packages/contracts — zod schemas; bench — benchmarks. The catalog has 64 products, 3 are out of stock, and 8 require buying batteries or wrapping paper separately. Each message triggers 8 steps, of which the model handles only 3: parsing the message, selecting products, and writing the reply. Medians across 6 conversations: parsing ~3.3 s, selection ~4.3 s, reply ~2 s, while the SQL filter takes ~22 ms and price calculation ~9 ms. An answer to a question arrives in 3–7 seconds, a product suggestion in 9–15 seconds. Across six customers, the assistant never exceeded the budget and never quoted an incorrect amount. ### 💡 Why It Matters The case demonstrates a working scheme for separating responsibilities: the model handles only dialogue and selection, while verifiable operations are done by code. The assistant communicates with the store via two requests — product search with filters for age, price, and delivery time, and order total calculation — so it can be connected to a real online store through an API. Customer-facing questions live in a config file, so adapting it for another store only requires changing that file, not the code. The model's reply is verified before sending: the code finds all dollar amounts, cross-checks them against backend data, and if a made-up figure appears, replaces the reply with a template. ### 🧩 Context The first version of the assistant made mistakes with dates: if you wrote "by Monday" on a Saturday, the model would set the Monday a week later. So now the model doesn't calculate the date — it only extracts the word "Monday" from the message, while the code computes the date and the number of business days. During product selection, the model picks up to 5 top products from a list already filtered by an SQL query, following a JSON schema, and the code fits them into the budget, skipping anything that doesn't fit. A separate fitsRequest flag helps respond correctly when no suitable product exists: the model honestly marks it as unsuitable. First, the assistant offers one gift and two backup options, and if seve
⚡ The Gist in 5 Seconds - Author nmrcs built a gift shop assistant on a local qwen3.5 9B model: it clarifies the customer's request and suggests products, while code handles the totals, discounts, and cart.
- The source code is open in the nmrcs/cart-agent repository on npm workspaces, alongside a mock store you can download and run locally.
- Streaming the response text is nearly pointless: the model takes 2 seconds out of a 9–15 second total wait to write the answer, so the UI shows processing stages instead.
🔍 What Was Found The case study on Habr describes a dialogue: a customer writes "a gift for my brother, he's 30, loves board games, under $70." The assistant asks when the birthday is and learns there are 5 business days left until Friday.
Per the store's rules, such an order gets a 20% discount, so three gifts totaling $83.99 fit into the $70 budget, coming out to $67.19 after the discount.
The dialogue is handled by the qwen3.5 9B model, running in LM Studio in 4-bit format (about 6 GB) on a MacBook M3 Max.
Instead of LM Studio, you can plug in any OpenAI-compatible API — the address, model, and key are set in.env.
The architecture is a monorepo: apps/backend on NestJS 11, Prisma 7, and PostgreSQL; apps/harness — the assistant itself with no database of its own; apps/frontend on React 19, Vite, and HeroUI; packages/contracts — zod schemas; bench — benchmarks.