Jeff: a 0.8B model trained at home for fast decisions in tens of milliseconds
Published: 2026-09-29 · Author: AI Release · @ai_release1
⚡ The gist in 5 seconds - Jeff is an open 0.8B "System 1" model for fast decisions; on an NVIDIA GPU it responds in tens of milliseconds (around ~30 ms). - Version v1.3 is available: a new adapter-first base, 15 adapters, GGUF files for llama.cpp, and a Jeff-Code build; weights are on Hugging Face, docs at jeffhub.ai. - Limitation: v1.3 adapters only work with the v1.3 base; some adapters (sanctions, soc) are distributed under CC BY-NC 4.0, while aml and trading-desk have their own data terms. ### 🔍 What was found Jeff is an open 0.8B "System 1" model that, per the announcement, was trained at home and is claimed to be Jev-compatible. Its key idea is to sit in front of a large model like Qwen 3.8-27B: Jeff makes fast decisions, and the 27B model is only invoked when Jeff is unsure. Across eight adapters with identical test strings, average accuracy rises from 86.6% for Qwen alone to 94.6% for the Jeff+adapter combo and 95.0% with fallback to Qwen. Decision time on an Apple M4 Max drops from 8.1 s to 0.25 s (35× faster) and to 0.46 s (28× faster), respectively. Memory-wise: Qwen takes 28.6 GB, adding Jeff with all 15 adapters costs +2.09 GB, with three adapters +1.83 GB; the base itself takes 1.74 GB (measured on an RTX PRO 6000). A separate Jeff-Code build is a coding agent based on Pi with two adapters for Qwen 3.8-27B. On 1,242 paired tasks from six benchmarks, pass rates are nearly identical: 62.4% vs 62.8% (paired difference −0.2 points, 95% interval from −2.6 to +2.1). Meanwhile, each task completes on average 47% faster — 32% less time; the mean ratio is 0.68× of the baseline build, median 0.70×. On SWE-bench Verified it's 0.63×, SWE-rebench 0.66×, Terminal-Bench Pro 0.64×, Harbor Index 0.71×; no notable speedup on Terminal-Bench 2.0 (0.96×) or SkillsBench (0.91×). Total time across all tasks drops by 14% (0.86×). Simply disabling Qwen's "thinking" costs 7.6 points of quality, so it's Jeff that decides when the large model needs to think. ### 💡 Why it matters The practical benefit is saving time and memory while maintaining accuracy in low-latency scenarios. For example, the guard task (protection against prompt injections and jailbreaks): Qwen 3.8-27B gets 84.0% accuracy in 3.9 s, while Jeff with an adapter gets 98.7% in 0.07 s. A similar picture holds for triage, tool selection, spam, and navigation: Jeff consistently delivers 88–99% in 0.07–0.58 s. This makes it possible to run a small model locally in front of agentic scenarios, filtering, and routing, without keeping a heavy model on every step. ### 🧩 Context Jeff v1.3 is a new adapter-first base release: 15 adapters, each trained for one epoch, with accuracy and calibration error reported for each on held-out tests. For example, the sanctions adapter hits 99.96% accuracy, guard 98.2%, legal-clauses 83.6% across 100 clause types. The methodology is fixed in the description: Apple M4 Max 128 GB, both models on MLX, Qwen 3.8-27B in 8-bit, samples of 300 strings per task (500 for emotion and legal-clauses). v1.2 adapters remain available under their old names. According to Hacker News, the post p
⚡ The gist in 5 seconds - Jeff is an open 0.8B "System 1" model for fast decisions; on an NVIDIA GPU it responds in tens of milliseconds (around ~30 ms).
- Version v1.3 is available: a new adapter-first base, 15 adapters, GGUF files for llama.cpp, and a Jeff-Code build; weights are on Hugging Face, docs at jeffhub.ai.
- Limitation: v1.3 adapters only work with the v1.3 base; some adapters (sanctions, soc) are distributed under CC BY-NC 4.0, while aml and trading-desk have their own data terms.
🔍 What was found Jeff is an open 0.8B "System 1" model that, per the announcement, was trained at home and is claimed to be Jev-compatible.
Its key idea is to sit in front of a large model like Qwen 3.8-27B: Jeff makes fast decisions, and the 27B model is only invoked when Jeff is unsure.
Across eight adapters with identical test strings, average accuracy rises from 86.6% for Qwen alone to 94.6% for the Jeff+adapter combo and 95.0% with fallback to Qwen.
Decision time on an Apple M4 Max drops from 8.1 s to 0.25 s (35× faster) and to 0.46 s (28× faster), respectively.
Memory-wise: Qwen takes 28.6 GB, adding Jeff with all 15 adapters costs +2.09 GB, with three adapters +1.83 GB; the base itself takes 1.74 GB (measured on an RTX PRO 6000).