LibLayaX: fast local AI decisions via a C API inside your application
Published: 2026-10-03 · Author: AI Release · @ai_release1
⚡ The gist in 5 seconds - The point: LibLayaX version 1.0.15 runs the Laya model directly in your application's process as a regular library: an AI decision becomes a function call, with no server or cloud. - Where it's available: Windows x64/ARM64, Linux x86-64/ARM64, macOS on Apple Silicon and Intel. CPU works everywhere; GPU works via Vulkan on Windows, Linux, and Apple Silicon. - The limitation: the library doesn't include the model itself: the weights must be downloaded separately from Hugging Face — about 800 MB for the English version. The project is unofficial and not endorsed by the authors of Laya or laya.cpp. ### 🔍 What was found The DaragonTech/LibLayaX project is an unofficial C API library built on the laya.cpp engine. It runs Laya — a transformer neural network for decision-making (ModernBERT-large for English, mmBERT for 100+ languages) — which returns an answer with a probability in a single pass over the text. Laya doesn't generate text like chat models; it immediately outputs a label or a number: "is this a refund request — yes/no," "which intent is this," "how angry is the customer on a scale." A single pass instead of token generation is what provides the main speed advantage. In the laya-bench benchmark with the English model and batches of 16 questions, an NVIDIA RTX 5080 Laptop GPU on Vulkan fp16 delivers about 670 questions per second (1.5 ms per question), fp32 — about 230. An Intel Core Ultra 9 275HX on CPU — about 21 questions per second; an Apple M3 Ultra via Vulkan on Metal in bf16 — 170, on CPU — 14. A single question on the M3 Ultra takes 29 ms on GPU and 76 ms on CPU. Model loading takes about a second, and on CPU it consumes about 1.7 GB of memory. In half precision, probabilities deviate from CPU results by roughly 0.003. ### 💡 Why it matters This approach turns AI into a tool that lives inside the software rather than in a chat window: routing a support ticket, checking a message, deciding whether a document needs review. It's fast enough to sit in the request path or process a backlog of thousands of items, and the answer comes as a ready-made label or number, not as text to parse. Ten simple C functions with JSON input and output can be called from C, C++, Delphi, Rust, Lua, C#, Python, Go, Java, and any other language capable of loading a shared library. A single set of functions and identical JSON requests across all platforms means the integration is written once and ported to the other targets. The text never leaves the machine, everything works offline, there are no per-call fees, and no separate process to deploy. ### 🧩 Context Existing Laya implementations aren't designed for embedding into third-party applications: the original Python library requires Python and PyTorch, while laya.cpp — a native C++ port on ggml — is a command-line utility and HTTP server. An application would have to launch, monitor, and maintain a separate process. LibLayaX closes this gap by turning the laya.cpp engine into a single library file. The model is not included, though: the weights are published by the Laya project on Hugging Face in the convaiinnovations/laya repository, and five files are required, inc
⚡ The gist in 5 seconds - The point: LibLayaX version 1.0.15 runs the Laya model directly in your application's process as a regular library: an AI decision becomes a function call, with no server or cloud.
- Where it's available: Windows x64/ARM64, Linux x86-64/ARM64, macOS on Apple Silicon and Intel.
CPU works everywhere; GPU works via Vulkan on Windows, Linux, and Apple Silicon.
- The limitation: the library doesn't include the model itself: the weights must be downloaded separately from Hugging Face — about 800 MB for the English version.
The project is unofficial and not endorsed by the authors of Laya or laya.cpp.
🔍 What was found The DaragonTech/LibLayaX project is an unofficial C API library built on the laya.cpp engine.
It runs Laya — a transformer neural network for decision-making (ModernBERT-large for English, mmBERT for 100+ languages) — which returns an answer with a probability in a single pass over the text.
Laya doesn't generate text like chat models; it immediately outputs a label or a number: "is this a refund request — yes/no," "which intent is this," "how angry is the customer on a scale." A single pass instead of token generation is what provides the main speed advantage.