l2hyunwoo/webgpu-inference: Kotlin GPU Runtime for Android on WebGPU
Published: 2026-09-26 · Author: AI Release · @ai_release1
⚡ The gist in 5 seconds - What it is: An experimental Kotlin runtime, l2hyunwoo/webgpu-inference, has appeared for GPU inference on Android, built on AndroidX WebGPU. It supports static FP32 models, custom WGSL kernels, reusable sessions, and explicit GPU buffer budgets. - Availability: The code is published on GitHub, but the library is not released to Maven Central and is not intended for production — building from source is required. - Key limitation: APIs may change, stability and compatibility are unconfirmed; there is no FP16, no dynamic shapes, and no direct ONNX/TFLite/GGUF loading. ### 🔍 What was found The l2hyunwoo/webgpu-inference repository contains a compact GPU inference runtime for Android, written in Kotlin and using AndroidX WebGPU 1.0.0-alpha05. The feature set includes static FP32 tensors, model creation in Kotlin or via an exported graph (in the latter case — one input and one output), multiple named inputs/outputs in Kotlin, registration of custom WGSL kernels, session and intermediate buffer reuse, and control over the maximum GPU buffer size (256 MB in the example). The WebGPU Denoise Lab demo app performs tiled photo denoising with NAFNet: the user selects a JPEG/PNG, runs processing, compares the result with a slider, and saves a PNG. The demo runs with a 25% inference duty cycle. In a benchmark on an SM-S948N over 100 NAFNet FP32 inferences at 256×256 resolution, the median time was 266.0 ms, p95 was 506.2 ms, and initialization took 2.29 s. A debug APK and a patched WebGPU alpha05 were used; all 100 results passed verification against a PyTorch reference. The measurement includes input loading, command encoding, GPU execution, and result readback, but excludes photo decoding, tiling, pauses, and PNG export — it is not the full image processing time. It is also emphasized that this is data from a single device and session, not a comparison with other runtimes. ### 💡 Why it matters The project gives Android developers a working example of integrating WebGPU for local neural network inference. Instead of being tied to specific libraries or system APIs, it offers a unified Kotlin interface: model definition, weight loading, session parameter configuration, and execution with results returned as FloatArray. This simplifies experimenting with GPU compute in mobile apps and provides a benchmark for evaluating performance (266 ms median on 256×256 NAFNet in a debug build). The availability of custom WGSL kernels and buffer control offers flexibility for optimization. ### 🧩 Context The repository was created for experimentation, so the library is not published to external repositories, and pretrained weights are not included — they must be prepared separately. The library code is distributed under MIT, but separate terms apply to model sources and weights. The current version does not support FP16 or dynamic tensor shapes, provides no public APIs for GPU-resident data, and does not directly load ONNX, TFLite, or
⚡ The gist in 5 seconds - What it is: An experimental Kotlin runtime, l2hyunwoo/webgpu-inference, has appeared for GPU inference on Android, built on AndroidX WebGPU.
It supports static FP32 models, custom WGSL kernels, reusable sessions, and explicit GPU buffer budgets.
- Availability: The code is published on GitHub, but the library is not released to Maven Central and is not intended for production — building from source is required.
- Key limitation: APIs may change, stability and compatibility are unconfirmed; there is no FP16, no dynamic shapes, and no direct ONNX/TFLite/GGUF loading.
🔍 What was found The l2hyunwoo/webgpu-inference repository contains a compact GPU inference runtime for Android, written in Kotlin and using AndroidX WebGPU 1.0.0-alpha05.
The feature set includes static FP32 tensors, model creation in Kotlin or via an exported graph (in the latter case — one input and one output), multiple named inputs/outputs in Kotlin, registration of custom WGSL kernels, session and intermediate buffer reuse, and control over the maximum GPU buffer size (256 MB in the example).
The WebGPU Denoise Lab demo app performs tiled photo denoising with NAFNet: the user selects a JPEG/PNG, runs processing, compares the result with a slider, and saves a PNG.