Unbound: small uncensored models that run on your phone and in your browser
Unbound is a family of small, uncensored language models built to run on the device in front of you, plus the iOS app and browser client that run them. Everything stays local: the model downloads once, and chat happens without an API.
The models
There are three released checkpoints on Hugging Face:
- Unbound E2B (
evalengine/unbound-e2b), a finetune ofgoogle/gemma-4-E2B-it. About 3.4 GB at Q4_K_M, sized for phones. - Unbound E4B (
evalengine/unbound-e4b), the larger Gemma 4 sibling. About 5.3 GB at Q4_K_M, better suited to laptops, tablets and flagship phones. - Unbound Q-0.8B (
evalengine/unbound-q-0.8b), a text-only finetune of Qwen3.5-0.8B at about 530 MB.
The recipe combines LoRA SFT with Unsloth and TRL, and weight-orthogonalisation abliteration with Heretic. Compliance training data was distilled from the AEON teacher model. For Q-0.8B the data was audited row by row and 48 rows with major fabrications were removed before training.
How we scored it
Every experiment had to move three axes the right way: refusal rate on AdvBench (520 prompts, LLM judge, built to catch "sneaky" refusals), capability on lm-evaluation-harness tasks against the track's own base model, and KL divergence from base. Experiments ran as an agent-driven loop: change a config, train, evaluate, keep or revert, and log the result.
Results from the model cards:
Model | Refusal (base to Unbound) | BBH macro (base to Unbound) |
|---|---|---|
E2B | 98.46% to 4.42% | 41.07% to 39.97% |
E4B | 98.08% to 2.69% | 54.26% to 53.45% |
Q-0.8B | 90.58% to 5.00% | 38.21% to 41.29% |
Hallucination on harmful prompts went up for all three (15.96% for E2B, 13.08% for E4B, 35.77% for Q-0.8B), and the cards say so. For the 0.8B model, a decontamination experiment indicated that this ceiling comes from model size, not from the teacher data.
The app and the web client
The mobile app is built on Expo and React Native with llama.rn, a llama.cpp binding that uses Metal on iOS. It downloads the E2B Q4_K_M GGUF once, streams tokens into a markdown chat UI, and supports cancelling mid-generation. It is on the App Store.
unbound.evalengine.ai runs inference entirely in the browser with wllama (llama.cpp compiled to WebAssembly), caching model files in OPFS. Multi-threaded inference needs SharedArrayBuffer, so the site sends cross-origin isolation headers. Without them, the heap drops from about 4 GB to about 2 GB and larger models fail to load. The web client also supports custom GGUF URLs, and generates images locally over WebGPU.
Try it
1ollama pull evalengine/unbound-e2b2ollama run evalengine/unbound-e2b
These models have reduced safety filtering and can produce harmful, false or biased output. They are provided as-is.
