EVAL Engine

Decision-4B and Decision-0.8B: small models that choose instead of chat

Decision-4B and Decision-0.8B are open-weight models that pick an option. You give them a state, a question and a list of options, and they return one option and a probability for each. They don't generate text. The answer comes from a single forward pass.

Interface

Input is JSON with a state, a question and between 2 and 24 options. Each option has a letter label, a semantic key and a description. Yes/no questions and rubric scores are written as options too. The system prompt tells the model to treat the state as data rather than instructions, and to return only a letter.

1{
2 "state": "Customer message: My card was charged twice for the same subscription.",
3 "question": "Which listed support intent best matches this message?",
4 "options": [
5 {"label": "A", "key": "duplicate_charge", "description": "Charged more than once."},
6 {"label": "B", "key": "cancel_subscription", "description": "Wants to end a subscription."}
7 ]
8}

The chosen option is the letter with the highest logit at the last position. The probabilities are a softmax over the listed letters only.

Training

Both are rank-8 LoRA adapters trained for one epoch on 74,308 examples from twelve public sources. Loss is computed only on the answer letter and EOS. Decision-4B is based on Qwen3.5-4B (65 MB adapter, 2.7 GB Q4_K_M GGUF). Decision-0.8B is based on Qwen3.5-0.8B (21.7 MB adapter, F16 and Q8_0 GGUFs). Both were trained on an RTX PRO 6000 Blackwell.

Results

On a 2,800-case held-out panel covering nine task families (five public datasets, CLINC150, GoEmotions, PAWS, VitaminC and HelpSteer2, plus four synthetic rule workflows):

Model

Family mean

Accuracy

Decision-4B

76.4%

79.1%

Decision-0.8B

62.6%

68.8%

Original Qwen3.5-4B

51.9%

58.3%

Original Qwen3.5-0.8B

39.8%

43.9%

Decision-4B scores 84.1% on the language tasks but 66.8% on the rule workflows (quorum, veto, exception, fallback). On a separate 500-case holdout, Decision-0.8B scored 79.6%, against 56.2% for its base model.

The model cards are clear about caveats. These panels come from the same public sources used in training, so they measure held-out examples from familiar distributions. The models are English only, have a 2,048-token context, and are weaker on response-quality grading and multi-rule policies. The 0.8B adapter is not calibrated. Neither model is meant for unattended high-stakes decisions.

A small demo: Snake

The Unbound app and unbound.evalengine.ai include a Snake game played by a decision model on the device. Code filters out moves that would collide on the next tick, and the model only picks when more than one safe move remains. In five seeded games, Decision-4B collected 122 food by summing probability per direction. Decision-0.8B did better by taking its top answer among safe moves (53 food, up from 32). Measured in Chrome on WebGPU, Decision-0.8B takes 372 ms per decision.

Get it

evalengine/decision-4b and evalengine/decision-0.8b are on Hugging Face under Apache 2.0, with GGUF builds in the matching -gguf repos.