← Back to ArticlesCorrect selection or refusal on 24 frozen cases
Relative response speed (Luna = 100)
Equivalent 21-question requests per $1
Discovery
Copy article
Exploring Decision Models with a Lightweight Benchmark
A lightweight decision model benchmark comparing Jev, GPT-6 Luna, Liquid d1 and more on accuracy, latency, cost and usable application decisions.
Developed by Robert E. Beckner III (Merlin) | rbeckner.com
A decision model reads a situation and chooses among options supplied by an application. I asked Codex to run a lightweight, exploratory benchmark using our actual 21-question workload and 24 frozen synthetic messages.
Jev won this comparison for our tested application. It matched the best native primary score, 24/24, returned complete responses at 243 ms median and 324 ms p95, and admitted 12/12 eligible requests under the application's existing rules, with no negative admitted.
GPT-6 Luna Decisions won on speed: 181 ms median and 307 ms p95, with the same 24/24 primary score. Its median response was 25.4% quicker than Jev's, but its reported charge was 1.99× Jev's and the existing rules admitted 10/12 eligible requests. Correct selections and usable decisions are different measurements.
Liquid d1 won on provider charge: approximately $0.304 per 1K requests, versus Jev's $0.480, a 36.7% reduction on this recorded workload. Its median response took nearly twice as long, and it scored 23/24 on the primary task.
Perplexity Decider V1 offered a cheaper perfect primary score: 24/24 at a 276 ms median, with about 9.1% lower reported charge than Jev. The existing probability rules admitted 9/12 eligible requests, so it still needs a separate calibration experiment.
My recommendation is to keep Jev serving this application. That is a decision from this experiment, not a claim that Jev wins every workload.
These findings describe a small, agent-labelled sample at the application's existing operating point. Independent annotation and production burst testing remain further qualification steps.
How we benchmarked 6 decision models on 21 questions
We compared Jev 1.13, Liquid d1, Clef Flash, Kev 4B, Perplexity Decider V1 27B and GPT-6 Luna Decisions through their native decision endpoints.
Each main request carried 21 typed questions, including the kind of request, whose records hold the answer, the function to use, and slots and dates. The function question offered 56 options. Every model received the same state, options and instructions. These native decision endpoints return typed answers and probability distributions rather than generated prose. The workload tests routing and classification: which function to call, whose records to read and when to escalate.
We froze 24 new messages before observing outputs: 12 requests for existing records and 12 that should be escalated. Positives covered cash, outstanding invoices, tasks and calendar records, including Spanish. Negatives covered another company's own records, actions, analysis, combined reads and small talk. Each label had a written rationale. Codex authored and reviewed these cases; they had no independent annotator.
The primary score measures correct function selection and semantic refusal: a positive needs the right function and an agreeing request/ownership read; a negative must decline that combination. All 21 answers must name allowed options. We did not independently score every slot, execute record lookups or judge delivered answers.
Calls ran serially from New York City through the production gateway. The initial 4 native candidates were interleaved in rotating model order. At my request, Decider and then Luna followed as separate 26-call arms on the exact frozen cases, without rerunning the baselines. Gateway caching was disabled, and every recorded call made 1 physical attempt. This was not a burst-capacity test. The later arms also leave time-of-run and provider-load variation uncontrolled.
Accuracy: Luna, Jev and Decider tied at the top
Chart data
| Correct selection or refusal | |
|---|---|
| Luna | 24 |
| Jev | 24 |
| Decider | 24 |
| d1 | 23 |
| Kev | 22 |
| Clef | 21 |
All 6 native models completed the cohort. The chart runs from most to fewest correct selections or refusals; response time breaks the tie. Chart labels shorten Liquid d1 to d1, Clef Flash to Clef, Kev 4B to Kev and GPT-6 Luna Decisions to Luna. These counts measure the primary decision, while the admission rules below determine what the application can use.
Latency: Luna had the fastest median and p95
Chart data
| Relative response speed | |
|---|---|
| Luna | 100 |
| Jev | 74.65 |
| Decider | 65.62 |
| d1 | 38.36 |
| Clef | 19.02 |
| Kev | 9.15 |
Taller bars mean faster responses. The speed index is 100 × Luna's median divided by each model's median, so Luna scores 100 and a model taking twice as long scores 50. It is a comparison of response times, not measured requests per second. Timings include network and gateway overhead. Percentiles use nearest-rank observations; 24 calls provide exploratory timings, not a production tail guarantee.
| Native model | Median | p95 | Reported charge per 1K requests |
|---|---|---|---|
| Luna Decisions | 181 ms | 307 ms | $0.956 |
| Jev | 243 ms | 324 ms | $0.480 |
| Decider V1 | 276 ms | 429 ms | $0.437 |
| Liquid d1 | 472 ms | 962 ms | $0.304 |
| Clef Flash | 952 ms | 1,168 ms | $1.158 |
| Kev 4B | 1,980 ms | 2,111 ms | $0.339 |
Cost: Liquid d1 was the least expensive
Chart data
| Requests per dollar | |
|---|---|
| d1 | 3288 |
| Kev | 2948 |
| Decider | 2290 |
| Jev | 2082 |
| Luna | 1046 |
| Clef | 863 |
Taller bars mean more equivalent requests per dollar. We divide $1 by each model's average reported charge for the 24-case cohort, then round to a whole request. This is a cost-efficiency projection for the same workload, not a throughput measurement or future bill. d1's reduction is meaningful for repeated decisions, provided the extra latency and classification difference are acceptable for the application. Kev's low charge did not compensate for its slow responses and low admission under the existing rules.
Each request contains 21 questions. The table's 1K column and the chart both scale the recorded charges. OpenRouter reported the charges directly. Providers counted the same payload differently: Jev reported 274,501 input tokens across the cohort, versus d1's 182,505. A price per million tokens alone therefore does not predict the cost of the same application request.
The admission rules determine what the application can use
Our existing reader combines probability mass, agreement between questions and independent vetoes before admitting a request. At its unchanged operating point, Jev and d1 admitted 12/12 positives; Luna admitted 10/12, Decider 9/12, Clef Flash 3/12, and Kev 0/12. None admitted a negative.
This is a compatibility diagnostic. The cut was established for Jev; it has not been calibrated for the other models. Lower admission does not prove that a model is incapable, and lowering the cut on this already-seen sample would not establish safety. A new model needs its own development calibration and an untouched evaluation set.
The result matters to an application that escalates uncertain decisions to a larger model: inexpensive classification saves little when its correct selections cannot be admitted. Coverage, wrong admissions and the cost of escalation belong in the comparison together.
Drex selected correctly on 21 of 23 matched cases
We also ran Drex DLM, an approximately 8.19B-parameter decision model, as Q8_0 weights on Apple Silicon through the publisher's custom Metal runtime. It received the same 21 questions and frozen messages. This is a separate accuracy arm; its hardware timings and zero provider invoice do not enter the API speed or cost charts.
One main request was stopped by the resource guard before an answer and remains unscored. The table compares the 23 completed inputs with those exact inputs from the API cohorts. No case was replayed.
| Native decision model | Correct / 23 | Eligible requests admitted / 11 |
|---|---|---|
| Jev | 23 | 11 |
| Luna Decisions | 23 | 9 |
| Decider V1 | 23 | 8 |
| Liquid d1 | 22 | 11 |
| Drex Q8_0 | 21 | 4 |
| Kev 4B | 21 | 0 |
| Clef Flash | 20 | 3 |
Drex's two errors chose one function for a request asking for two different reads. The application needed to escalate the whole request. Its admission rules blocked both errors, and no model admitted a negative on this matched subset. That protection came from the application, not a perfect primary classifier.
Drex passed both short context probes. This does not establish long-context correctness or parity with unquantized weights. Its weights carry CC BY-NC 4.0; this was noncommercial research. The result adds a useful comparison, but does not justify replacing Jev. Our earlier Apple Silicon serving tests cover the separate question of running local inference efficiently.
Keep Jev; qualify Luna and Decider separately
Jev delivered the best combination for the tested application: tied native primary correctness, quick responses and full positive admission at the existing operating point. Luna won speed; d1 won charge. Luna and Decider both matched the primary score, but each needs separate calibration before it can replace Jev. Their tradeoffs differ: Luna was quicker and more expensive; Decider was cheaper and slower. Clef Flash and Kev did not improve this workload as configured.
The next comparison needs independently reviewed labels, candidate-specific calibration developed separately from an untouched evaluation set, complete slot/date checks, and realistic bursts through the record executor and delivered answer. No replacement was promoted from this small study.
One context limit also remains unresolved. OpenRouter's Clef Flash listing warns of text-state truncation around 2K tokens. Both of our roughly 2.8K–3.1K-token evidence-position probes succeeded on all 6 native models. That limited result did not reproduce the boundary and does not establish long-context safety.
Key attributions
By me. The experiment direction, the candidate comparison, and the insistence that the test match the application's decisions and name its winner.
By AI. Codex built the provider comparison, froze the cases, preserved usage and responses, and derived the charts.
Arrived at together. A decision model earns its place through correct admitted work, complete response time and measured cost on the same application request.
#AI#benchmarking#performance