Updated on 28 September 2026 with the new releases (Claude Sonnet 5.5 and Nex N2.5 Pro), a wider role policy for planning and writing, and refreshed figures and explorer data.

I use pi as a coding agent, and for a long time it ran with one default model. My earlier experiment with reasoning levels measured how much performance, cost, and latency move when one setting changes. pi-jev-router makes the model itself a decision per task. Jev classifies the work. A local selector estimates each available model’s quality, cost, and latency, then chooses from their Pareto frontier.

This is the second version of this post. The first, from 20 September, picked the frontier’s knee point because it needs no weights. A week of live use and 72 more benchmark runs showed what that costs. The rule changed.

Classify, then choose the model

Jev is TypeSafe’s first System One model. Its outputs are structured decisions. Choice selects from a fixed list, Score evaluates an ordered rubric, and Noul returns a probability for a yes-or-no question. It does not generate the code or prose that completes the task.

The extension sends Jev a bounded task envelope instead of the full conversation. In one request it asks five things: the work category, the reasoning difficulty, the consequences of failure, whether the brief is sufficient, and whether the work should be decomposed. The first version acted on the category and the risk score. Now all five answers act. Complexity scales the cost weight. A confident “clarify” on the brief stops the router from switching, because a different model does not make an ill-defined task well-defined. A high decompose probability adds a hint to delegate. The category list gained planning and writing, because “write the release notes” used to classify as code.

A bare “continue” no longer reaches Jev. Of the first 151 logged decisions, 90 came back unclear, and most of those were continuations. On such prompts the router keeps the current model and skips the call.

Jev classifies a task, a role policy filters the eligible models, and a value function picks: planning goes to GLM-5.3 Flash, code to ling-3.0-flash or Opus 5.5 when the task is hard, and writing to DeepSeek V4 Flash

Routing happens before a task starts. The extension defaults to shadow mode, which records a recommendation without switching models. Automatic mode applies the recommendation at the task boundary. Every decision goes to a ledger with the pick, the model that ran, and the version of the router that decided. I added that last field after one pi session kept deciding with the 20 September code for six days while the checkout moved on.

Estimating the frontier

The selector first removes models that cannot serve the task. A model is out if it lacks tool support, cannot take the input modalities, has expired, or has too little context for the estimate plus 30 percent headroom. A zero price is out too. It means the price is unpublished. It does not mean the model is free. Planning and code require reasoning support. Higher-risk tasks require three recorded runs of that work kind, and a measured pass rate under 0.65 disqualifies a model. If these gates leave no candidates, they relax in a fixed order, and the recommendation records which one relaxed.

Each remaining model gets a quality score \(q\), an estimated task cost \(c\), and an estimated latency \(t\). The quality prior comes from Artificial Analysis for planning and code, and from EQ-Bench Creative Writing v3 Elo for writing. The first version min–max normalised those scores within the catalogue and used the result as \(q\). That put measured and unmeasured models on different scales, a pass rate for one and a rank for the other. Sonnet at 8 of 8 planning runs scored below its own prior. GPT-6 Sol, at Sonnet’s price with a higher index, was dominated only because it had no runs.

The prior is now mapped onto the range of measured pass rates. With \(r\) the min–max rank of the benchmark index, and \(q_{\mathrm{lo}}\) and \(q_{\mathrm{hi}}\) the lowest and highest pass rates among models with a benchmark entry and at least three runs,

\[ q_{\mathrm{prior}} = q_{\mathrm{lo}} + (q_{\mathrm{hi}} - q_{\mathrm{lo}})\, r . \]

I tried a least-squares fit first. On four noisy pass rates the slope came out negative for code, and every prior collapsed to one value. The linear map keeps the rank order, and order is all a rank contains.

Recorded results then replace the prior. For \(s\) verifier passes in \(n\) runs,

\[ q = \frac{s + 2\,q_{\mathrm{prior}}}{n + 2} . \]

That is a posterior mean with two pseudo-observations at the prior. An untested model sits at its prior. A clean record lifts it. Failures pull it down hard. I would still not read \(q = 0.8\) as an 80 percent chance of completing the next task. A run is scored when the evidence is read, so the turn budget can change without rewriting old result files. More than twelve turns, or a timeout, counts as a failure.

Cost starts with catalogue token prices, the estimated input size, and an allowance of four turns with 1,500 output tokens each. It blends toward the observed mean task cost with the weight \(n/(n+5)\). Latency blends observed means with a shared prior for the work kind. A model is dominated when another feasible model has at least as much estimated quality, no greater cost, and no greater latency, with a strict improvement on at least one axis. Removing the dominated models leaves the Pareto frontier.

Why the knee lost its job

The knee is the frontier point farthest above the chord between the cheapest and the best member, in normalised quality against log cost. It needs no weights. I wrote a replay tool that runs every logged decision through the current selector. The replay showed two problems.

The knee ignores the task. All 24 code decisions in the ledger got the same model at every complexity from 0.2 to 2.5, because the knee reads only the frontier’s shape. Jev’s complexity score was computed and stored and never reached the pick. The knee also depends on which extreme models exist that week. A $0 stealth preview at one end of the chord moved the code route without any change in the evidence.

The pick is now the weighted value function that used to be the fallback:

\[ U_i = q_i - \lambda_k(x)\,\frac{\log_2(c_i/c_{\min})}{10} - \mu_k\,\frac{\log_2(t_i/t_{\min})}{3}, \qquad \lambda_k(x) = \lambda_k\left(1.5 - \frac{x}{3}\right). \]

Here \(\lambda_k\) and \(\mu_k\) are the cost and latency weights for work kind \(k\), and \(x\) is Jev’s complexity score from 0 to 3. The denominators are reference spans: ten doublings from the cheapest to the dearest capable model, three from a fast model to a slow one. A mechanical step is 1.5 times as cost averse as the kind’s default. An ambiguous cross-system task is half as cost averse. A code task rated 2 or harder uses the planning weights. With the code weights, no complexity let a 100-times price step buy a tenth of quality. The knee is still computed and logged as a diagnostic.

The weights are policy, and so is the candidate set. Before the frontier is built, a role policy decides which models may take the work:

Work kindEligible modelsThinking
PlanningArtificial Analysis intelligence index of 44.5 or more; or 40 with three measured planning runs; or, without an index, four measured runs, all passedhigh
Codeevery feasible modelmedium
Writingany model with an EQ-Bench Elo of 1760 or more, or three measured writing runs, plus a Simplified Technical English directivelow

Each threshold sits in a gap of a measured distribution, and the source records the gap next to the constant. The writing Elo line exists because an unrated model got the optimistic prior and won on price alone. gpt-5-nano took every writing task for an afternoon. Until 28 September writing was also limited to OpenAI models, and the planning lines stood at 48.5 and 44.5. That left seven eligible writing models and eight for planning, so I dropped the vendor rule and lowered both planning lines. Writing now has 24 eligible models and planning 14.

Quality against estimated task cost for planning, code, and writing, with eligible models in grey, Pareto frontiers in blue, and the value-function pick in orange; code shows a second pick for a hard task

On the current catalogue, planning goes to z-ai/glm-5.3-flash, easy code to inclusionai/ling-3.0-flash, hard code to Opus 5.5, and writing to deepseek/deepseek-v4-flash-0731. The planning frontier has three members: GLM-5.3 Flash at 5 of 5 for about a cent a task, and GPT-6 Sol and Opus 5.5 at 4 of 4 for 15 to 35 times as much. Under the planning weights the cheap measured model wins. Opus is one weight change away, and whether to make it is the open question of the wider gate.

What changes when the selection rule changes?

An epsilon-constraint is easy to interpret: choose the cheapest model whose quality score exceeds a chosen floor. If that score were a calibrated success probability, the floor could express a reliability requirement. With the current blended score, it is only a threshold on a proxy.

The explorer below compares the router’s value function with the knee, the epsilon-constraint, weighted sums, Chebyshev scalarisations, compromise programming, and TOPSIS. It uses the catalogue and evidence as of 28 September, filtered by the role policy. Value function is the router’s rule. The complexity slider sets \(\lambda\) the way Jev’s score does, and for code it switches to the planning weights at 2. Knee point shows the old rule, with the knee marked by a star. Every candidate within a work kind has the same latency estimate in this snapshot. Hover a point to see the model. The code panel leaves out the 115 eligible models that have no index; they share one fallback prior and told you nothing but their price.

A weighted sum can select only supported efficient points in the coordinates being scalarised. Chebyshev methods can also recover unsupported Pareto points with suitable weights. That distinction matters when a frontier is non-convex. Several methods returning the same model at one setting says little about its global shape, and some agreement is built into the explorer: its utopian-reference achievement scalarising function reduces to the augmented Chebyshev expression.

The knee avoids tuning a quality floor or exchange-rate weights for each work kind. It also cannot express anything about the task. I now maintain two numbers per work kind, placed once and checked by replay.

Seventy-two more runs

The first version measured four cheap models and Sonnet on five tasks that every model passed. Since then the task set has grown. It has code tasks with a specific failure mode, ceiling tasks scored against tests the model never sees, a runbook under strict sentence limits, and two planning tasks whose verifiers check structure: the expand-and-contract order of a column migration, and step order and owners in an incident runbook. I also measured the models the role policy routes to. Opus 5.5, GPT-6 Sol, Gemini 3.8 Flash and GPT-5.6 Luna ran on 26 September for $5.52. Fireworks’ Ember-1 ran on 27 September for $1.08. Claude Sonnet 5.5 and Nex N2.5 Pro ran on 28 September, the day Sonnet 5.5 reached OpenRouter, for $0.33 together.

ModelPlanningHard codeWriting
anthropic/claude-opus-5.54/410/102/2
openai/gpt-6-sol4/410/106/6
fireworks/ember-14/410/102/2
google/gemini-3.8-flash2/42/102/2
openai/gpt-5.6-luna——4/4
anthropic/claude-sonnet-5.54/410/102/2
nex-agi/nex-n2.5-pro1/48/101/2

These counts are scored the way the router reads them. The runner’s own pass flag differs where a run went over the twelve-turn budget or timed out. That covers most of Gemini’s code runs, and three of Nex N2.5 Pro’s: it passed ten of ten hard code tasks by the verifier, at under a cent each, but it is slow and hit the timeout three times.

Mean cost per task for code and planning across seven measured models, from Opus 5.5 at the top to Ling-3.0 Flash at the bottom, with the number of passes for each

The frontier models cost far more than the cheap ones for the same pass count. GPT-6 Sol solves a hard code task for about $0.05 and Opus 5.5 for $0.14. Ling does it for $0.001. The suite still contains no code task that Ling fails and Opus passes, and I have not found one that is also cheap to verify. So the benchmark cannot yet show where the expensive models earn their price.

Twelve of the frontier runs failed on first scoring. Seven of those artifacts were correct plans that the verifier rejected on phrasing. A plan that says “write full_name, first_name, and last_name together” is dual-writing, but the rule did not know that form. I spent three rounds widening the rules. Each round, a reviewer found a paraphrase of a wrong plan that now passed. So I restored the strict verifiers and recorded the correct artifacts as adjudications. The row keeps its pass and the verifier’s real output, with a note on which rule rejected it and why the plan is right. For a benchmark that feeds a router, a false fail costs less than a false pass. Sonnet 5.5’s one rejected plan was the ninth adjudication: a forward reference, “so that release 3 can stop writing the column”, matched the stop-write rule two steps before the real write stop.

What the ledger says

The replay is the before-and-after check. Under the 20 September rule, the ledger’s planning tasks went to z-ai/glm-5.3-flash, code to Ling, and writing to deepseek-v4-flash-0731. Under the rule of 27 September, planning went to Opus 5.5, 21 of 24 code tasks to Ling and three hard ones to Opus 5.5, and writing to GPT-5.6 Luna. Under the current rule, two of those picks are back where the knee had them, by policy this time: with the gates widened on 28 September, GLM-5.3 Flash takes planning through its five measured passes and DeepSeek V4 Flash takes writing through its twelve, and the code picks are unchanged. Of 147 replayed rows, 88 abstain because Jev was confident the brief needed clarification. Most of those were “continue” prompts logged before the continuation rule existed.

These are recommendations. Outside tests, the router has run in shadow mode. The live trial of the new rule is the next step. Whether a $0.14 code task on Opus saves any turns over a $0.001 task on Ling is the question the benchmark cannot yet answer.

The source, selector tests, replay and evaluation harness are on GitHub, with the adjudications in eval/results/ADJUDICATIONS.md. In pi, /router shadow records recommendations without applying them, and /router frontier shows the candidates behind the latest decision.