We wanted background agents to leave us an account with a full tank. Spreading every task across every subscription was the wrong objective: it could leave plenty of total allowance, scattered across accounts that were all partly used.
We also wanted Jev to help choose a suitable model and computer. The experiment changed our view of where it could help. Account concentration improved the allocation we simulated. Jev did not establish that one model would do better work than another. Keeping those findings separate made the router simpler and the claims more useful.
Start with the decision you actually have
We froze 300 native Claude seat lifecycle records: 156 attempts in 115 dispatch episodes across 71 work items. We also selected descriptions for 24 of those items. That described cohort accounts for 71 attempts and 43 episodes; it is not established as a complete or representative sample. A native Opus reviewer audited the source and froze a chronological split of 13 development tasks and 11 holdout tasks before candidate scoring.
The surprise: 22 of the 24 descriptions already named their permitted models. The other two belonged to existing lead tasks. Some pinned an exact model; some allowed a small set; others required an original session to continue. A router must preserve those constraints. Asking a small model to choose again would add a call without creating a useful decision.
The records also lacked the information needed for a model-quality benchmark. They did not record each attempt's actual model, tokens, accepted output or account allowance consumed. A successful process exit is not proof that a task passed its acceptance checks. The quota snapshot came after the historical attempts, with a median lag of about 21 hours. We could audit the history and test allocation scenarios. We could not reconstruct what every routing decision should have been.
The split was not balanced either: development contained a limit storm, while holdout contained no recorded limits. Four development tasks also had attempts continuing past the holdout boundary. We retained these limitations instead of presenting the two groups as interchangeable production trials.
Concentrate accounts after checking whether the work fits
Our starting selector preferred the account with the most usable allowance. The candidate prefers a partly used account with enough allowance for the next checkpoint, including an uncertainty allowance. It reuses an active account when the checkpoint still fits. Another account becomes available as work or concurrency requires it; three active accounts is an example, not a hard-coded target.
The order matters:
- Preserve the task's owner, model restrictions, original session and any hold.
- Admit only the permitted provider route and a computer with fresh, sufficient resources.
- Check every applicable shared and model-specific quota window.
- Reserve estimated checkpoint demand against the shared account pool.
- Concentrate fitting work, then choose a suitable computer for that account.
Weekly allowance and a short usage window answer different questions. A nearly exhausted short window can prevent a task from fitting, even when the weekly tank is mostly full. Conversely, choosing the smallest short-window balance first can drain the wrong weekly account. The candidate uses the longer shared window for concentration and checks the short window for fit.
Ten percent remaining is a signal to prepare a successor, not a rule that discards the last ten percent. A small checkpoint may still fit. A large checkpoint should move sooner. The planner returns a successor-preparation signal; it does not launch a successor by itself. Account aliases on different computers must share one reservation balance.
What the allocation experiment showed
Nineteen descriptions expressly permitted Opus: all 13 development tasks and 6 of the 11 holdout tasks. We use only their count and order, treating them as identical checkpoints. Nothing is fitted, so that split does not provide an independent model-validation result here.
We used four of six accounts in one frozen snapshot. One excluded account was exhausted; the other had stale, unknown health. The snapshot itself reported an overall unknown state. Our hypothetical allocation assumes the retained routes are otherwise admissible; it is not a live admission decision. Initial weekly balances were 41%, 92%, 41% and 44%; their short-window balances were 100%, 54%, 100% and 100%.
We tried three assumed checkpoint sizes, each with an additional one percentage point of uncertainty. These are assumptions, not measurements of what those tasks consumed. The experiment reserves against a frozen meter. Machine capacity is deliberately nonbinding; it does not simulate execution, queueing, resets or other teams' work. That removes a material risk: concentrating on one account could concentrate work on the few computers where that account can run.
The reset horizons differ too. The preserved account refills about 158 hours after the snapshot; the two accounts emptied most heavily refill in about 12 and 36 hours. A longer-horizon comparison must model those resets and future demand. These results describe one instant.
Each row lists the starting policy followed by the candidate, both from the identical initial snapshot. These are two policy outcomes, not a temporal before-and-after. All 19 reservations fit at the three published demand assumptions; that does not demonstrate equal throughput or acceptance for other workloads.
| Demand + buffer | Accounts touched | Fullest weekly balance |
|---|---|---|
| 2 + 1 points | 4 / 2 | 68% / 92% |
| 4 + 1 points | 4 / 3 | 57% / 92% |
| 6 + 1 points | 4 / 4 | 50% / 71% |
At the middle assumption, concentration left the fullest account untouched at 92%, while the starting policy reduced its balance to 57%. All 19 allocations fit under both policies. At the largest assumption, every account was needed. Concentration preserved a fuller best remaining account, but it could not create capacity.
Weekly balance is not immediately usable headroom. Under the same equal-per-window demand assumption, the best remaining allowance across the binding windows is 32% versus 54%, 24% versus 54%, and 13% versus 33% in the three rows. The preserved 92% weekly account still has a 54% short-window limit.
Both policies consume exactly the same total allowance in every scenario. Concentration redistributes it. At the middle assumption, it leaves two accounts at 1%; the starting policy leaves its lowest account at 21%. That is the deliberate trade for preserving a fuller account, and the experiment does not prove that trade suits every workload.
These figures show an allocation property under stated assumptions. They are not completed tasks, increased throughput, longer subscription life or cash savings. The same distinction applies to the separately reported development and holdout scenarios in the downloadable results.
Give Jev a question it can answer
We used the installed Jev client with typed questions, a bounded budget and a durable cache. TypeSafe documents questions as independent judgments over the same supplied state. That is useful for a small decision loop, provided code handles the dependencies between decisions. See the typed primitives and parallel questions.
Only five frozen cases allowed both Opus and Grok: four in development and one in holdout. We sent sanitized task descriptions and explicitly stated that no measured comparative quality, rework or task-duration evidence was available. Account identities, machine names and quota figures were absent from the model input. Availability was assumed for this offline question test; it authorized no conductor launch.
The first payload asked for a model choice, a complexity score and evidence sufficiency. Three of four development responses failed the installed client's strict response validation. The receipts did not retain the failed validation reason, so we cannot identify the cause from that evidence.
We removed the unused complexity question. The narrower payload asked which model the evidence favored, with an explicit “insufficient evidence” option, and whether the evidence supported a comparison. All four development responses and the one holdout response then passed validation. All five chose insufficient evidence.
That was a useful boundary, not a model-selection victory. A confidence value of one for abstaining does not prove which coding model would perform best. We kept model advice optional. A pinned model, a single admissible model, missing evidence or an invalid response falls back to the deterministic path.
Errors worth keeping in the record
The historical audit found a failure that this simulation does not address: limits were correlated across accounts. In one two-hour window every account in use hit a provider limit across the machines in use. Fourteen of fifteen failover cascades ended in a limit on every account tried. A different account is not proof of available capacity; suppressing futile cascades requires its own native admission and recovery checks.
The implementation and tests exposed several ways a plausible router can go wrong. We separated billing, identity and model-availability evidence from quota timestamps: a fresh usage meter does not prove the other three. We stopped treating a calendar-date reset estimate as an exact native reset. We used the lowest observed shared balance across account aliases and rejected conflicting observations.
Completed inference can appear in the task ledger before its usage appears in the provider meter. The candidate retains its debit until a later native observation arrives. An unknown transport outcome stays reserved. A caller's “terminal verified” flag cannot by itself release a reservation; the hashed receipt must bind the selected task and route.
We also found an error in our first allocation harness. It lowered a previously unused short-window balance while leaving its reset unknown. The router correctly refused that inconsistent observation. We preserved the failed experiment and changed the harness to model reservations against the unchanged snapshot. This prevented a test-fixture mistake from becoming a claim about the policy.
Forty-one focused checks pass in the candidate, including concurrent claims, account aliases, stale data, short-window exhaustion, ten-percent behavior, model restrictions and response failures. That is implementation evidence. It is not proof that the whole fleet has adopted the router.
Release status and the next useful test
The planner, reservation ledger, Jev question adapter and an opt-in native conductor pilot adapter are built. The adapter preserves the installed conductor's admission and resume checks. It changes account ranking and checkpoint-fit checks in memory; it does not rewrite the installed conductor.
The independent Grok source-review job remains held after its native tool refusal. It has not been replaced or retried through another route. Shared concurrent quota reservations are not yet connected to every native launcher, and the account policy is not an estate-wide production rollout. Those are release boundaries, not paperwork to waive.
The next useful evidence is accepted work through the integrated path, with checkpoint demand compared against observed usage and rework. Model selection needs suitable tasks run under comparable conditions and independently assessed against the same acceptance criteria. More questions or repeated guesses cannot substitute for that evidence.
The lesson so far is practical: use ordinary code for ownership, billing, quota arithmetic and resource admission. Preserve explicit model choices. Use Jev for a remaining bounded judgment when there is evidence to judge. Measure the result, retain the failures, and remove a question when it contributes no useful decision.
Reproduce the allocation numbers
The public package uses anonymous account labels and omits private work descriptions, account identities, credentials and machine addresses. It reproduces allocation arithmetic, not historical task performance. The private cohort descriptions are not published, so the attached corpus hash cannot independently establish the task-selection count for a public reader. We supply the anonymous count as an input. The public routing core is the same core exercised by the focused implementation tests; some tests additionally cover private native adapters.
- Experiment results and assumptions
- Recompute script
- Tested routing core used by the script
- Anonymous allocation inputs
- Focused implementation test log
For related work, see our Jev memory-filter trial, six bounded Jev decisions and conductor workflow.
Contributors

James Brady
Problem framing · Routing priorities · Publish decision
GPT-6 AstraAI
Implementation · Evaluation · Drafting
Claude Opus 5AI
Independent source audit · Frozen corpus and labels
Jev 1.13.0AI
Typed evaluation responses