WRECK VALLEY / 0.61.0 / RESEARCH + BUILD REPORT / 20 SEP 2026

Jev chooses the tactic.
The city puts it to the test.

A playable private duel against one model-directed rival, with visible decisions, measured costs and automatic spending limits.

1 vs 1Private, six-minute experimental rounds
$0.05Production match cap · $1 shared UTC-day cap
427Passing automated tests, including 14 new Jev tests
$0.001619Confirmed cost of probe, scenarios and both recorded local sessions
Jev duel in Port Nova with live model tactic and cost HUD
The new lobby option starts a private duel. The HUD identifies fresh model tactics, decision latency and match spending.

What the research supports

OpenRouter lists Jev 1.13 as a low-cost structured-decision model: $0.042 per million input tokens, zero output-token charge, 32K context. Its endpoint metadata reports text-to-decisions. TypeSafe’s API takes a state plus typed questions; the Choice primitive returns an allowed option with confidence and probabilities.

TypeSafe’s Doom demonstration uses structured state and acknowledges that a conventional bot could perform better. It does not establish skill in this driving game. The vendor’s documented limitations support keeping arithmetic and precise control in code.

Assessment: Jev is a plausible inexpensive tactical selector. Our tests establish a working integration; they do not establish that it is a stronger or more enjoyable opponent than ordinary AI. A confident, correctly typed answer can still be a poor decision.

We verified POST https://openrouter.ai/api/alpha/decisions directly with the configured key. The first repair probe used 412 input tokens, cost $0.000017304 and took 608 ms. The request alias typesafe/jev-1.13 resolved to typesafe/jev-1.13-20260917. This is an alpha route and a minor-version alias; provider changes remain possible.

How this opponent works

01 · ObserveCompact local radar state; code computes distances and bearings.
02 · ReserveSQLite checks global and match budgets atomically.
03 · DecideJev chooses a valid tactic, once per second.
04 · DriveExisting controller steers and collides using ordinary physics.

Choices include pursuit, interception, repair, public objectives and patrol. No screenshot analysis, generated coordinates, extra grip or power. Radar reaches 180 m and includes nearby rival health; this is not a camera-only agent. Names, profile IDs, chat, exact coordinates and future inputs are excluded from provider requests.

One request may run per room; two rooms may run at once. Calls have a 1.5-second deadline and decisions expire after three seconds. Death, respawn, round changes and departure invalidate stale work. Errors back off. Model calls stop during loading and intermission. A local controller takes over visibly on failures or limits. All experimental records stay separate from normal modes.

Live model. Real client inputs.

The clip is an excerpt of the recorded browser run at its original timing. Keyboard events passed through the production client and authoritative simulation. No position, health or match-clock overrides were used. Chromium used SwiftShader and low graphics; the recording is neither a hardware frame-rate benchmark nor human playtesting.

24/24Decisions applied in the revised run
3 mClosest sampled rival distance, including fallback period
$0.000663Revised run confirmed model cost

What the first recording taught us

The first run selected the distant hot zone on 27 of 29 calls, despite that zone being unable to pay without a nearby rival. Its closest sampled approach was 67 m. This exposed an incomplete state description. We added candidate distances, the hot-zone scoring rule and an eligibility filter, and oriented the opening road spawn toward the encounter.

The revised run selected pursuit on 21 of 24 calls. The rival travelled approximately 883 m across 87 seconds of sampled gameplay, including the local-fallback period. This is a short diagnostic run, with changed conditions; it is not a controlled win-rate comparison. No superiority claim follows from it.

A small, observable experiment

Production defaults reserve $1 per UTC day across all players and $0.05 per match. The demo operator pays. The browser test deliberately used a much smaller $0.002 match cap to exercise cutoff. It paused new calls after 24 attempts, at $0.000663 confirmed spend.

Test match cap $0.0020 s87 s

The gap below the cap is deliberate: the next request needs a full $0.001344 reservation, the verified maximum 32K input charge. Confirmed provider cost replaces that reservation after a response. Missing billing and timeouts keep it counted. Restarting or rejoining cannot erase the global ledger.

Budget cutoff shown as local fallback in the game
The rival continues driving after cutoff, with an explicit local-fallback label. This capture uses the smaller test-only budget.

At the revised run’s average cost, 360 calls would be approximately $0.009940 for a six-minute round, or $0.099401 per continuous hour at one call/second. These are estimates from a small sample; state size, provider prices and failed calls can change the result.

Cost monitor showing confirmed cost, reservations, remaining budget and private match history
The monitor refreshes every five seconds. It shows production ledger calls when opened on the live demo; this screenshot shows isolated test-ledger totals.

Global daily aggregates are public. Match history belongs to the current browser’s driver profile. Credentials never reach the client. Reported totals are application telemetry, not a provider invoice or whole-account balance. Prices are checked before play and every ten minutes; unknown prices pause calls. An OpenRouter key spending limit provides an independent cap if upstream billing differs from metadata.

Validation and review

CheckResult
Complete automated suite427 passed; zero failures
Lint and production buildPassed
Cost and lifecycle regression coverageAtomic reservations, unknown billing, restart, UTC rollover, timeout, backoff, late answers, death, respawn, disconnect, round change, pricing failure
Mode isolationOne Jev rival, global room cap, private histories, independent records; ordinary AI still has seven rivals
Snapshot financial precisionSub-cent amounts survive encode/decode without positional rounding
Live browserActual inputs and rival movement, live decisions, automatic cutoff, desktop monitor, mobile monitor without page overflow; no JavaScript errors
Two live recordings combined53 attempted / 53 applied; p50 328 ms, p95 424 ms

Twelve synthetic checks

12/12 choices matched the expected tactic across four simple scenarios, each repeated three times. “Expected” is a developer judgment. The scenarios are not a held-out benchmark and do not measure full-game skill. Confirmed cost: $0.000277.

Show all scenario outcomes
ScenarioExpectedActualLatencyCost
Critical damage, nearby repairrepairrepair572 ms$0.000024
Healthy, vulnerable rival directly aheadpursuepursue304 ms$0.000025
No radar contact, public objectiveobjectiveobjective294 ms$0.000021
Protected rival, no attack availableobjectiveobjective323 ms$0.000022
Critical damage, nearby repairrepairrepair259 ms$0.000024
Healthy, vulnerable rival directly aheadpursuepursue270 ms$0.000025
No radar contact, public objectiveobjectiveobjective261 ms$0.000021
Protected rival, no attack availableobjectiveobjective343 ms$0.000022
Critical damage, nearby repairrepairrepair314 ms$0.000024
Healthy, vulnerable rival directly aheadpursuepursue328 ms$0.000025
No radar contact, public objectiveobjectiveobjective284 ms$0.000021
Protected rival, no attack availableobjectiveobjective303 ms$0.000022

Download sanitized measurements

What we still need to learn

Can the model win more often, get stuck less, or feel more interesting than our conventional rival? This build does not answer those questions. The next useful study is a matched-seed, full-round comparison followed by human playtests. The controller and observation design can dominate the result; a valid model response alone is not success.