WRECK VALLEY / 0.61.0 / RESEARCH + BUILD REPORT / 20 SEP 2026
Jev chooses the tactic.
The city puts it to the test.
A playable private duel against one model-directed rival, with visible decisions, measured costs and automatic spending limits.

What the research supports
OpenRouter lists Jev 1.13 as a low-cost structured-decision model: $0.042 per million input tokens, zero output-token charge, 32K context. Its endpoint metadata reports text-to-decisions. TypeSafe’s API takes a state plus typed questions; the Choice primitive returns an allowed option with confidence and probabilities.
TypeSafe’s Doom demonstration uses structured state and acknowledges that a conventional bot could perform better. It does not establish skill in this driving game. The vendor’s documented limitations support keeping arithmetic and precise control in code.
We verified POST https://openrouter.ai/api/alpha/decisions directly with the configured key. The first repair probe used 412 input tokens, cost $0.000017304 and took 608 ms. The request alias typesafe/jev-1.13 resolved to typesafe/jev-1.13-20260917. This is an alpha route and a minor-version alias; provider changes remain possible.
How this opponent works
Choices include pursuit, interception, repair, public objectives and patrol. No screenshot analysis, generated coordinates, extra grip or power. Radar reaches 180 m and includes nearby rival health; this is not a camera-only agent. Names, profile IDs, chat, exact coordinates and future inputs are excluded from provider requests.
One request may run per room; two rooms may run at once. Calls have a 1.5-second deadline and decisions expire after three seconds. Death, respawn, round changes and departure invalidate stale work. Errors back off. Model calls stop during loading and intermission. A local controller takes over visibly on failures or limits. All experimental records stay separate from normal modes.
Live model. Real client inputs.
The clip is an excerpt of the recorded browser run at its original timing. Keyboard events passed through the production client and authoritative simulation. No position, health or match-clock overrides were used. Chromium used SwiftShader and low graphics; the recording is neither a hardware frame-rate benchmark nor human playtesting.
What the first recording taught us
The first run selected the distant hot zone on 27 of 29 calls, despite that zone being unable to pay without a nearby rival. Its closest sampled approach was 67 m. This exposed an incomplete state description. We added candidate distances, the hot-zone scoring rule and an eligibility filter, and oriented the opening road spawn toward the encounter.
The revised run selected pursuit on 21 of 24 calls. The rival travelled approximately 883 m across 87 seconds of sampled gameplay, including the local-fallback period. This is a short diagnostic run, with changed conditions; it is not a controlled win-rate comparison. No superiority claim follows from it.
A small, observable experiment
Production defaults reserve $1 per UTC day across all players and $0.05 per match. The demo operator pays. The browser test deliberately used a much smaller $0.002 match cap to exercise cutoff. It paused new calls after 24 attempts, at $0.000663 confirmed spend.
The gap below the cap is deliberate: the next request needs a full $0.001344 reservation, the verified maximum 32K input charge. Confirmed provider cost replaces that reservation after a response. Missing billing and timeouts keep it counted. Restarting or rejoining cannot erase the global ledger.

At the revised run’s average cost, 360 calls would be approximately $0.009940 for a six-minute round, or $0.099401 per continuous hour at one call/second. These are estimates from a small sample; state size, provider prices and failed calls can change the result.

Global daily aggregates are public. Match history belongs to the current browser’s driver profile. Credentials never reach the client. Reported totals are application telemetry, not a provider invoice or whole-account balance. Prices are checked before play and every ten minutes; unknown prices pause calls. An OpenRouter key spending limit provides an independent cap if upstream billing differs from metadata.
Validation and review
| Check | Result |
|---|---|
| Complete automated suite | 427 passed; zero failures |
| Lint and production build | Passed |
| Cost and lifecycle regression coverage | Atomic reservations, unknown billing, restart, UTC rollover, timeout, backoff, late answers, death, respawn, disconnect, round change, pricing failure |
| Mode isolation | One Jev rival, global room cap, private histories, independent records; ordinary AI still has seven rivals |
| Snapshot financial precision | Sub-cent amounts survive encode/decode without positional rounding |
| Live browser | Actual inputs and rival movement, live decisions, automatic cutoff, desktop monitor, mobile monitor without page overflow; no JavaScript errors |
| Two live recordings combined | 53 attempted / 53 applied; p50 328 ms, p95 424 ms |
Twelve synthetic checks
12/12 choices matched the expected tactic across four simple scenarios, each repeated three times. “Expected” is a developer judgment. The scenarios are not a held-out benchmark and do not measure full-game skill. Confirmed cost: $0.000277.
Show all scenario outcomes
| Scenario | Expected | Actual | Latency | Cost |
|---|---|---|---|---|
| Critical damage, nearby repair | repair | repair | 572 ms | $0.000024 |
| Healthy, vulnerable rival directly ahead | pursue | pursue | 304 ms | $0.000025 |
| No radar contact, public objective | objective | objective | 294 ms | $0.000021 |
| Protected rival, no attack available | objective | objective | 323 ms | $0.000022 |
| Critical damage, nearby repair | repair | repair | 259 ms | $0.000024 |
| Healthy, vulnerable rival directly ahead | pursue | pursue | 270 ms | $0.000025 |
| No radar contact, public objective | objective | objective | 261 ms | $0.000021 |
| Protected rival, no attack available | objective | objective | 343 ms | $0.000022 |
| Critical damage, nearby repair | repair | repair | 314 ms | $0.000024 |
| Healthy, vulnerable rival directly ahead | pursue | pursue | 328 ms | $0.000025 |
| No radar contact, public objective | objective | objective | 284 ms | $0.000021 |
| Protected rival, no attack available | objective | objective | 303 ms | $0.000022 |
What we still need to learn
Can the model win more often, get stuck less, or feel more interesting than our conventional rival? This build does not answer those questions. The next useful study is a matched-seed, full-round comparison followed by human playtests. The controller and observation design can dominate the result; a valid model response alone is not success.