WRECK VALLEY · SELF-PLAY LAB · 0.61.6

Learn from the fight.

A local combat policy plays accelerated matches against earlier checkpoints and fixed opponent styles. Outcomes change its learned weights. A separate tournament selects a candidate; unseen paired matches decide promotion.

2,944Training encounters
51.3 hSimulated training time
63.5%Held-out win score; draws count ½
$0.00Training API spend · local CPU
PROMOTED · selfplay-g18

This checkpoint was included in the release under the original gate. See the evaluation correction below.

Evaluation correction · neural RL audit

These original results swapped starting positions but always processed the candidate first. An identical-driver check exposed an order advantage. The test harness now swaps both positions and processing order; the historical 55.6% result below should not be treated as an unbiased improvement estimate. See the corrected neural evaluation. Live vehicle physics is unchanged.

What learns

This trains our local combat controller. Jev’s hosted model weights do not change. Jev keeps choosing high-level tactics; during a fresh attack request, the learned policy selects an approach four times per second. It sees relative position, headings, speeds, closing speed, health and nitro. The seven approaches are tracking, short leading, interception, driving through the target path, left/right flanks and a braking bait.

The algorithm is a cross-entropy evolution strategy: mutate policy weights, play matches, rank by score-based reward, average elite policies and retain previous opponents. This is outcome-based self-play training, rather than PPO or fine-tuning the hosted Jev model. The network is a small linear policy with 84 learned weights. It runs locally with no extra model requests.

Training and promotion

Training win scores use changing starting positions and league opponents. This curve is diagnostic; it is not evidence of steadily increasing strength. Checkpoints are saved atomically and the trainer can resume from its saved population state.

  1. Train against the stock controller, fixed approaches and archived self-play opponents.
  2. Select a checkpoint in a separate validation tournament.
  3. Evaluate unseen seeds, twice per start with starting pose and velocity swapped.
  4. Promote only if both the overall league and stock-opponent paired bootstrap 95% interval are entirely above 50%. A checkpoint hash binds the evaluation to the exact weights.

Final league estimate: 63.5% (58.9%–68.5%), 192 paired starts. Against stock: 55.6% (52.6%–58.7%), 512 pairs.

OpponentMatchesWinsDrawsLosses
stock-06159654042
fixed-intercept9665031
fixed-drive_through9641055
fixed-flank_left9684012

Full-round transfer

128 fresh six-minute games against the stock driver: 57.0% win score. Paired bootstrap interval: 48.4%–66.4%. This interval includes 50%; this is a transfer/regression check rather than conclusive evidence of superior full-game play. Human opponents were not tested.

The stock-only short-encounter confirmation contains 1024 additional fresh matches. Its estimate is reported separately from the balanced league table above.

Replay the actual simulation

These top-down replays come from the real city simulation, including road geometry, collision damage, pickups and traffic. A win and a loss are selected when available. Cyan is the candidate; coral is its opponent. Use the time slider to inspect approaches and health.

Loading replay…

3D self-play recording

Controlled self-play fixture using the production renderer and physics. Both cars receive local automated inputs; model transport is a zero-cost stub. Chromium/SwiftShader at low quality is not a human playtest or hardware performance benchmark. This recording is separate from the seeded evaluation matches. Reviewed frames show repeated close approaches and pickup collection; the cars also slow beside roadside obstacles and need recovery.

Across the entire recorded fixture, the candidate issued 220 nitro and 26 drift input substeps, and finished with 624 points against 145. The clip is an excerpt.

Separate real-Jev integration test: 19/19 applied model responses, no provider or JavaScript errors, and the isolated test budget stopped paid control correctly. This integration test cost $0.004834; the offline training itself made no model calls.

What the first experiment taught us

The last checkpoint of the initial 1,152-match run won only 35.4% on its first unseen evaluation and was rejected. A separate checkpoint tournament found an earlier version that performed better. Its next unseen evaluation reached 63.5% overall, but uncertainty against the stock opponent was still too large to pass. Wider training includes ambient traffic and more paired starts per generation. Four additional generations trained on full six-minute rounds. The checkpoint tournament can retain an earlier policy when later training does not improve it. Failed evaluations remain in the development artifacts.

Limits and safeguards

The curriculum starts with short encounters of up to 35 seconds, then includes complete six-minute games. Short held-out encounters last up to 60 seconds and end at the first wreck; complete rounds are evaluated separately. It does not measure human win rates, all vehicles, the island or every long-term strategy. Opponent tactics are deterministic local approximations; the self-play matches do not call Jev’s API. Improving this benchmark does not make the opponent unbeatable.

The learned policy can propose approach targets and a slower braking approach. Shared health, damage, traction, nitro, collision and vehicle-speed physics remain unchanged. Repair commitment, protection, ramp handling and obstacle navigation still override it. Training runs offline in bounded worker jobs; the game server loads only a frozen promoted checkpoint. Ordinary AI does not use this policy.

488 automated tests, lint and build passed. Regression checks cover deterministic replay, swapped starting states, no network calls during training, invalid checkpoints, action masking, repair/protection guards, promotion rejection and checkpoint substitution.

Research basis

TypeSafe documents Jev as a structured decision model accepting text state. The local trainable controller fits that architecture: TypeSafe System One. Competitive self-play research motivates a pool of past opponents to reduce overfitting: Competitive self-play. Reward-driven parameter search is an established alternative to gradient-based RL: Evolution strategies.