WRECK VALLEY · RESEARCH EXPERIMENT
Teaching Jev’s driver through self-play.
Real gradient-based reinforcement learning over the actual game simulation. A separate neural network learns adjustments to the existing driver through competitive self-play. Jev’s hosted weights are unchanged.
The frozen driver is enabled for Nightfall GT versus Nightfall GT duels after passing independent short, full-round and mixed-situation tests plus live-runtime checks. Other vehicle matchups retain the established driver. These simulator results do not establish a human win rate.
What was trained?
This run updates our own local neural driver. It does not change TypeSafe’s hosted Jev weights. The hosted model still selects high-level tactics in live games; self-play uses local tactical rules and makes no model API calls.
What the network controls
Pick a destination and baseline driving inputs. Local rules stand in for hosted Jev during offline training.
49 state features → two 64-unit layers → steering, throttle/brake, nitro and drift adjustments at 5Hz.
Apply legal inputs at 60Hz. Repair, recovery, shields, ramps and airborne handling retain the existing driver.
The actor contains 8,270 weights and biases. A separate value estimator learns expected future reward during training. This is a feed-forward neural model, trained with backpropagation using PPO in Stable Baselines3. It is not a fine-tuned language model, and it currently learns driving adjustments rather than new strategic destinations.
Independent short encounters
512 unseen encounters, up to 60 seconds or first wreck. Each starting scenario is played twice with poses and physics processing order exchanged. Wins follow actual game score; draws count half a win. Intervals resample whole paired starts.
| Opponent | Wins / losses / draws | Win score | 95% paired interval |
|---|---|---|---|
| Deployed combat champion | 160 / 95 / 1 | 62.7% | 57.4%–68.4% |
| Stock driver | 179 / 77 / 0 | 69.9% | 64.5%–75.0% |
Complete six-minute games
256 fresh full rounds with repairs and respawns. These are local simulation benchmarks, not human win rates or tests of hosted Jev’s strategic choices.
| Opponent | Wins / losses / draws | Win score | 95% paired interval |
|---|---|---|---|
| Deployed combat champion | 91 / 37 / 0 | 71.1% | 62.5%–78.9% |
| Stock driver | 98 / 30 / 0 | 76.6% | 69.5%–83.6% |
Inspect a win and a loss
Cyan is the neural candidate; coral is the current champion. The replay shows actual applied controls and health. These cases are illustrations selected from the evaluation, not additional independent evidence.
Loading…
Continued training, harder starts
Resumed the original 131,072-step checkpoint with its optimizer, random state and frozen opponent league. Added 393,216 steps and 1,550 completed training matches. Half the new starts keep the original setup; the rest cover repair pressure, limited nitro, faster approaches and wounded rivals. Starting resources and physics processing order exchange with player positions. Hosted Jev is not fine-tuned.
Checkpoint selection on separate validation matches
| Checkpoint | Combined win score | vs champion | vs stock |
|---|---|---|---|
| ppo-524288 · selected | 64.8% | 64.1% | 65.6% |
| ppo-393216 | 58.6% | 59.4% | 57.8% |
| ppo-262144 | 57.8% | 51.6% | 64.1% |
| ppo-131072 | 49.2% | 51.6% | 46.9% |
Selection uses 32 paired starts per opponent on the mixed-v1 curriculum. The selected actor is frozen before fresh test seeds are evaluated. Draws count as half a win. These are simulator results, not human win rates.
Independent mixed-situation test
| Opponent | Wins / losses / draws | Win score | 95% paired interval |
|---|---|---|---|
| Deployed combat champion | 76 / 52 / 0 | 59.4% | 52.3%–65.6% |
| Stock driver | 81 / 47 / 0 | 63.3% | 57.0%–69.5% |
Fresh resource-pressure and fast-approach starts; 256 matches, paired starting positions and processing order.
Original neural checkpoint on the same fresh short test
| Opponent | Wins / losses / draws | Win score | 95% paired interval |
|---|---|---|---|
| Deployed combat champion | 131 / 125 / 0 | 51.2% | 45.3%–56.6% |
| Stock driver | 141 / 115 / 0 | 55.1% | 49.6%–60.2% |
Checkpoint ppo-131072, evaluated after selection. Comparison is against the same stock and champion opponents; it is not a direct match between the two neural checkpoints.
Change versus the original neural checkpoint, in percentage points (paired 95% interval): champion: 11.5 (3.9 to 19.5); stock: 14.8 (7.8 to 21.1).
Applied controls in independent short matches
| Opponent | N₂O | Drift | Braking | Changed baseline input |
|---|---|---|---|---|
| champion | 12.8% | 18.8% | 16.9% | 39.2% |
| stock | 13.1% | 15.5% | 14.7% | 34.8% |
Percentages of applied simulation input substeps. These show what reached the car controls; using a control more often does not by itself demonstrate better timing or stronger play.
Recorded 3D fixture
Actual renderer and physics; both cars automated. The camera follows the baseline rival; the car labelled AI · JEV 1.13 uses the trained driver. Close-contact recovery and occasional camera obstruction remain visible. Chromium/SwiftShader at low quality, not human playtesting or a GPU performance benchmark. The candidate adjusts the current champion’s inputs; no paid model calls are made. The clip is separate from the evaluation matches.
Across the entire fixture: 1705 changed input substeps, 302 nitro substeps, 193 drift substeps, 244 braking substeps. A substep is 1/60 second; counts do not measure whether each action was useful.
Live input monitor

The LIVE · LEARNED indicator identifies the local neural layer. Hosted model decisions and API costs remain separate.
Verification and cost
The continuation ran in 16.6 minutes on this CPU host. All retained completed training matches cover 29.0 simulated hours in completed matches. No GPU or model API was used. The network’s weights changed, and Python/JavaScript logits and actions agree across 128 test observations. Checkpoint resume was exercised separately.
610 automated tests passed, plus lint and the recorded browser fixture. Checks cover real physics determinism, observation validity, repair/recovery guards, malformed models/actions, terminal scoring, processing-order fairness and export parity. The live driver shares the tested JavaScript inference implementation. It loads a hash-verified frozen artifact and has no Python or PyTorch dependency. Training never runs in the game server.
Correction to the earlier benchmark
The previous self-play arena exchanged starting poses but always processed the candidate first. An identical-driver check exposed an order advantage. Both test harnesses now swap processing order too, and the original promotion gate rejects legacy results. The earlier report’s 55.6% figure is preserved with a visible correction and should not be treated as an unbiased improvement estimate. Live vehicle physics was not changed.
What should improve next?
- Broaden training situations and run longer, across multiple training seeds. More steps alone are not evidence of improvement.
- Teach a tactical head when to attack, intercept or repair. The present neural layer has limited authority over those decisions.
- Use successful human demonstrations to teach useful driving before self-play, then evaluate against held-out human sessions.
- Test different vehicles, evasive opponents, difficult corners, low-health starts and coastal terrain before expanding the rollout to those cases.
The current training covers one car and a limited set of road starts. The independent tests establish an advantage over the tested local opponents. Human play and other vehicle classes still need evaluation.