WRECK VALLEY · RESEARCH EXPERIMENT
A neural driver that learns.
Real gradient-based reinforcement learning over the actual game simulation. A separate neural network learns adjustments to the existing driver through competitive self-play. Jev’s hosted weights are unchanged.
The candidate did not clear all independent performance gates. It stays experimental; the live game retains the existing 0.61.6 controller. Training changed real neural weights, but that alone does not establish better gameplay.
Can we fine-tune Jev?
TypeSafe explicitly says Jev is not fine-tuned or LoRA-adapted with customer data: all accounts share the same model weights. We can train our own downstream network instead. This pilot does exactly that, using structured game state and the current controller’s recommendations. Official TypeSafe documentation.
What the network controls
Pick a destination and baseline driving inputs. Local rules stand in for hosted Jev during offline training.
49 state features → two 64-unit layers → steering, throttle/brake, nitro and drift adjustments at 5Hz.
Apply legal inputs at 60Hz. Repair, recovery, shields, ramps and airborne handling retain the existing driver.
The actor contains 8,270 weights and biases. A separate value estimator learns expected future reward during training. This is a feed-forward neural model, trained with backpropagation using PPO in Stable Baselines3. It is not a fine-tuned language model, and it currently learns driving adjustments rather than new strategic destinations.
Independent short encounters
256 unseen encounters, up to 60 seconds or first wreck. Each starting scenario is played twice with poses and physics processing order exchanged. Wins follow actual game score; draws count half a win. Intervals resample whole paired starts.
| Opponent | Wins / losses / draws | Win score | 95% paired interval |
|---|---|---|---|
| Current champion (0.61.6) | 72 / 56 / 0 | 56.3% | 48.4%–63.3% |
| Old stock driver (0.61.5) | 70 / 58 / 0 | 54.7% | 46.1%–63.3% |
Complete six-minute games
128 fresh full rounds with repairs and respawns. These are local simulation benchmarks, not human win rates or tests of hosted Jev’s strategic choices.
| Opponent | Wins / losses / draws | Win score | 95% paired interval |
|---|---|---|---|
| Current champion (0.61.6) | 34 / 30 / 0 | 53.1% | 42.2%–64.1% |
| Old stock driver (0.61.5) | 38 / 26 / 0 | 59.4% | 46.9%–71.9% |
Checkpoint selection
The candidate was chosen from a separate validation set, before either test set was evaluated. The untrained network keeps the baseline controls and scores exactly 50% against the identical champion after processing order is swapped. Improvements against only the older stock driver do not qualify as a stronger replacement.
| Checkpoint | Validation vs champion | Validation vs stock |
|---|---|---|
| ppo-131072 | 39.1% | 65.6% |
| ppo-0 | 50.0% | 50.0% |
| ppo-32768 | 50.0% | 50.0% |
| ppo-65536 | 43.8% | 53.1% |
| ppo-98304 | 46.9% | 50.0% |
Inspect a win and a loss
Cyan is the neural candidate; coral is the current champion. The replay shows actual applied controls and health. These cases are illustrations selected from the evaluation, not additional independent evidence.
Loading…
Recorded 3D fixture
Actual renderer and physics; both cars automated. Chromium/SwiftShader at low quality, not human playtesting or a GPU performance benchmark. The candidate adjusts the current champion’s inputs; no paid model calls are made. The clip is separate from the evaluation matches.
Across the entire fixture: 926 changed input substeps, 555 nitro substeps, 107 drift substeps, 520 braking substeps. A substep is 1/60 second; counts do not measure whether each action was useful.
Verification and cost
Training completed in 4.9 minutes on this CPU host, covering 7.3 simulated hours in completed matches. No GPU or model API was used. The network’s weights changed, and Python/JavaScript logits and actions agree across 128 test observations. Checkpoint resume was exercised separately.
495 automated tests passed, plus lint and the recorded browser fixture. Checks cover real physics determinism, observation validity, repair/recovery guards, malformed models/actions, terminal scoring, processing-order fairness and export parity. The production driver has no dependency on the experimental training code.
Correction to the earlier benchmark
The previous self-play arena exchanged starting poses but always processed the candidate first. An identical-driver check exposed an order advantage. Both test harnesses now swap processing order too, and the original promotion gate rejects legacy results. The earlier report’s 55.6% figure is preserved with a visible correction and should not be treated as an unbiased improvement estimate. Live vehicle physics was not changed.
What should improve next?
- Broaden training situations and run longer, across multiple training seeds. More steps alone are not evidence of improvement.
- Teach a tactical head when to attack, intercept or repair. The present neural layer has limited authority over those decisions.
- Use successful human demonstrations to teach useful driving before self-play, then evaluate against held-out human sessions.
- Test different vehicles, evasive opponents, difficult corners, low-health starts and coastal terrain before considering deployment.
The current experiment covers one car and a limited set of road starts. It demonstrates a working neural training pipeline and exposes its performance limits. It does not establish stronger human-level opposition.