- Project number
- Project / 001
- Year
- 2026
- Category
- Product
- Status
- Under construction
Sepang Vision Lab
A motorsport digital twin and race-strategy research platform built around Sepang International Circuit. You do not drive the car. You act as the race engineer.
- Status
- Under construction
- Role
- Solo: product, engineering, data, ML
- Timeline
- 2026 · 17 milestones
- Data
- 2017 Malaysian GP · 1,024 lap times
- Stack
- React Three Fiber · FastAPI · scikit-learn · MediaPipe
01 /Overview
A strategy workstation, not a racing game.
Sepang Vision Lab puts an interactive digital twin of Sepang next to the tools a strategist would want: historical replay, telemetry, lap-time models, pit-strategy and Monte Carlo simulation, weather scenarios, webcam hand tracking and an AI race engineer.
It is deliberately not a driving game, a fantasy manager or a chatbot with racing trivia. The question it explores is how far honest, public race data can go before you have to start guessing, and how to show clearly where that line is.
02 /The problem
Strategy is decided under uncertainty.
Race strategy is made from partial information: lap times, gaps, pit windows, weather that might arrive. Public data covers only a slice of it. There is timing, but no pedal telemetry, no tyre compounds and no car positions between timing lines.
Most visualisations fill those gaps quietly with plausible-looking numbers. I wanted a system that used what the data actually supports, modelled what it could, and labelled everything else as an assumption.
03 /The idea
Build in layers.
- Observe
- Predict
- Simulate
Each capability depends on the one before it, so the project was built strictly in that order. No layer started until the one beneath it worked and was tested.
- Circuit
- Car movement
- Race state
- Replay
- Telemetry
- ML
- Strategy
- Monte Carlo
- Weather
- Computer vision
- AI race engineer
04 /Current build
What runs today.
Sepang Vision Lab is under construction. These are captures of the current local build on the 2017 Malaysian Grand Prix timing archive. Movement is interpolated between timing lines, not GPS, and the screens say so.

01Historical workspace at lap 35: the 3D circuit, every car with timing coverage, the selected driver and the replay controls.
Current build · 2017 Malaysian GP timing, reconstructed movement
02Tyre and stint analysis: an observed pace trend with its limits stated beside it.
Current build · 2017 Malaysian GP timing, reconstructed movement
03Lap-time ML: the previous-lap baseline beats every trained model on test error, and the panel says so.
Current build · 2017 Malaysian GP timing, reconstructed movement
05 /System architecture
Ingest once. Replay offline.
Step 1: SOURCE
Jolpica F1 API
Ergast-compatible results, laps and pit stops.
Step 2: PROCESS
Ingest
Paginated, paced, cached. Validates lap sequences and the winner’s total time.
Step 3: DATABASE
Parquet cache
Raw-response hashes and retrieval time kept as provenance.
Step 4: API
FastAPI
Pydantic-validated replay, models and strategy endpoints.
Step 5: OUTPUT
Digital twin
React Three Fiber scene; readouts refresh at 10 Hz.
Step 6: USER
Race engineer
Selects drivers, seeks the race, branches strategies.
Circuit geometry
Community GeoJSON, cross-checked against the operator’s circuit map. The projected path is ~5,549 m against a published 5.543 km (0.11% over). Smoothed with a centripetal Catmull–Rom spline.
Historical replay
Cars move by linear reconstruction between recorded lap crossings. It is not GPS, and it is labelled that way. A retired car is never extrapolated around the track.
06 /Data / intelligence
The baseline won.
The lap-time experiment predicts lap N from the three previous laps, the lap number and the last recorded position. Splits are chronological across the whole field (train by 45 min, validate to 65 min, test after), so no model ever sees the future.
Random Forest won on validation MAE and was selected. On the test window it did not beat the naive previous-lap baseline, and the app says so prominently instead of hiding it.
| Model | Validation MAE (s) | Test MAE (s) | Test RMSE (s) | Test R² |
|---|---|---|---|---|
| Previous-lap baseline | 1.584 | 0.795 | 2.162 | −0.368 |
| Linear Regression | 1.634 | 1.098 | 1.709 | 0.145 |
| Random Forest (selected) | 1.505 | 1.160 | 1.731 | 0.123 |
| XGBoost | 1.843 | 1.566 | 1.934 | −0.095 |
Single race only; not validated on unseen races.
Stint pace
Median-of-pairwise-slopes regression per stint. Pit laps and the following lap are excluded but stay visible. Tyre-only degradation is reported as not identifiable from this data.
Monte Carlo
1,000–10,000 seeded scenarios per comparison. Every strategy shares the same random draws (common random numbers), so the differences come from the strategy, not the noise.
07 /Build
Constrain what can be trusted.
01Strategy Lab
Freezes a hypothetical branch from a driver’s last completed lap and compares stay out, pit now and a delayed stop. All defaults are editable and labelled illustrative.
02Weather Lab
Dry → light rain → wet → drying, with five deterministic tyre plans (slicks, early or delayed inters, early or delayed wets) and editable per-lap penalties.
03Hand tracking
MediaPipe HandLandmarker runs locally in a web worker, capped at 15 fps. No camera request is made until the user explicitly enables it.
04AI race engineer
The model may call exactly one tool, which re-runs the captured comparison. It cannot change inputs, and only simulator-derived numbers are shown. The no-AI path is labelled no-AI.
09 /Outcome
Seventeen milestones in.
- Milestones
- 17
- Recorded laps
- 1,024
- 20 drivers · 20 pit records
- Scenarios / run
- 10,000
- Monte Carlo, seeded
- Geometry error
- 0.11%
- vs published length
Still open: the gesture model is waiting on real labelled recordings, and the live AI race-engineer path still needs end-to-end verification. Both are marked as open in the project rather than shown as finished.
10 /What I learned
A baseline is a result.
01Beat the naive model first
A previous-lap guess outperformed every trained model on test. Showing that is more useful than a flattering chart.
02Label the unknown
Saying “tyre degradation is not identifiable here” is a finding, not a gap.
03Layers keep you honest
Building strictly bottom-up meant every new feature stood on something already tested.
04Give AI a narrow job
Letting the model explain the simulator, but never invent numbers, made the AI feature trustworthy.
Independent motorsport technology and research project. Motorsport names and trademarks belong to their respective owners. No official affiliation or endorsement is implied.
11 /Next project
BalangA Malaysian probability party game: read the jar of 25 kuih tokens, bet on secret prediction cards, and play friends in peer-to-peer rooms with no backend.

