Files
DarkflameServer/docs/LiveUpdate.md
Aaron Kimbrell 5658fb036f docs: world hot reload
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
2026-09-30 02:51:21 -05:00

177 lines
12 KiB
Markdown

# Live updates
Moving every running server onto a new build of the server binaries without taking the server down. Players are moved
to new world instances started from the new binaries; the other servers restart one by one. Master keeps running.
Code: `dMasterServer/LiveUpdateMachine.h` (the order, no master state; unit tested), `dMasterServer/LiveUpdateCoordinator`
(master's glue), `dMasterServer/MigrationCoordinator` and `dGame/dUtilities/WorldMigration` (moving one instance's
players, see [SeamlessTransfer.md](SeamlessTransfer.md)), `dNet/master/LiveUpdate.h` (messages).
## Starting one
Put the new build in place (the binaries master starts are the ones in its own directory), then:
| Trigger | Who |
| --- | --- |
| Dashboard, home page, **Live update** card | `server_live_update` (default GM 9) |
| `/liveupdate start [warn seconds]`, `/liveupdate cancel`, `/liveupdate status` | GM 9 (paired with `server_live_update`) |
| `kill -USR2 <master pid>` (not on Windows) | whoever can signal master |
Nothing happens without one of these. Master refuses when one is already running, when it is shutting down, or when
`WorldServer`, `AuthServer`, `ChatServer` (and `DashboardServer` / `UgcServer` when enabled) are missing or empty in
the binary directory (a build still being written).
Cancelling starts nothing new; what is under way finishes.
## Sequence
1. **Database**: the new build's migrations (`MigrationRunner::RunMigrations`, `RunSQLiteMigrations`). Failure stops the
update before anything else is touched. The running servers must cope with the new schema until they are replaced.
2. **UGC server, auth, chat** (together):
* UGC: `LIVE_UPDATE_RETIRE`. It drops its queue (the rows stay pending in the database), finishes and records the jobs
it is running, then exits. After `live_update_ugc_drain_timeout` it gets `SHUTDOWN` (running jobs are made again).
* Auth: `SHUTDOWN`. Logins fail until the new one is up (a few seconds); players already in game are not affected.
* Chat: `LIVE_UPDATE_RETIRE`. It sends its teams to master (`CHAT_HANDOFF`) and exits without logging anyone out.
* Master starts the new process when the old one disconnects, as it always does. A server counts as replaced once a
new one connects. One that doesn't come back within `live_update_service_timeout` is started again (3 tries).
* When the new chat server connects, master gives it the teams, and once it is up sends `CHAT_SERVER_READY` to every
world: each connects at once and sends its loaded players again (`LoginSessionNotify` with `resync`). The new chat
server takes them over without logging a login, and reads their friends lists.
3. **World instances**, after chat is back and 3 s for the worlds to reconnect. Character selection first, then the
busiest; `live_update_parallel_worlds` at a time (waiting instances don't take a slot). Instances started after the
update began are already on the new build and are not touched. What happens to each is decided when its turn comes:
| Instance | Plan |
| --- | --- |
| Nobody there | Stopped. Zones in `prestart_worlds` (and character selection) get a new instance first; the old one stops once it is ready. |
| Public world with players | Replaced: a new instance starts, players are moved, the old one stops. |
| Property (clone) | Saved and frozen first (`MIGRATE_PREPARE`), then replaced as above. |
| Private instance | Replaced by a new private instance with the same password. |
| Activity zone (any `Activities.instanceMapID`: races, minigames) | Draining: nobody new goes there; its players finish. After `live_update_activity_wait` whoever is left is moved to a new instance (the activity is lost). |
| Character selection | A new one starts at once and takes all logins. The old one drains; after `live_update_char_select_wait` whoever is still there is moved to the new one. |
4. **Dashboard**, last: `SHUTDOWN`; master starts the new one. Sessions survive (JWT, secret in `dashboard_jwt_secret`
or `jwt_secret`). The new dashboard asks master for the status when it connects.
## States
A world: `pending` → (`preparing`: property being saved) → `starting` (new instance launching) → `ready` (new instance
up, players warned) → `draining` (players being moved) → `stopping` (old instance shutting down) → `stopped`.
`waiting`: an activity zone or character selection waiting for its players to leave by themselves.
A server: `pending` → `stopping` (UGC: `draining`) → `starting` → `stopped` (chat: `ready` for 3 s in between).
`failed`: left as it was; a world keeps running on the old build and takes players again (a property some players
already went to stays with the new instance). `skipped`: not running, not enabled, or cancelled before its turn.
The update: `running` → `done` (possibly with failed rows), or `failed` (database migrations), `cancelling` →
`cancelled`.
The dashboard's world list shows instances being emptied as **Moving players** (`ServerListResponse` state `DRAINING`).
Master logs every change of every row (`Live update N: ...`).
## Routing during an update
`InstanceManager::FindInstance` (`InstanceMigration::AcceptsNewPlayers`) skips draining instances, so new zone requests,
logins and friend transfers go to new instances (started if needed, waiting for them to be ready). A private
instance's password finds its replacement (`FindPrivateInstance` skips draining ones). While a property is being saved
its visitors still go to the old instance, where nobody can build.
## Moving players
Each move is an instance migration (`MigrationCoordinator::Start` with `Options::liveUpdate`), see
[SeamlessTransfer.md](SeamlessTransfer.md):
* Players get the game's Mythran Maintenance Alert, then after `live_update_warn_seconds` (or the value picked for this
update) up to 10 a second are saved, locked and sent `TRANSFER_TO_WORLD` with the Mythran shift flag.
* Dead or building players wait up to `live_update_player_wait`, then go anyway. An open trade is cancelled.
* Where they stood is carried (`CarriedPlayerState` position) and applied when the new instance creates them, also on
properties and Moon Base where the saved character doesn't keep it. The pet that was out is summoned again.
* Character selection has no characters loaded: its users are just sent to the new one, which sends them their
characters (no maintenance notice).
### Properties
The new instance loads the property from the database when it starts, so before it is started the old one:
1. tells builders "building on this property ends in N seconds" and waits up to `live_update_property_build_wait` for
nobody to be building;
2. takes anyone still building out of build mode, saves the property, and freezes it: nobody can build, place, pick up,
claim or BBB-save there, and it is never saved again (disconnects and shutdown included), so the new instance's
saves can't be overwritten;
3. answers `MIGRATE_STATUS` `PREPARED`; master starts the new instance and the move goes on as above.
If the move is cancelled before anyone reached the new instance, the property is unfrozen.
## Messages (appended to `MessageType::Master`)
| Message | Direction | Payload |
| --- | --- | --- |
| `MIGRATE_PREPARE` | master → property world | `MigratePrepare` (migration ID, max wait) |
| `LIVE_UPDATE_REQUEST` | dashboard / world (GM) → master | `LiveUpdateRequest` (start, cancel, status; warn seconds; who) |
| `LIVE_UPDATE_STATUS` | master → dashboard; → worlds for the GM who asked | `LiveUpdateStatus` (phase, every row) |
| `LIVE_UPDATE_RETIRE` | master → chat, UGC | none |
| `CHAT_HANDOFF` | chat → master → next chat | `ChatHandoff` (teams) |
| `CHAT_SERVER_READY` | master → worlds | none |
Changed, compatibly: `MigrationStatus` states `PREPARING` and `PREPARED` (appended), `MigratePlayersOrder.maxWaitSeconds`
and `CarriedPlayerState` position (appended, read only when present), `ChatPackets::LoginSessionNotify.resync` (written
only when set), `ServerListResponse` state `DRAINING` (appended).
Master is not replaced, so the master ↔ server messages of the running master must still be understood by the new
binaries: add fields at the end and read them only when present. A change master itself needs takes a normal restart.
## Settings (`masterconfig.ini`, read when an update starts)
| Setting | Default | |
| --- | --- | --- |
| `live_update_warn_seconds` | 10 | Warning before players are moved (0-300) |
| `live_update_parallel_worlds` | 4 | World instances replaced at once |
| `live_update_player_wait` | 30 | Dead or building players, seconds |
| `live_update_property_build_wait` | 60 | Property builders, seconds |
| `live_update_char_select_wait` | 60 | Character selection, seconds |
| `live_update_activity_wait` | 1800 | Activity zones, seconds |
| `live_update_ugc_drain_timeout` | 300 | UGC server finishing its jobs, seconds |
| `live_update_service_timeout` | 30 | A server stopping or coming back, seconds (3 starts) |
| `live_update_run_migrations` | 1 | Run database migrations first |
`prestart_worlds` decides which zones always keep an instance.
## What a player sees
* In a world: the Mythran Maintenance Alert, a loading screen of the same zone, "Mythran Dimensional Shift Succeeded!",
standing where they were. Chat history, open windows, team, friends and pet stay.
* On a property: the same; builders are told building ends first and leave build mode.
* In a race or minigame: nothing until it ends (they leave normally); only after `live_update_activity_wait` are they
moved, losing the activity.
* At character selection: nothing, unless still there after `live_update_char_select_wait`; then a reconnect to the new
character selection, which lists their characters again.
* Logging in: a few seconds in which auth doesn't answer (the client reports a connection error; logging in again works).
* Friends and whispers: a few seconds without them while chat restarts.
## Limitations
* **Master is not updated.** A new master would need every server to survive master's absence and re-register (instance
table, clone/private/password, caps, player counts, session keys). Worlds shut down after 5 s without master, and the
session keys of logged-in accounts live only in master. Updating master takes a normal restart.
* A loading screen is always shown; the experimental seamless mode of instance migrations is not used here.
* Each world's simulation starts fresh: enemies, smashables, quick builds, dropped loot and scripted events restart.
* Lost when moved: an open trade (cancelled first), build mode in progress (players wait for it first), possession and
mounts, an activity lobby, anything else not in the saved character.
* A BBB model not saved when a property's builders are taken out of build mode is lost (the BBB autosave remains).
* Chat: messages, whispers and team invites sent in the seconds chat is down are lost. A player who logs out while chat is
down stays in their team until it next changes.
* Auth: no second auth process on the same port (RakNet binds the port exclusively); logins pause while it restarts.
* Players already on their way to an old instance when it starts draining are moved once they arrive; one arriving after
the old instance finished is disconnected when it shuts down (and logs in again).
* A world whose migration fails keeps running the old build; start another update (only instances started before it
began are replaced) or shut it down from the dashboard.
* The new build's database migrations run while the old servers still run; a migration that breaks the old code breaks
them until they are replaced.
* Transfers were not tested with the game client when this was written.
## Reloading one zone
When only zone files changed (not the binaries), master replaces just the instances that loaded them, with the same
per-instance moves: see [WorldHotReload.md](WorldHotReload.md).