docs/LiveUpdate.md: how to start one, the order, each kind of instance, the states, routing, moving players and properties, messages, settings, what a player sees and the limits. SeamlessTransfer.md, Commands.md and Dashboard.md point to it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
12 KiB
Live updates
Moving every running server onto a new build of the server binaries without taking the server down. Players are moved to new world instances started from the new binaries; the other servers restart one by one. Master keeps running.
Code: dMasterServer/LiveUpdateMachine.h (the order, no master state; unit tested), dMasterServer/LiveUpdateCoordinator
(master's glue), dMasterServer/MigrationCoordinator and dGame/dUtilities/WorldMigration (moving one instance's
players, see SeamlessTransfer.md), dNet/master/LiveUpdate.h (messages).
Starting one
Put the new build in place (the binaries master starts are the ones in its own directory), then:
| Trigger | Who |
|---|---|
| Dashboard, home page, Live update card | server_live_update (default GM 9) |
/liveupdate start [warn seconds], /liveupdate cancel, /liveupdate status |
GM 9 (paired with server_live_update) |
kill -USR2 <master pid> (not on Windows) |
whoever can signal master |
Nothing happens without one of these. Master refuses when one is already running, when it is shutting down, or when
WorldServer, AuthServer, ChatServer (and DashboardServer / UgcServer when enabled) are missing or empty in
the binary directory (a build still being written).
Cancelling starts nothing new; what is under way finishes.
Sequence
-
Database: the new build's migrations (
MigrationRunner::RunMigrations,RunSQLiteMigrations). Failure stops the update before anything else is touched. The running servers must cope with the new schema until they are replaced. -
UGC server, auth, chat (together):
- UGC:
LIVE_UPDATE_RETIRE. It drops its queue (the rows stay pending in the database), finishes and records the jobs it is running, then exits. Afterlive_update_ugc_drain_timeoutit getsSHUTDOWN(running jobs are made again). - Auth:
SHUTDOWN. Logins fail until the new one is up (a few seconds); players already in game are not affected. - Chat:
LIVE_UPDATE_RETIRE. It sends its teams to master (CHAT_HANDOFF) and exits without logging anyone out. - Master starts the new process when the old one disconnects, as it always does. A server counts as replaced once a
new one connects. One that doesn't come back within
live_update_service_timeoutis started again (3 tries). - When the new chat server connects, master gives it the teams, and once it is up sends
CHAT_SERVER_READYto every world: each connects at once and sends its loaded players again (LoginSessionNotifywithresync). The new chat server takes them over without logging a login, and reads their friends lists.
- UGC:
-
World instances, after chat is back and 3 s for the worlds to reconnect. Character selection first, then the busiest;
live_update_parallel_worldsat a time (waiting instances don't take a slot). Instances started after the update began are already on the new build and are not touched. What happens to each is decided when its turn comes:Instance Plan Nobody there Stopped. Zones in prestart_worlds(and character selection) get a new instance first; the old one stops once it is ready.Public world with players Replaced: a new instance starts, players are moved, the old one stops. Property (clone) Saved and frozen first ( MIGRATE_PREPARE), then replaced as above.Private instance Replaced by a new private instance with the same password. Activity zone (any Activities.instanceMapID: races, minigames)Draining: nobody new goes there; its players finish. After live_update_activity_waitwhoever is left is moved to a new instance (the activity is lost).Character selection A new one starts at once and takes all logins. The old one drains; after live_update_char_select_waitwhoever is still there is moved to the new one. -
Dashboard, last:
SHUTDOWN; master starts the new one. Sessions survive (JWT, secret indashboard_jwt_secretorjwt_secret). The new dashboard asks master for the status when it connects.
States
A world: pending → (preparing: property being saved) → starting (new instance launching) → ready (new instance
up, players warned) → draining (players being moved) → stopping (old instance shutting down) → stopped.
waiting: an activity zone or character selection waiting for its players to leave by themselves.
A server: pending → stopping (UGC: draining) → starting → stopped (chat: ready for 3 s in between).
failed: left as it was; a world keeps running on the old build and takes players again (a property some players
already went to stays with the new instance). skipped: not running, not enabled, or cancelled before its turn.
The update: running → done (possibly with failed rows), or failed (database migrations), cancelling →
cancelled.
The dashboard's world list shows instances being emptied as Moving players (ServerListResponse state DRAINING).
Master logs every change of every row (Live update N: ...).
Routing during an update
InstanceManager::FindInstance (InstanceMigration::AcceptsNewPlayers) skips draining instances, so new zone requests,
logins and friend transfers go to new instances (started if needed, waiting for them to be ready). A private
instance's password finds its replacement (FindPrivateInstance skips draining ones). While a property is being saved
its visitors still go to the old instance, where nobody can build.
Moving players
Each move is an instance migration (MigrationCoordinator::Start with Options::liveUpdate), see
SeamlessTransfer.md:
- Players get the game's Mythran Maintenance Alert, then after
live_update_warn_seconds(or the value picked for this update) up to 10 a second are saved, locked and sentTRANSFER_TO_WORLDwith the Mythran shift flag. - Dead or building players wait up to
live_update_player_wait, then go anyway. An open trade is cancelled. - Where they stood is carried (
CarriedPlayerStateposition) and applied when the new instance creates them, also on properties and Moon Base where the saved character doesn't keep it. The pet that was out is summoned again. - Character selection has no characters loaded: its users are just sent to the new one, which sends them their characters (no maintenance notice).
Properties
The new instance loads the property from the database when it starts, so before it is started the old one:
- tells builders "building on this property ends in N seconds" and waits up to
live_update_property_build_waitfor nobody to be building; - takes anyone still building out of build mode, saves the property, and freezes it: nobody can build, place, pick up, claim or BBB-save there, and it is never saved again (disconnects and shutdown included), so the new instance's saves can't be overwritten;
- answers
MIGRATE_STATUSPREPARED; master starts the new instance and the move goes on as above.
If the move is cancelled before anyone reached the new instance, the property is unfrozen.
Messages (appended to MessageType::Master)
| Message | Direction | Payload |
|---|---|---|
MIGRATE_PREPARE |
master → property world | MigratePrepare (migration ID, max wait) |
LIVE_UPDATE_REQUEST |
dashboard / world (GM) → master | LiveUpdateRequest (start, cancel, status; warn seconds; who) |
LIVE_UPDATE_STATUS |
master → dashboard; → worlds for the GM who asked | LiveUpdateStatus (phase, every row) |
LIVE_UPDATE_RETIRE |
master → chat, UGC | none |
CHAT_HANDOFF |
chat → master → next chat | ChatHandoff (teams) |
CHAT_SERVER_READY |
master → worlds | none |
Changed, compatibly: MigrationStatus states PREPARING and PREPARED (appended), MigratePlayersOrder.maxWaitSeconds
and CarriedPlayerState position (appended, read only when present), ChatPackets::LoginSessionNotify.resync (written
only when set), ServerListResponse state DRAINING (appended).
Master is not replaced, so the master ↔ server messages of the running master must still be understood by the new binaries: add fields at the end and read them only when present. A change master itself needs takes a normal restart.
Settings (masterconfig.ini, read when an update starts)
| Setting | Default | |
|---|---|---|
live_update_warn_seconds |
10 | Warning before players are moved (0-300) |
live_update_parallel_worlds |
4 | World instances replaced at once |
live_update_player_wait |
30 | Dead or building players, seconds |
live_update_property_build_wait |
60 | Property builders, seconds |
live_update_char_select_wait |
60 | Character selection, seconds |
live_update_activity_wait |
1800 | Activity zones, seconds |
live_update_ugc_drain_timeout |
300 | UGC server finishing its jobs, seconds |
live_update_service_timeout |
30 | A server stopping or coming back, seconds (3 starts) |
live_update_run_migrations |
1 | Run database migrations first |
prestart_worlds decides which zones always keep an instance.
What a player sees
- In a world: the Mythran Maintenance Alert, a loading screen of the same zone, "Mythran Dimensional Shift Succeeded!", standing where they were. Chat history, open windows, team, friends and pet stay.
- On a property: the same; builders are told building ends first and leave build mode.
- In a race or minigame: nothing until it ends (they leave normally); only after
live_update_activity_waitare they moved, losing the activity. - At character selection: nothing, unless still there after
live_update_char_select_wait; then a reconnect to the new character selection, which lists their characters again. - Logging in: a few seconds in which auth doesn't answer (the client reports a connection error; logging in again works).
- Friends and whispers: a few seconds without them while chat restarts.
Limitations
- Master is not updated. A new master would need every server to survive master's absence and re-register (instance table, clone/private/password, caps, player counts, session keys). Worlds shut down after 5 s without master, and the session keys of logged-in accounts live only in master. Updating master takes a normal restart.
- A loading screen is always shown; the experimental seamless mode of instance migrations is not used here.
- Each world's simulation starts fresh: enemies, smashables, quick builds, dropped loot and scripted events restart.
- Lost when moved: an open trade (cancelled first), build mode in progress (players wait for it first), possession and mounts, an activity lobby, anything else not in the saved character.
- A BBB model not saved when a property's builders are taken out of build mode is lost (the BBB autosave remains).
- Chat: messages, whispers and team invites sent in the seconds chat is down are lost. A player who logs out while chat is down stays in their team until it next changes.
- Auth: no second auth process on the same port (RakNet binds the port exclusively); logins pause while it restarts.
- Players already on their way to an old instance when it starts draining are moved once they arrive; one arriving after the old instance finished is disconnected when it shuts down (and logs in again).
- A world whose migration fails keeps running the old build; start another update (only instances started before it began are replaced) or shut it down from the dashboard.
- The new build's database migrations run while the old servers still run; a migration that breaks the old code breaks them until they are replaced.
- Transfers were not tested with the game client when this was written.