Files
DarkflameServer/docs/LiveUpdate.md
2026-09-30 04:29:54 -05:00

15 KiB

Live updates

Moving every running server onto a new build of the server binaries without taking the server down. Players are moved to new world instances started from the new binaries; the other servers restart one by one. Master keeps running. Properties are the exception: they are never moved, and update once everyone has left them.

Code: dMasterServer/LiveUpdateMachine.h (the order, no master state; unit tested), dMasterServer/LiveUpdateCoordinator (master's glue), dMasterServer/OutdatedInstances.h (instances on the old build: routing, property reminders, stopping when empty; unit tested), dMasterServer/MigrationCoordinator and dGame/dUtilities/WorldMigration (moving one instance's players, see SeamlessTransfer.md), dNet/master/LiveUpdate.h (messages).

Starting one

Put the new build in place (the binaries master starts are the ones in its own directory), then:

Trigger Who
Dashboard, home page, Live update card server_live_update (default GM 9)
/liveupdate start [warn seconds], /liveupdate cancel, /liveupdate status GM 9 (paired with server_live_update)
kill -USR2 <master pid> (not on Windows) whoever can signal master

Nothing happens without one of these. Master refuses when one is already running, when it is shutting down, or when WorldServer, AuthServer, ChatServer (and DashboardServer / UgcServer when enabled) are missing or empty in the binary directory (a build still being written).

Cancelling starts nothing new; what is under way finishes.

Sequence

  1. Database: the new build's migrations (MigrationRunner::RunMigrations, RunSQLiteMigrations). Failure stops the update before anything else is touched. The running servers must cope with the new schema until they are replaced. Once the database is up to date, master marks every world instance that was running when the update started outdated (see Old instances): from then on nobody new is sent to one.

  2. UGC server, auth, chat (together):

    • UGC: LIVE_UPDATE_RETIRE. It drops its queue (the rows stay pending in the database), finishes and records the jobs it is running, then exits. After live_update_ugc_drain_timeout it gets SHUTDOWN (running jobs are made again).
    • Auth: SHUTDOWN. Logins fail until the new one is up (a few seconds); players already in game are not affected.
    • Chat: LIVE_UPDATE_RETIRE. It sends its teams to master (CHAT_HANDOFF) and exits without logging anyone out.
    • Master starts the new process when the old one disconnects, as it always does. A server counts as replaced once a new one connects. One that doesn't come back within live_update_service_timeout is started again (3 tries).
    • When the new chat server connects, master gives it the teams, and once it is up sends CHAT_SERVER_READY to every world: each connects at once and sends its loaded players again (LoginSessionNotify with resync). The new chat server takes them over without logging a login, and reads their friends lists.
  3. World instances, after chat is back and 3 s for the worlds to reconnect. Character selection first, then the busiest; live_update_parallel_worlds at a time (waiting instances don't take a slot). Instances started after the update began are already on the new build and are not touched. What happens to each is decided when its turn comes:

    Instance Plan
    Nobody there Stopped. Zones in prestart_worlds (and character selection) get a new instance first; the old one stops once it is ready.
    Public world with players Replaced: a new instance starts, players are moved, the old one stops.
    Property (clone) Not in the plan: never moved. It stays on the old build until everyone left, then stops (see Properties).
    Private instance Replaced by a new private instance with the same password.
    Activity zone (any Activities.instanceMapID: races, minigames) Draining: nobody new goes there; its players finish. After live_update_activity_wait whoever is left is moved to a new instance (the activity is lost).
    Character selection A new one starts at once and takes all logins. The old one drains; after live_update_char_select_wait whoever is still there is moved to the new one.
  4. Dashboard, last: SHUTDOWN; master starts the new one. Sessions survive (JWT, secret in dashboard_jwt_secret or jwt_secret). The new dashboard asks master for the status when it connects.

States

A world: pending → (preparing: property being saved) → starting (new instance launching) → ready (new instance up, players warned) → draining (players being moved) → stopping (old instance shutting down) → stopped. waiting: an activity zone or character selection waiting for its players to leave by themselves.

A server: pending → stopping (UGC: draining) → starting → stopped (chat: ready for 3 s in between).

failed: left as it was; a world keeps running on the old build. It stays outdated: it takes nobody new and stops once empty (public instances of prestart_worlds zones excepted: shut those down from the dashboard). skipped: not running, not enabled, or cancelled before its turn.

The update: running → done (possibly with failed rows), or failed (database migrations), cancelling → cancelled.

The dashboard's world list shows instances being emptied as Moving players (ServerListResponse state DRAINING) and outdated ones as Draining (old version) with how many players are still there (ServerListResponse per-instance outdated, appended after the endpoints and read only when present). Master logs every change of every row (Live update N: ...).

Old instances

An instance started before the update (old binary) or on zone files that changed since (WorldHotReload.md) is outdated (Instance::GetIsOutdated, InstanceView::outdated).

  • Routing. InstanceManager::FindInstance (InstanceMigration::AcceptsNewPlayers) skips draining and outdated instances, so zone transfers, logins, property visits, friend and team joins go to new instances (started if needed, waiting for them to be ready). A private instance's password finds its replacement (FindPrivateInstance skips draining and outdated ones).
  • Stopping. Master checks outdated instances once a second (InstanceManager::UpdateOutdatedInstances). One with nobody in it and nobody on the way (no players, held seats, pending transfers or affirmations) is shut down (OutdatedInstances::ShouldStop). Left to the update itself: instances being emptied (draining), character selection, and public instances of prestart_worlds zones (they get their new instance first).
  • A zone's new instance. When somebody went to a zone after the update began, master already started its new instance; the update then just stops the zone's empty old ones instead of starting another.

Properties

A property (any clone instance) is never replaced or moved, by a live update or a world reload: builders may have work in progress that isn't saved, and moving them would lose it.

  1. It is marked outdated like every other instance: nobody new goes there.
  2. Its players get a server announcement (the popup and a chat line, ANNOUNCE sent by master to that world): "A server update is available. This property keeps running on the old version until everyone has left it: leave and come back to get the update. Nothing you built is lost." Once at first, then every 10 minutes while anyone is still there (OutdatedInstances::NoticeDue).
  3. When the last player leaves, it stops (and saves, as any world does when it shuts down).

One instance per property. A request for a property whose old instance is still running (still occupied, or still shutting down) gets a new instance that waits: master adds it (instance ID, port) and queues the request on it, but only starts its world server once the old instance has disconnected (OutdatedInstances::MustWaitForOld, InstanceManager::StartWaitingInstances). The new world loads the property from the database after the old one saved it for the last time, so two worlds never both save the same property's models. The visitor waits for that (their transfer is answered when the new instance is ready); the dashboard shows the waiting instance as starting.

Moving players

Each move is an instance migration (MigrationCoordinator::Start with Options::liveUpdate), see SeamlessTransfer.md:

  • Players get the game's Mythran Maintenance Alert, then after live_update_warn_seconds (or the value picked for this update) up to 10 a second are saved, locked and sent TRANSFER_TO_WORLD with the Mythran shift flag.
  • Dead or building players wait up to live_update_player_wait, then go anyway. An open trade is cancelled.
  • Where they stood is carried (CarriedPlayerState position) and applied when the new instance creates them, also on properties and Moon Base where the saved character doesn't keep it. The pet that was out is summoned again.
  • Character selection has no characters loaded: its users are just sent to the new one, which sends them their characters (no maintenance notice).

Saving and freezing a property (MIGRATE_PREPARE)

Still in the protocol, no longer sent by live updates or world reloads (properties are not moved). When sent, the old instance tells builders building ends in N seconds, waits up to the given time, takes anyone still building out of build mode, saves the property and freezes it (never saved again), and answers MIGRATE_STATUS PREPARED. A cancelled move unfreezes it.

Messages (appended to MessageType::Master)

Message Direction Payload
MIGRATE_PREPARE master → property world MigratePrepare (migration ID, max wait); not sent by live updates any more
ANNOUNCE (existing) master → an outdated property's world Announcement (the update reminder)
LIVE_UPDATE_REQUEST dashboard / world (GM) → master LiveUpdateRequest (start, cancel, status; warn seconds; who)
LIVE_UPDATE_STATUS master → dashboard; → worlds for the GM who asked LiveUpdateStatus (phase, every row)
LIVE_UPDATE_RETIRE master → chat, UGC none
CHAT_HANDOFF chat → master → next chat ChatHandoff (teams)
CHAT_SERVER_READY master → worlds none

Changed, compatibly: MigrationStatus states PREPARING and PREPARED (appended), MigratePlayersOrder.maxWaitSeconds and CarriedPlayerState position (appended, read only when present), ChatPackets::LoginSessionNotify.resync (written only when set), ServerListResponse state DRAINING (appended) and per-instance outdated (after the endpoints, read only when present), WorldFilesStatus instance flag 4 (outdated).

Master is not replaced, so the master ↔ server messages of the running master must still be understood by the new binaries: add fields at the end and read them only when present. A change master itself needs takes a normal restart.

Settings (masterconfig.ini, read when an update starts)

Setting Default
live_update_warn_seconds 10 Warning before players are moved (0-300)
live_update_parallel_worlds 4 World instances replaced at once
live_update_player_wait 30 Dead or building players, seconds
live_update_property_build_wait 60 Property builders, seconds (only used by MIGRATE_PREPARE, which live updates no longer send)
live_update_char_select_wait 60 Character selection, seconds
live_update_activity_wait 1800 Activity zones, seconds
live_update_ugc_drain_timeout 300 UGC server finishing its jobs, seconds
live_update_service_timeout 30 A server stopping or coming back, seconds (3 starts)
live_update_run_migrations 1 Run database migrations first

prestart_worlds decides which zones always keep an instance.

What a player sees

  • In a world: the Mythran Maintenance Alert, a loading screen of the same zone, "Mythran Dimensional Shift Succeeded!", standing where they were. Chat history, open windows, team, friends and pet stay.
  • On a property: nothing changes; the server announcement "A server update is available…" at once and every 10 minutes. Leaving and coming back once nobody is left there puts them on the new build.
  • Visiting a property whose old instance still has players: the transfer waits until those players left (the visitor stays where they are meanwhile).
  • In a race or minigame: nothing until it ends (they leave normally); only after live_update_activity_wait are they moved, losing the activity.
  • At character selection: nothing, unless still there after live_update_char_select_wait; then a reconnect to the new character selection, which lists their characters again.
  • Logging in: a few seconds in which auth doesn't answer (the client reports a connection error; logging in again works).
  • Friends and whispers: a few seconds without them while chat restarts.

Limitations

  • Master is not updated. A new master would need every server to survive master's absence and re-register (instance table, clone/private/password, caps, player counts, session keys). Worlds shut down after 5 s without master, and the session keys of logged-in accounts live only in master. Updating master takes a normal restart.
  • A loading screen is always shown; the experimental seamless mode of instance migrations is not used here.
  • Each world's simulation starts fresh: enemies, smashables, quick builds, dropped loot and scripted events restart.
  • Lost when moved: an open trade (cancelled first), build mode in progress (players wait for it first), possession and mounts, an activity lobby, anything else not in the saved character.
  • A property with somebody on it keeps the old build as long as they stay; visitors wait for it to empty (no timeout, no message to the visitor while they wait).
  • Chat: messages, whispers and team invites sent in the seconds chat is down are lost. A player who logs out while chat is down stays in their team until it next changes.
  • Auth: no second auth process on the same port (RakNet binds the port exclusively); logins pause while it restarts.
  • Players already on their way to an old instance when it starts draining are moved once they arrive; one arriving after the old instance finished is disconnected when it shuts down (and logs in again).
  • A world whose migration fails keeps running the old build; start another update (only instances started before it began are replaced) or shut it down from the dashboard.
  • The new build's database migrations run while the old servers still run; a migration that breaks the old code breaks them until they are replaced.
  • Transfers were not tested with the game client when this was written.

Reloading one zone

When only zone files changed (not the binaries), master replaces just the instances that loaded them, with the same per-instance moves: see WorldHotReload.md.