mirror of
https://github.com/DarkflameUniverse/DarkflameServer.git
synced 2026-10-02 02:43:44 +00:00
212 lines
15 KiB
Markdown
212 lines
15 KiB
Markdown
# Live updates
|
|
|
|
Moving every running server onto a new build of the server binaries without taking the server down. Players are moved
|
|
to new world instances started from the new binaries; the other servers restart one by one. Master keeps running.
|
|
Properties are the exception: they are never moved, and update once everyone has left them.
|
|
|
|
Code: `dMasterServer/LiveUpdateMachine.h` (the order, no master state; unit tested), `dMasterServer/LiveUpdateCoordinator`
|
|
(master's glue), `dMasterServer/OutdatedInstances.h` (instances on the old build: routing, property reminders, stopping
|
|
when empty; unit tested), `dMasterServer/MigrationCoordinator` and `dGame/dUtilities/WorldMigration` (moving one instance's
|
|
players, see [SeamlessTransfer.md](SeamlessTransfer.md)), `dNet/master/LiveUpdate.h` (messages).
|
|
|
|
## Starting one
|
|
|
|
Put the new build in place (the binaries master starts are the ones in its own directory), then:
|
|
|
|
| Trigger | Who |
|
|
| --- | --- |
|
|
| Dashboard, home page, **Live update** card | `server_live_update` (default GM 9) |
|
|
| `/liveupdate start [warn seconds]`, `/liveupdate cancel`, `/liveupdate status` | GM 9 (paired with `server_live_update`) |
|
|
| `kill -USR2 <master pid>` (not on Windows) | whoever can signal master |
|
|
|
|
Nothing happens without one of these. Master refuses when one is already running, when it is shutting down, or when
|
|
`WorldServer`, `AuthServer`, `ChatServer` (and `DashboardServer` / `UgcServer` when enabled) are missing or empty in
|
|
the binary directory (a build still being written).
|
|
|
|
Cancelling starts nothing new; what is under way finishes.
|
|
|
|
## Sequence
|
|
|
|
1. **Database**: the new build's migrations (`MigrationRunner::RunMigrations`, `RunSQLiteMigrations`). Failure stops the
|
|
update before anything else is touched. The running servers must cope with the new schema until they are replaced.
|
|
Once the database is up to date, master marks every world instance that was running when the update started
|
|
**outdated** (see [Old instances](#old-instances)): from then on nobody new is sent to one.
|
|
2. **UGC server, auth, chat** (together):
|
|
* UGC: `LIVE_UPDATE_RETIRE`. It drops its queue (the rows stay pending in the database), finishes and records the jobs
|
|
it is running, then exits. After `live_update_ugc_drain_timeout` it gets `SHUTDOWN` (running jobs are made again).
|
|
* Auth: `SHUTDOWN`. Logins fail until the new one is up (a few seconds); players already in game are not affected.
|
|
* Chat: `LIVE_UPDATE_RETIRE`. It sends its teams to master (`CHAT_HANDOFF`) and exits without logging anyone out.
|
|
* Master starts the new process when the old one disconnects, as it always does. A server counts as replaced once a
|
|
new one connects. One that doesn't come back within `live_update_service_timeout` is started again (3 tries).
|
|
* When the new chat server connects, master gives it the teams, and once it is up sends `CHAT_SERVER_READY` to every
|
|
world: each connects at once and sends its loaded players again (`LoginSessionNotify` with `resync`). The new chat
|
|
server takes them over without logging a login, and reads their friends lists.
|
|
3. **World instances**, after chat is back and 3 s for the worlds to reconnect. Character selection first, then the
|
|
busiest; `live_update_parallel_worlds` at a time (waiting instances don't take a slot). Instances started after the
|
|
update began are already on the new build and are not touched. What happens to each is decided when its turn comes:
|
|
|
|
| Instance | Plan |
|
|
| --- | --- |
|
|
| Nobody there | Stopped. Zones in `prestart_worlds` (and character selection) get a new instance first; the old one stops once it is ready. |
|
|
| Public world with players | Replaced: a new instance starts, players are moved, the old one stops. |
|
|
| Property (clone) | Not in the plan: never moved. It stays on the old build until everyone left, then stops (see [Properties](#properties)). |
|
|
| Private instance | Replaced by a new private instance with the same password. |
|
|
| Activity zone (any `Activities.instanceMapID`: races, minigames) | Draining: nobody new goes there; its players finish. After `live_update_activity_wait` whoever is left is moved to a new instance (the activity is lost). |
|
|
| Character selection | A new one starts at once and takes all logins. The old one drains; after `live_update_char_select_wait` whoever is still there is moved to the new one. |
|
|
|
|
4. **Dashboard**, last: `SHUTDOWN`; master starts the new one. Sessions survive (JWT, secret in `dashboard_jwt_secret`
|
|
or `jwt_secret`). The new dashboard asks master for the status when it connects.
|
|
|
|
## States
|
|
|
|
A world: `pending` → (`preparing`: property being saved) → `starting` (new instance launching) → `ready` (new instance
|
|
up, players warned) → `draining` (players being moved) → `stopping` (old instance shutting down) → `stopped`.
|
|
`waiting`: an activity zone or character selection waiting for its players to leave by themselves.
|
|
|
|
A server: `pending` → `stopping` (UGC: `draining`) → `starting` → `stopped` (chat: `ready` for 3 s in between).
|
|
|
|
`failed`: left as it was; a world keeps running on the old build. It stays outdated: it takes nobody new and stops once
|
|
empty (public instances of `prestart_worlds` zones excepted: shut those down from the dashboard). `skipped`: not running,
|
|
not enabled, or cancelled before its turn.
|
|
|
|
The update: `running` → `done` (possibly with failed rows), or `failed` (database migrations), `cancelling` →
|
|
`cancelled`.
|
|
|
|
The dashboard's world list shows instances being emptied as **Moving players** (`ServerListResponse` state `DRAINING`)
|
|
and outdated ones as **Draining (old version)** with how many players are still there (`ServerListResponse` per-instance
|
|
`outdated`, appended after the endpoints and read only when present).
|
|
Master logs every change of every row (`Live update N: ...`).
|
|
|
|
## Old instances
|
|
|
|
An instance started before the update (old binary) or on zone files that changed since ([WorldHotReload.md](WorldHotReload.md))
|
|
is **outdated** (`Instance::GetIsOutdated`, `InstanceView::outdated`).
|
|
|
|
* **Routing.** `InstanceManager::FindInstance` (`InstanceMigration::AcceptsNewPlayers`) skips draining and outdated
|
|
instances, so zone transfers, logins, property visits, friend and team joins go to new instances (started if needed,
|
|
waiting for them to be ready). A private instance's password finds its replacement (`FindPrivateInstance` skips
|
|
draining and outdated ones).
|
|
* **Stopping.** Master checks outdated instances once a second (`InstanceManager::UpdateOutdatedInstances`). One with
|
|
nobody in it and nobody on the way (no players, held seats, pending transfers or affirmations) is shut down
|
|
(`OutdatedInstances::ShouldStop`). Left to the update itself: instances being emptied (draining), character
|
|
selection, and public instances of `prestart_worlds` zones (they get their new instance first).
|
|
* **A zone's new instance.** When somebody went to a zone after the update began, master already started its new
|
|
instance; the update then just stops the zone's empty old ones instead of starting another.
|
|
|
|
### Properties
|
|
|
|
A property (any clone instance) is never replaced or moved, by a live update or a world reload: builders may have
|
|
work in progress that isn't saved, and moving them would lose it.
|
|
|
|
1. It is marked outdated like every other instance: nobody new goes there.
|
|
2. Its players get a server announcement (the popup and a chat line, `ANNOUNCE` sent by master to that world): "A
|
|
server update is available. This property keeps running on the old version until everyone has left it: leave and
|
|
come back to get the update. Nothing you built is lost." Once at first, then every 10 minutes while anyone is still
|
|
there (`OutdatedInstances::NoticeDue`).
|
|
3. When the last player leaves, it stops (and saves, as any world does when it shuts down).
|
|
|
|
**One instance per property.** A request for a property whose old instance is still running (still occupied, or
|
|
still shutting down) gets a new instance that waits: master adds it (instance ID, port) and queues the request on it,
|
|
but only starts its world server once the old instance has disconnected (`OutdatedInstances::MustWaitForOld`,
|
|
`InstanceManager::StartWaitingInstances`). The new world loads the property from the database after the old one saved
|
|
it for the last time, so two worlds never both save the same property's models. The visitor waits for that (their
|
|
transfer is answered when the new instance is ready); the dashboard shows the waiting instance as starting.
|
|
|
|
## Moving players
|
|
|
|
Each move is an instance migration (`MigrationCoordinator::Start` with `Options::liveUpdate`), see
|
|
[SeamlessTransfer.md](SeamlessTransfer.md):
|
|
|
|
* Players get the game's Mythran Maintenance Alert, then after `live_update_warn_seconds` (or the value picked for this
|
|
update) up to 10 a second are saved, locked and sent `TRANSFER_TO_WORLD` with the Mythran shift flag.
|
|
* Dead or building players wait up to `live_update_player_wait`, then go anyway. An open trade is cancelled.
|
|
* Where they stood is carried (`CarriedPlayerState` position) and applied when the new instance creates them, also on
|
|
properties and Moon Base where the saved character doesn't keep it. The pet that was out is summoned again.
|
|
* Character selection has no characters loaded: its users are just sent to the new one, which sends them their
|
|
characters (no maintenance notice).
|
|
|
|
### Saving and freezing a property (`MIGRATE_PREPARE`)
|
|
|
|
Still in the protocol, no longer sent by live updates or world reloads (properties are not moved). When sent, the old
|
|
instance tells builders building ends in N seconds, waits up to the given time, takes anyone still building out of
|
|
build mode, saves the property and freezes it (never saved again), and answers `MIGRATE_STATUS` `PREPARED`. A cancelled
|
|
move unfreezes it.
|
|
|
|
## Messages (appended to `MessageType::Master`)
|
|
|
|
| Message | Direction | Payload |
|
|
| --- | --- | --- |
|
|
| `MIGRATE_PREPARE` | master → property world | `MigratePrepare` (migration ID, max wait); not sent by live updates any more |
|
|
| `ANNOUNCE` (existing) | master → an outdated property's world | `Announcement` (the update reminder) |
|
|
| `LIVE_UPDATE_REQUEST` | dashboard / world (GM) → master | `LiveUpdateRequest` (start, cancel, status; warn seconds; who) |
|
|
| `LIVE_UPDATE_STATUS` | master → dashboard; → worlds for the GM who asked | `LiveUpdateStatus` (phase, every row) |
|
|
| `LIVE_UPDATE_RETIRE` | master → chat, UGC | none |
|
|
| `CHAT_HANDOFF` | chat → master → next chat | `ChatHandoff` (teams) |
|
|
| `CHAT_SERVER_READY` | master → worlds | none |
|
|
|
|
Changed, compatibly: `MigrationStatus` states `PREPARING` and `PREPARED` (appended), `MigratePlayersOrder.maxWaitSeconds`
|
|
and `CarriedPlayerState` position (appended, read only when present), `ChatPackets::LoginSessionNotify.resync` (written
|
|
only when set), `ServerListResponse` state `DRAINING` (appended) and per-instance `outdated` (after the endpoints, read
|
|
only when present), `WorldFilesStatus` instance flag `4` (outdated).
|
|
|
|
Master is not replaced, so the master ↔ server messages of the running master must still be understood by the new
|
|
binaries: add fields at the end and read them only when present. A change master itself needs takes a normal restart.
|
|
|
|
## Settings (`masterconfig.ini`, read when an update starts)
|
|
|
|
| Setting | Default | |
|
|
| --- | --- | --- |
|
|
| `live_update_warn_seconds` | 10 | Warning before players are moved (0-300) |
|
|
| `live_update_parallel_worlds` | 4 | World instances replaced at once |
|
|
| `live_update_player_wait` | 30 | Dead or building players, seconds |
|
|
| `live_update_property_build_wait` | 60 | Property builders, seconds (only used by `MIGRATE_PREPARE`, which live updates no longer send) |
|
|
| `live_update_char_select_wait` | 60 | Character selection, seconds |
|
|
| `live_update_activity_wait` | 1800 | Activity zones, seconds |
|
|
| `live_update_ugc_drain_timeout` | 300 | UGC server finishing its jobs, seconds |
|
|
| `live_update_service_timeout` | 30 | A server stopping or coming back, seconds (3 starts) |
|
|
| `live_update_run_migrations` | 1 | Run database migrations first |
|
|
|
|
`prestart_worlds` decides which zones always keep an instance.
|
|
|
|
## What a player sees
|
|
|
|
* In a world: the Mythran Maintenance Alert, a loading screen of the same zone, "Mythran Dimensional Shift Succeeded!",
|
|
standing where they were. Chat history, open windows, team, friends and pet stay.
|
|
* On a property: nothing changes; the server announcement "A server update is available…" at once and every 10
|
|
minutes. Leaving and coming back once nobody is left there puts them on the new build.
|
|
* Visiting a property whose old instance still has players: the transfer waits until those players left (the visitor
|
|
stays where they are meanwhile).
|
|
* In a race or minigame: nothing until it ends (they leave normally); only after `live_update_activity_wait` are they
|
|
moved, losing the activity.
|
|
* At character selection: nothing, unless still there after `live_update_char_select_wait`; then a reconnect to the new
|
|
character selection, which lists their characters again.
|
|
* Logging in: a few seconds in which auth doesn't answer (the client reports a connection error; logging in again works).
|
|
* Friends and whispers: a few seconds without them while chat restarts.
|
|
|
|
## Limitations
|
|
|
|
* **Master is not updated.** A new master would need every server to survive master's absence and re-register (instance
|
|
table, clone/private/password, caps, player counts, session keys). Worlds shut down after 5 s without master, and the
|
|
session keys of logged-in accounts live only in master. Updating master takes a normal restart.
|
|
* A loading screen is always shown; the experimental seamless mode of instance migrations is not used here.
|
|
* Each world's simulation starts fresh: enemies, smashables, quick builds, dropped loot and scripted events restart.
|
|
* Lost when moved: an open trade (cancelled first), build mode in progress (players wait for it first), possession and
|
|
mounts, an activity lobby, anything else not in the saved character.
|
|
* A property with somebody on it keeps the old build as long as they stay; visitors wait for it to empty (no timeout,
|
|
no message to the visitor while they wait).
|
|
* Chat: messages, whispers and team invites sent in the seconds chat is down are lost. A player who logs out while chat is
|
|
down stays in their team until it next changes.
|
|
* Auth: no second auth process on the same port (RakNet binds the port exclusively); logins pause while it restarts.
|
|
* Players already on their way to an old instance when it starts draining are moved once they arrive; one arriving after
|
|
the old instance finished is disconnected when it shuts down (and logs in again).
|
|
* A world whose migration fails keeps running the old build; start another update (only instances started before it
|
|
began are replaced) or shut it down from the dashboard.
|
|
* The new build's database migrations run while the old servers still run; a migration that breaks the old code breaks
|
|
them until they are replaced.
|
|
* Transfers were not tested with the game client when this was written.
|
|
|
|
## Reloading one zone
|
|
|
|
When only zone files changed (not the binaries), master replaces just the instances that loaded them, with the same
|
|
per-instance moves: see [WorldHotReload.md](WorldHotReload.md).
|