Skip to content

Backend Architecture

Overview

The Python backend follows Cosmic Python ("Architecture Patterns with Python") adapted for a single-user application. Code is split into layers with a strictly enforced dependency direction:

  • services/ — orchestration. Business logic and the public callable surface.
  • adapters/ — I/O. Everything that touches the network, the filesystem, the clock, or Steam.
  • domain/ — pure compute. Functions in, values out; no I/O, no state mutation, no service/adapter imports.
  • lib/ — cross-cutting utilities independent of every other layer.
  • models/ — data shapes (TypedDicts, dataclasses) independent of every other layer.
  • host/ — the process itself. Not a layer of the application at all: see below.

Services depend on Protocols (defined in services/protocols/), never on concrete adapter classes. Adapters implement those Protocols. bootstrap/ is the composition root — the only place where concrete adapters meet services. main.py owns the process lifecycle and the callable surface; it holds no business logic.

class Plugin:
    # No base classes — pure composition
    # Owns: the lifecycle (_main / _unload) and the callable surface
    # Delegates: all business logic to services, all I/O to adapters

The host (backend/host/)

The backend runs in a process of its own — ADR-0036. host/ is everything that used to be a plugin loader's job: the single-instance lock, the start-up order, a loopback port, the panel bundle, and one WebSocket carrying calls and events. Standard library only.

It is deliberately not a layer of the application, and the boundary runs in both directions. host/ imports nothing from services, adapters, bootstrap or domain; and nothing but main.py imports host — both are .importlinter contracts. The second half is the one that would rot without a check: the first service that wanted to send an event would reach in here for the sink, and the composition root would stop being the only place that knows a transport exists.

Module Owns
protocol.py The four message kinds and the transport-reason vocabulary
access.py The three admission checks — Host, Origin, Token — in that order
dispatch.py Name resolution onto the plugin object, the exception boundary, the answer cap
connection.py One live WebSocket: frame reading, the frame cap, the heartbeat, calls in flight
events.py Where an event leaves the process, and whether anybody heard it
server.py The bind and its fallback, the file route, the upgrade, newest-connection-wins
single_instance.py The flock beside the database, and the port file beside the session
logging_setup.py The root logger, and the token redaction on its file handler
runtime.py The start-up order and the shutdown
status.py What the panel may ask about the run hosting it

Two pure halves live in lib/ instead, because they are checkable against tables of bytes rather than against a running server: lib/websocket_frames.py (the RFC 6455 codec) and lib/http_messages.py (request heads in, response heads out).

The protocol. JSON text frames, both directions, each message stating its kind as a readable word:

call   {type, id, method, args}     args are positional
reply  {type, id, result}
error  {type, id, reason, message, traceback?}
event  {type, name, payload}

error.reason is the transport layer — method_unknown, payload_too_large, backend_exception, malformed_message, connection_lost. A callable's own failure is a successful transport and arrives inside result in the {success, reason, message} shape scripts/check_failure_shape.py guards; that gate does not see host/, so keeping the two apart is prose and review. Reachable methods are exactly the endpoints on Plugin: the public methods marked @route, def or async def alike. An endpoint's answer is awaited only when it is awaitable. The set is the one scripts/check_callable_manifest.py derives, asserted equal by tests/host/test_dispatch.py.

Two size caps, two purposes. ~12 MiB on one call's encoded answer, refused as an ordinary error for that call alone; 16 MiB on the connection's frames, judged on the announced length before a byte is buffered, whose breach closes the socket. Breaking the second rejects every call in flight, which is why one oversized cover image may not reach it.

No reply store. A call whose answer was in flight when the socket went is not redelivered — its task is cancelled, and the caller's own pending register answers it connection_lost. A lost answer fails visibly; it never disappears.

Events leave through a seam main.py is handed, never a module-level call. Plugin._event_sink is a PluginEventSink — one emit(name, payload) that answers whether anybody heard. Nothing is buffered: an event with no panel attached is dropped with a log line, because every event this backend sends is a statement about now and "sync finished" delivered three hours later lands in a session that never started one.

The answer is what a prune claim on an event needs. Five events carry one, because their Steam-side work outlives the backend's: the funnel every service emits through, _emit_with_prune_continuation, attaches it to prune_complete and migration_relaunch_options; the library service takes it itself for sync_complete and sync_stale when it emits them, and the download service for a bound ROM's download_complete. What the frontend does under the two library-sync claims is Steam Non-Steam Shortcuts — Explicit cleanup of vanished versions. The funnel hands the sink's answer back, so a service's EventEmitter answers whether anybody heard too. A claim handed to a panel that is not there holds off every removed-game cleanup until it expires — so whichever side took it gives it straight back the moment the answer says nobody heard. One consequence follows and nothing checks it: an event carrying a claim is awaited, never scheduled as a task (the answer would arrive after the claim was handed out).

The start-up order is what makes the port file meaningful:

lock (with a short retry window)   ← beside the database, which is what it protects
  → schema migration + start-up routines
  → bind the port (27737, then the next free)
  → write the port file
  → migrate_legacy_credentials     ← the only start-up step with network I/O

So "the port file is there" means "the backend is ready", and there is no readiness flag for a call to wait on. A second backend never reaches the port fallback — the lock refused it first — so a fallback always means some other program holds the port. The port file is a hint; connecting to it is the proof.

A start-up routine's failure is counted, not fatal (bootstrap/startup.py). One edge binds two of them: prune_stale_installed_roms runs only after detect_retrodeck_path_change succeeded, because the prune reads the pending homes the detection writes.

Dependency Diagram

host/ (the process: lock, port, protocol, lifetime — imports none of the below)
    ↑ dispatches onto
main.py (Plugin — lifecycle + callable routing; the only importer of host/)
    ↓ calls
bootstrap/ (composition root: adapters.bootstrap() builds adapters, services.wire_services() builds services)
    ↓ creates
┌─────────────────────────────────────────────────────────┐
│ Adapters (own all I/O — implement Protocols)            │
│   RommHttpAdapter / RommApiAdapter — RomM REST          │
│   SteamConfigAdapter — Steam VDF, grid dir, Steam Input │
│   SteamGridDbAdapter / SgdbArtworkCacheAdapter — SGDB   │
│   PersistenceAdapter (+ persister adapters) — JSON I/O  │
│   SqliteUnitOfWork (+ repository adapters) — SQLite I/O │
│   CoverArtFileStore / DownloadFile                      │
│   FirmwareFile / MigrationFile / RomFile / SaveFile     │
│   RetroDeckPaths                                        │
│   AtlasCatalogueAdapter (ES-DE catalogue, vendored)     │
│   EsFindRulesAdapter (ES-DE es_find_rules.xml)          │
│   PlatformCoreReaderAdapter (settings platform_cores)   │
│   SystemClock / SystemUuidGen / AsyncioSleeper          │
│   HostnameAdapter / PathProbe / DebugLogger             │
└────────────────────────┬────────────────────────────────┘
                         │ injected via *ServiceConfig
┌────────────────────────▼────────────────────────────────┐
│ Services (depend on Protocols, not concrete adapters)   │
│   LibraryService        SaveService                     │
│   DownloadService       PlaytimeService                 │
│   FirmwareService       SteamGridService                │
│   MetadataService       AchievementsService             │
│   MigrationService      GameDetailService               │
│   ArtworkService        RomRemovalService               │
│   ShortcutRemovalService  SettingsService               │
│   CoreService           ConnectionService               │
│   StartupHealingService LaunchGateService               │
│   SessionLifecycleService                               │
│   RomAdoptionService    RomInstallRecorder              │
└────────────────────────┬────────────────────────────────┘
                         │ depend on
┌────────────────────────▼────────────────────────────────┐
│ Protocols (services/protocols/) — grouped topically;    │
│   see Protocol Interfaces below                         │
└─────────────────────────────────────────────────────────┘

Domain (domain/) — pure compute, imported by services and adapters; imports nothing above it.

Arrow direction: depends-on (A -> B means A uses B).

The XxxServiceConfig constructor pattern

Every service takes a single config keyword argument — a frozen dataclass named <ServiceName>Config. All dependencies live in the config: Protocol-typed adapters, infrastructure seams (event loop, logger, Clock, UuidGen, Sleeper), persistence callbacks, and settings-derived values. There are no bare-param or mixed constructors.

sync_service = LibraryService(
    config=LibraryServiceConfig(
        romm_api=...,           # Protocol-typed adapter
        steam_config=...,       # Protocol-typed adapter
        clock=...,              # Clock Protocol
        uuid_gen=...,           # UuidGen Protocol
        sleeper=...,            # Sleeper Protocol
        uow_factory=...,        # UnitOfWorkFactory Protocol (roms / sync_runs / kv_config / rom_metadata)
        artwork=...,            # cross-service Protocol-typed peer
        # ...
    ),
)

Outer services keep the Service token in both names (SteamGridService + SteamGridServiceConfig). Sub-services may use role-based names without the token when it reads more naturally (SyncEngine + SyncEngineConfig, SyncOrchestrator + SyncOrchestratorConfig).

Module Responsibilities

Services (backend/services/)

Six services are packages of sub-services — library/, saves/, firmware/, prune/, rom_adoption/ and migration/; protocols/ is a package too, but it holds the Protocol interfaces rather than a service. The rest are single modules. A service over ~1000 LOC is the decomposition signal.

Module Domain
library/ LibraryService façade — fetch ROMs, preview/apply sync, per-unit shortcut delivery, roms/SyncRun writes + queries (decomposed; see below)
saves/ SaveService aggregate — .srm upload/download, conflict detection, slots, versions (decomposed; its sub-services are described on Save File Sync Architecture)
downloads.py DownloadService — ZIP extraction, M3U, progress, bounded-concurrency download queue (Semaphore(2) + reserved-bytes pre-flight); cancel/cleanup never deletes a live install. The occupancy pre-flight sits beside the disk-space one and refuses rather than writing over content the plugin did not place (DownloadTargetGateFn → RomAdoptionService)
rom_adoption/ RomAdoptionService (service.py) + AdoptionRenamer (renamer.py) over the shared Target (_target.py). The service owns the download-target gate (describe what occupies the path, refuse with the comparison, or clear it on an explicit replace), the candidate search behind it, adoption of on-disk content as an install, and the user-triggered content check against RomM's per-file checksums. The same filter answers the game-detail read through the AdoptionCandidateProbeFn seam — a boolean over the leaner list_top_level_names listing (no size-or-mtime stat per entry), stopping at the name match so the archive reads that rank candidates are skipped, and swallowing every failure into "no candidate" so an unreadable folder never makes a game look uninstallable. The two answer from different knowledge on purpose — a roms row against the server payload — and are not held to agreeing: they have diverged on the served shape, the platform folder, the matched name and the directory listing itself. So the click-time search re-runs and is the authority, and the button's promise rests on the last answer in its chain rather than on an invariant over two searches. Every entry is judged by what it is, without following it: a file, a directory or a symlink, and anything else — FIFO, socket, device node — has no kind. One function (_kind_of) decides that and every read of the filesystem asks it, listings and single-path description alike; re-deriving the rule per door is what left a named pipe adoptable at the target path after the listings already excluded it. The two reads differ only in what they do with a kindless entry — a listing drops it, describe_path reports it with kind: None, because something that is there must not come back as nothing for the next write to act on. A symlink is its own kind rather than whatever it resolves to, because claim_source refuses a symlink on every uninstall, so an install row pointing at one could never be removed. That chain, most specific first: adoption_candidates (served shape, adoptable kind), unusable_namesake (the other shape, or a link — one refusal, because the user's choice is the same), and candidate_vanished (the backstop: the page reported a copy, neither of the above applies, so the download stops rather than starting silently — which also covers the file deleted between page open and press). The same rule reaches the occupied-target dialog through describe_path, which answers existence with lstat: a link or a kindless entry at the ROM's own path is reported as occupying it and marked adoptable: False, which also stops the finalize os.replace from overwriting a dangling link in silence. adoptable is one predicate (adoptable_content) rather than a flag each site computes — the gate, the validation immediately before the move and the last check after it all ask it, so a candidate that turns into a link between the dialog opening and the user confirming is refused by the acting site and not only by a disabled button. Both the validation before the move and the last check after the sibling supersede are one function (_unadoptable_refusal), so the pair of answers cannot drift apart: content that vanished is nothing_to_adopt, content still sitting there that no install row may point at — the wrong shape, a link, or no kind at all — is unexpected_content_kind, because "the files are no longer there" said of a file that merely became a link sends the user looking for something that did not happen. Before any of it, the directory must be an ES-DE system (SystemKnownFn over the same es_systems.xml the accept-list comes from) — an unmapped RomM slug is otherwise taken verbatim as a directory name, and an empty accept-list would then remove the extension filter as well. When the target path is free the same gate lists the platform directory's top level (never descending — a single multi-file install can hold tens of thousands of files), keeps what ES-DE's live per-system accept-list admits, subtracts every path a rom_installs row accounts for, and matches on a tag-stripped normalized name (domain/rom_candidates.py); a match refuses the download with a short list ranked by free evidence — a single-member archive's central-directory CRC32 against RomM's file-level digest, an exact fs_size_bytes match, or the name alone — and states when the list was capped. Both exits of that dialog then go through AdoptionRenamer, which is why they cannot drift: Use These Files renames the ROM to the canonical name together with every save and savestate named after it, Download Instead removes the ROM (same is_safe_rom_path guard as the target-path replace) and carries only the saves — same plan (domain/adoption_rename.py), same collision question asked before the first file moves, same answer applied to the same whole set. A discarded multi-file candidate carries nothing, deliberately: its saves are named after a launch file that is still inside an unfetched archive. The move itself is AdoptionMoveStore's link-then-unlink with a rename-with-rollback fallback for a directory or a cross-filesystem set; an Overwrite's replaced targets — always saves or savestates, never the ROM — go through MatrixExecutor.quarantine_local_file rather than any delete of the adoption's own, so the .romm-backup discipline stays single-sourced (#965). A removal that fails after the carry names the saves that already moved, because that abort is not clean. The two companion directories are the resolver's answers — the save directory and the savestate directory, asked of the emulator the ROM launches with, for the launch file under its old name and under its new one (the new one need not exist; the resolver places a game by the path's own coordinates). A directory nothing could establish, and an emulator that keeps no savestates, carry nothing. Order is validate → carry → supersede → record. A directory's files are held to the place RomM names: RomFile.is_top_level compares rom.full_path against file_path, so the two are one coordinate system and the ROM-relative path is a subtraction — an entry the payload does not locate falls back to a by-name search, and one whose relative path escapes the ROM directory is refused rather than looked up. Anything whose name is one RomM reads as an archive is held to its contents instead, because RomM's digest for such a file describes what is inside (its current scanner accumulates over every member, the pre-4.9.0 one took the largest) while file_size_bytes stays the container's: neither the container's digest nor its size is compared, and the local side comes from the ZIP central directory (name, uncompressed size, CRC32, all free). One member inside — the same bytes under every rule RomM has used — is compared against the file-level digest; several members are compared one by one when the payload carries archive_members, and are unverifiable when it does not, since the single number is then either a composite or the largest member with nothing saying which. A size or CRC32 disagreement is reported without decompressing; the MD5 is what earns a match. A container this adapter cannot open leaves the entry unverifiable rather than accused — except a single-member archive whose member size the loose file on disk matches exactly, which is the unpacked case. Each difference is one {name, detail} line, and unverifiable carries a message per reason (no checksums published / could not be read / one checksum for the whole archive). Never deletes on its own initiative; both replace legs share one containment-guarded removal, and only the target-path leg keeps the carve-out that leaves a single file for os.replace to swap atomically. The candidate search is skipped on a resume — that decision was taken when the download was admitted, and the file the user declined must not refuse the transfer they started. Adoption runs the #1298 sibling supersede through the SiblingSupersedeFn seam (DownloadService owns the selection rule; ADR-0021 §5), ordered validate → supersede → record, so no other version is deleted until every refusal has been ruled out. See ADR-0028
rom_install_recorder.py RomInstallRecorder — the one writer of a rom_installs row and of the shortcut bake behind it (launchable verdict, the upsert, the fs_size_bytes write-back, the launch command resolved through the per-game core pin + disc pick, and the applied_launch_options memo). Both a completed download and an adoption reach an install through here, so an adopted row cannot drift from a downloaded one
firmware/ FirmwareService façade — joins the RomM firmware listing with what the installed emulators want (read live per platform through FirmwarePlatformResolver), downloads, per-core filtering; get_platform_firmware_status ships one platform's bios_level (unknown/ok/partial/missing via domain.bios_status.compute_bios_level) + required_count/required_downloaded/server_count/local_count/known_count/unknown_count/deletable_count/system_image so the platform detail reads the decision off the payload (#461), while get_firmware_status beside it names the platforms and pays no reading at all (decomposed; see below)
session_lifecycle.py SessionLifecycleService — post-exit orchestration (playtime + post-exit save sync + achievement sync + migration refresh)
migration/ MigrationService (service.py) — RetroDECK path-change detection + file migration — over FileMover (_moves.py), which moves one item to a path something may already occupy. A save directory that moves inside one home is not this package's: each game follows its own save directory the next time the plugin touches that game's saves (services/saves/save_directory.py). A home changed again before migrating accumulates a pending-home set (_previous + a _hops array) instead of overwriting the marker, and collection matches each tracked record by longest-prefix against every pending home, so files under an intermediate home are never stranded (#1042; pure transition/remap kernel in domain/migration_paths.py)
steamgrid.py SteamGridService — SteamGridDB fetch, cache, icons
artwork.py ArtworkService — cover art download into the per-ROM cover cache ({runtime}/covers/{rom_id}.png, the single source of truth for a ROM's cover), cover-cache invalidation against the persisted cover_source fingerprint (refresh_changed_covers, the per-unit pass that re-downloads changed covers and feeds the frontend's tile re-apply, #1386; a ts-only fingerprint change revalidates the cached bytes with a conditional request via the {rom_id}.cover-meta.json validator sidecar instead of re-downloading, #1454), publish of the active version's cache cover onto the Steam grid as {app_id}p.png (a copy so every sibling keeps its own file, ADR-0021 / #1346), the read-only get_artwork_base64 query, the cache-first fetch_cover_base64 used by the version picker (a cache miss downloads the ROM's cover from RomM — works for a server-only version with no local row), cache/staging cleanup + startup orphan pruning, the full grid-image sweep on shortcut removal (remove_artwork_files deletes every portrait/wide/hero/logo/icon × png/jpg/jpeg form for the removed appId, not just {app_id}p.png), and the user-triggered orphaned grid-image cleanup (cleanup_orphaned_grid_images, Data Management): deletes grid files whose parsed appId is in the non-Steam-shortcut range (domain/artwork_paths.is_shortcut_app_id, [0x80000000, 0xFFFFFFFF] — store-game custom art is never a candidate) and absent from the frontend's full live-shortcut scan (RomM-owned AND foreign), refusing outright (incomplete_scan) when any bound roms.shortcut_app_id is missing from the submitted live set — a provably partial scan deletes nothing; dry-run mode feeds the QAM's two-tap confirm count
game_detail.py GameDetailService — game detail page data aggregation. Network-free and bounded: for an uninstalled ROM it runs one stat on the computed target path and, only when that is free, one bare readdir through AdoptionCandidateProbeFn for a candidate elsewhere in the folder (#260) — no size-or-mtime stat per entry, since a name match reads neither. An install row answers both questions, so neither runs for an installed ROM. The whole read is offloaded to an executor by its callable: it opens a UoW and touches storage that may be asleep, and every game page opens it
game_process.py GameProcessService — the Stop Game ladder: find the RetroDECK flatpak's host processes, ask each to exit once, wait out a bounded grace window, force the survivors. Steam's TerminateApp cannot reach a portal-started flatpak, so the kill is backend-side; the never-re-request rule is a save-safety invariant (see State-aware Resume button)
playtime.py PlaytimeService — session recording into rom_playtime + the rom_playtime_sessions outbox, native play-session ingest (session-end fold+enqueue+flush to /api/play-sessions grouped per stored device_id with a bounded-retry quarantine + pull-only reconcile-on-view that restores total_seconds, session_count, and last_played via monotonic clamps, ADR-0018 / #903); owns the durable kv_config["playtime_scope_notice"] re-sign-in flag (a reconcile 403 sets it, a later 200 or a fresh sign-in clears it) surfaced via get_playtime_scope_notice
achievements.py AchievementsService — progress, caching, RA username
settings.py SettingsService — settings reads/writes, Steam Input config
data_inventory.py DataInventoryService — the Data Management page's own figures: the installs this device holds and the size RomM reported for them, read from the database, and the sealed recovery bundles and their size measured on disk
rom_removal.py RomRemovalService — ROM file deletion + rom_installs cleanup via the UoW; keeps the roms row, playtime, and saves per ADR-0007. Deletion runs under a descriptor-relative no-follow claim whose discipline follows the recovery bundle, not the caller: content-bound whenever a bundle exists to bind the hashes to, identity-only when none does — a user-initiated uninstall, or a cleanup run with recovery off. Only the identity-only form may adopt interrupted staging, and only it leases per unlink rather than over the whole tree — see Removed-game cleanup. One removal per rom_id at a time, across both entry points: a bulk uninstall reads its install list and claims every rom_id on the loop thread before dispatching its worker, so the single-ROM guard is what refuses a press during a bulk run, and an already-claimed ROM is what refuses the bulk run itself. Both refusals are in_progress; the bulk one carries no removal payload, which is the frontend's existing refusal discriminant. The claim set has a single writer — every mutation brackets the run_in_executor call instead of happening inside it — which is what keeps it lock-free, and services/ may not import threading anyway (.importlinter, no-stdlib-io-in-services). Start/finish logged with duration; per-file uninstall_progress frames for multi-file ROMs
prune/ PruneService façade — the explicit cleanup of local ROM entries RomM no longer has (decomposed; see Explicit removed-game cleanup)
cores.py CoreService — per-platform emulator-info lookup (get_platform_core_info — the classified emulators list + emulator_data_available), per-game core pin/clear (roms.emulator_override), per-platform core write (settings.json platform_cores) + fan-out re-bake; a pin's LABEL is validated against the bakeable options before any write and may name a standalone emulator or a libretro core; see Core and Emulator Selection
disc.py DiscService — the disc picker's two callables: get_disc_selection (whether a ROM is multi-disc, its launchable discs, the current and default target) and select_disc (pin a disc on the Rom aggregate or clear the pin, and return the re-baked launch command); see Multi-disc selection
disc_launch_resolver.py DiscLaunchResolver — the per-ROM read seam that folds the persisted disc pick over the live enumeration of the install's disc images and answers the path every launch bake writes; it never rewrites the install's file_path (see The read seam: DiscLaunchResolver)
version_switch.py VersionSwitchService — the game-detail version picker (ADR-0021). get_version_list(app_id) resolves the bound ROM's sibling group by roms.sibling_group_key (migration-010 index) and merges the local rows with RomM's live get_rom(bound).sibling_roms view — versions not yet synced are marked synced: false; returns each version's markers (active = the bound row, is_default = the resolution chain's pick with EMPTY install/binding filters, installed, and switchable) plus an additive server_query_failed flag when the server view can't be fetched. switchable is the SAME membership authority switch_version decides by (domain.sibling_group.target_in_sibling_group): a local target is a sibling_group_key match (the component keys encode membership, ADR-0022); a server-only target is judged by canonical compatibility — parse the bound key to {source}:{value} and require the target's id at that source to be absent-or-equal (the persisted key doubles as the group's canonical summary). A RomM sibling in a conflicting metadata match — a locally-synced ROM under a different sibling_group_key (#1359) or a never-synced ROM carrying a different id at the bound canonical source (#1360) — is listed but switchable: false and dropped from the default ranking, so the picker disables the row instead of offering a switch the backend would reject with not_in_group. A romm: bound key (no metadata source) admits no server-only target; a target that merely lacks the canonical id is in-group and adopts the bound group's key on switch (#1368 / ADR-0022). switch_version(app_id, target_rom_id, allow_stranded) checks the migration and prune rules at its entry (a removed-game cleanup's repoint calls switch_version_unchecked, which checks neither) and moves the group's binding to the target — a pure roms write (the repository's collision-unbind clears the old representative); no Steam call, no name change, no save migration (ADR-0021 §2/§4). Applies to downloaded games too (#1298): on success it returns target_installed + the target install's full launch_options ("" for an uninstalled target, the ADR-0009 placeholder) which the frontend confirm-writes onto the sticky shortcut. Switching away from a downloaded version whose local saves drift is a soft block (unsynced_saves, carrying server_reachable / unsynced_rom_id / unsynced_version_name) so the picker can offer sync-first / switch-anyway / cancel; allow_stranded=true is the "Switch anyway" override, honoured even offline. Guards, each canonical {success, reason, message}: not_found / download_in_progress (any group member has an active download — cancel first) / not_in_group / bound_elsewhere / unsynced_saves (soft) / server_unreachable / invalid_target (a server-only target that is a sibling but whose RomM detail the Rom aggregate rejects). The drift + reachability probes run in their own short reads before any write UoW opens; the write UoW re-checks membership + bound-elsewhere (TOCTOU) and the relaunch resolver re-bakes the command outside it (both open their own UoW; ADR-0006). A server-only target is persisted from its RomM detail first via domain.shortcut_data.extract_version_metadata, adopting the bound group's sibling_group_key (not its own coalesce-first key) so the next sync re-canonicalizes the whole component together (ADR-0022)
active_core_resolver.py ActiveCoreResolver — the single per-ROM read seam: active_core_for_rom(rom_id) folds the per-game DB override + the per-platform settings.json core over the live-es_systems.xml system default (three layers: per-game override → per-platform core → live es_systems default → None, the plain launch; no bundled snapshot). Each layer resolves a LABEL to a bakeable EmulatorInvocation (libretro or standalone) via get_emulator_options / label_to_invocation. Every per-game core read + every launch bake draws from it
relaunch_options_resolver.py RelaunchOptionsResolver — the one source of the current launch command for every installed and bound ROM (the RetroDECK-home migration's re-bake and the startup launch-options reconcile both draw from it) and of one ROM's bare launch target (the stop-game match); the rows are snapshotted in one short read UoW that closes before the bake resolves
shortcut_removal.py ShortcutRemovalService — shortcut removal; unbinds the ROM in roms (keeps the row per ADR-0007)
shortcut_relocation.py ShortcutRelocationService — which of our shortcuts name a launcher outside the launcher's home, handed to the frontend to repoint; a one-time transition with a recorded completion (see Repointing the shortcuts that already exist)
metadata.py MetadataService — ROM metadata reads from rom_metadata (7-day TTL): the per-ROM get_rom_metadata lookup, the paged get_metadata_cache_page(offset, limit) the frontend loads on plugin start (returns {items: {str(rom_id): entry}, total}, page and total read under one short read UoW), and the app_id → rom_id map the launcher resolves session ROMs through. Paging replaced a single whole-cache dump so a large library (5–10k ROMs) never pushes a multi-MB response as a single answer, which the host caps (host/dispatch.py) (#1025); the frontend pages at 500/page until total is reached
launch_gate.py LaunchGateService — pre-launch gate (rom lookup, install check, save status)
startup_healing.py StartupHealingService — prunes stale rom_installs rows against disk on load (via the UoW) + reconciles orphaned running SyncRuns (a hard crash leaves a running row → marked errored) + get_installed_relaunch_options() builds the startup launch-options reconcile items (see StartupHealingService notes)
leftover_tmp_cleanup.py LeftoverTmpCleanupService — removes the partial .tmp / .zip.tmp transfer files a stopped backend left under the ROM and BIOS directories, once at start-up
connection.py ConnectionService — connection test + RomM minimum-version gate + Client API Token lifecycle (mint/establish via credentials, validate/store a user-pasted token, or exchange a short-lived pairing code for a token — the latter two for OIDC accounts; host-bound to the minting origin; see ConnectionService notes)
update_check.py UpdateCheckService — whether a newer Tender release is out: the daily release check, the per-version dismissal and the daily-check switch, and Check now (see UpdateCheckService notes)
protocols/ Protocol interfaces grouped by concern (see Protocol Interfaces)

Version-list liveness

VersionSwitchService.get_version_list separates retained local membership from current RomM availability. On each lazy Game Page load, the bound detail request proves the bound id live and supplies the direct sibling ids; only non-bound local members absent from that direct sibling view are exact-id probed through RommRomReader.get_rom_once, the adapter's single-attempt short-timeout request_once path, concurrently on the existing executor fan-out. Only RommNotFoundError produces vanished: true. All other probe failures and malformed/falsy responses fail open. No result is persisted.

The response carries vanished on every version and bound_vanished even when multi_version is false. A singleton also carries its bound_version, which lets a definitively vanished synced binding expose scoped cleanup without inventing a multi-version picker. Retained vanished rows keep the independently-computed switchable membership verdict but are excluded from default ranking. A bound-id 404 is a successful entity verdict (bound_vanished: true, server_query_failed: false), not a global reachability signal. Artwork and save reads do not contribute availability evidence; fetch_cover_base64 remains a nullable data callable.

Version-switch target liveness

VersionSwitchService.switch_version obtains a fresh exact-id verdict immediately before every actual binding move onto an already-local target. The save-stranding guard runs first, so its initial soft block performs no target request; both follow-up choices (Sync now retry and Switch anyway) re-enter the service and receive the same liveness guard. allow_stranded bypasses only save stranding. The active-target no-op, invalid local context, active download, non-member target, and target bound to another shortcut all stop before the request because none can enter the binding write.

The local-target request uses RommRomReader.get_rom_once on the worker executor: one attempt, three-second timeout, no retry-progress event, and no open UoW. Only RommNotFoundError returns the canonical version_vanished refusal. Timeout, transport/DNS/SSL, authentication, server errors, and malformed or empty successful payloads are logged and fail open, preserving the fast offline switch path. The existing short write UoW still rechecks membership and bound-elsewhere after the request; the request narrows the race but cannot form a transaction across RomM and SQLite.

A server-only target still needs its full detail before it can be validated and persisted. That mandatory fetch keeps the normal retry/classification policy and is not preceded by a redundant exact-id probe; only its typed 404 is peeled into version_vanished. A refusal performs no binding, row/install/applied-launch-options write, launch-command resolution, event, cache invalidation, sync, or completion-stamp update.

Explicit removed-game cleanup (services/prune/)

Module Role
service.py Callable facade, atomic preview/run admission, and claimed frontend action leases
executor.py Serial per-group liveness, recovery, Steam-action, reconciliation, and finalization state machine
preview.py Local generation-gated candidate snapshot, complete affected-group disclosure, sizing, and paging
recovery.py Lossless aggregate snapshots, recovery artifact assembly, and sealed-state comparison
registry.py Short SQLite reads, action/final race validation, reconciliation, and final cascade delete
requests.py Preview/option decoding, bounded selections, and lossless bounded Steam snapshot validation

PruneService is the only path that deliberately deletes retained roms aggregate roots. A bulk preview is local-only: for each platform it requires a non-empty completed fetch generation and selects rows whose last_fetch_id differs, including NULL row generations. An inline preview may nominate one concrete retained ROM without generation evidence. Both forms return serialized-byte-budgeted pages plus an ephemeral fingerprint. A page may therefore contain fewer than the requested row limit, and the next offset advances by the rows actually returned. Pages include every member of an affected sibling group, with generation candidates marked separately, so whole-game deletion cannot reach an undisclosed row. start_prune consumes a finalized preview-bound installed-content selection. The frontend stages that selection in bounded pages, so wire bounds do not cap the total selected set. Start atomically refuses any registered conflicting callable and reserves the run before rebuilding the preview; concurrent starts cannot consume one token twice, and shutdown owns/cancels an admitted refresh before it can spawn a run. Each sync, download, migration, version-switch, save-write, session, uninstall, connection-identity change (including a successful connection test), or cache-mutation callable registers for its full lifetime before its first await. Detached status, download, and playtime tasks transfer the claim to their task lifetime. Core/disc writes, launch evaluation, and Steam Input application are included because they mutate recovered state. Frontend-owned shortcut removal, core/disc writes, version switches, uninstalls, SGDB/icon application, download completion, home migration, startup healing, pre-launch healing, and every post-sync Steam branch (launch options, collections, playtime, and overview metadata) hold tokenized conflict leases through their final Steam write and bounded release. Active continuations heartbeat those leases once per minute. A global frontend registry signals cooperative cancellation by component owner or on plugin dismount. Every backend wait captures the current plugin/owner mount generation before it starts; teardown tombstones that generation synchronously, so a lease-bearing response that arrives afterward is released without running its continuation. Only a genuine plugin/component remount opens a new generation. Cancellation stops every not-yet-started Steam mutation and lease renewal, but explicit backend release waits for any already-started non-cancellable Steam promise to settle. An unresolved operation stops renewing after a bounded five minutes; the backend's five-minute no-heartbeat expiry is the abandonment backstop if it never settles. Launch funnels carry the admission captured at the original Play action through every gate, modal, and launch-options confirmation wait. The version picker likewise rechecks its captured owner admission after save-sync and modal waits before any successor switch_version mutation, so an unmounted chain cannot resume under a new picker. Each non-empty sync_stale event carries its own lease through the paced removal tail; a later sync_complete lease overlaps and joins that same promise, so success composes both leases while a post-stale backend failure still leaves the tail covered. A terminal prune result that needs repoint publication likewise acquires its lease before event delivery while the old run is active; the frontend holds it across release acknowledgement and cover publication. Event delivery failure releases a token that never reached the frontend. This closes the reciprocal start/refusal race, and each path refuses while a prune claim is active. Migration and active-library-sync decorators additionally guard preview and start.

The executor processes sibling groups serially and catches ordinary exceptions per group. It rejects multiple shortcut bindings and active downloads, pins the preview's canonical RomM origin/token-origin/user namespace, probes every local group member with the single-attempt three-second get_rom_once, and treats only a typed RommNotFoundError from that same namespace as destructive authority. A connection/server/user change before, during, or after an exact-ID request is uncertain. Live, malformed, wrong-id, transport, authentication, timeout, server, and unknown outcomes retain data. Long recovery and frontend round trips are followed by another exact-ID proof before local source removal; every repoint re-proves both the vanished source and live target after the frontend action, including a vanished source that is not itself a generation candidate. The natural live repoint target uses the same filename-stem projection and resolve_group_representative ranking as the version picker with empty installed/bound preference sets. Repoint selection is independent of row deletion.

Recovery is coordinated through ordered ROM save locks, acquired with no UoW open. The lock set is retried until shared save ownership is stable. Short read UoWs close before filesystem work, and no RomM request, event emission, or frontend wait runs under those locks. The bundle is sealed before mutation. Aggregate, save-inventory, bundle/source, controller, and fresh frontend Steam state are validated before the irreversible Steam action. After that action the service reacquires the same recovery-set locks and repeats the guards immediately before local finalization; quarantine ownership is separately projected only for rows being deleted. The locks remain held through save quarantine, filesystem-only installed-content removal, and the final parent cascade. Source sets record expected presence or absence, the root's sealed no-follow identity and regular-file hash, and every descendant identity, mount ID, and regular-file hash. Every resolvable exclusive save path is represented even when it was absent. Traversal refuses nested mount transitions. After the anchored root rename, deletion obtains kernel read leases for every regular root/descendant before removing any entry and holds them through the last descriptor/hash check and unlink. An existing writer, an unsupported lease, or any inability to establish exclusion restores/retains the source rather than deleting bytes not in recovery. The controller VDF's held claimed inode uses the same writer exclusion from its final identity/hash validation through claim unlink on both normal publication and collision/rollback discard; setup failure preserves the claim and teardown failure after unlink is an ambiguous mutation. Content excluded from the bundle and recovery-off saves still receive a final complete no-follow claim; their remove/quarantine uses anchored parents and atomic no-replace rename rather than a raw path fallback, so a concurrently created backup is never overwritten. Writer-exclusion teardown faults after mutation are reported as ambiguous, not exact success. Controller rewrites revalidate the held claimed inode before every restore/discard branch and retain a newer claimed inode at a surfaced path when a concurrent Steam file wins publication. Every exclusive current-save path is expected absent after quarantine, regardless of whether it existed during inventory, and the whole set is collectively rechecked after all filesystem work and immediately before the aggregate cascade. Validation and claim decoding share one held bundle descriptor and return a digest that every later guard compares to the cached claims. Recovery root, staging, bundles, destination files, and the sealed bundle remain descriptor-anchored through copy, metadata application, hashing, cross-directory rename, validation, and both-parent fsync. Failure cleanup uses the same mount-aware descriptor-relative remover; if cleanup cannot prove a safe tree, it leaves and reports the full anchored staging path instead of recursively entering uncertain data. Adapter outcomes carry actual and durability-ambiguous changes into the mutation ledger even when a later item or parent fsync fails. The final SQLite delete is one short UoW that revalidates complete row/binding state and invalidates intersecting collection stamps; platform stamps are deliberately preserved (removed-game-cleanup.md).

Frontend Steam events use a claim/complete protocol. A claim checks the run, token, discriminant, appId, target, exact single-binding group, and current binding before the frontend rechecks the live tender-rom-launcher executable and mutates Steam. The lease is monotonic-clock bounded and rechecked after asynchronous validation; an identical repeat claim is idempotent, while mismatched or expired claims cannot authorize a mutation. Repoint commits through the normal version-switch authority first; shortcut removal is immediately reconciled to an unbound local row after Steam confirms absence. A claimed action whose completion is lost, or a RemoveShortcut/launch-options write that was attempted but could not be confirmed, is an explicit ambiguous partial. A pre-mutation refusal remains an ordinary failure. A later run can confirm a shortcut is already absent and reconcile without repeating Steam removal. A later guard failure is likewise an explicit partial result with the committed action and actual mutation categories recorded, not a claim that the group was unchanged. Cancellation can stop final guards and every later group; shielding starts only with the first irreversible local mutation and awaits that phase to a known result. Cancellation remains authoritative if that shielded child faults: the current group's truthful fault/ledger result is recorded and no later group starts. Terminal result strings/arrays are bounded and chunks are built to a serialized byte budget. Every action/progress/completion frame carries the originating preview ID, allowing the frontend to adopt a matching run even if the successful start response is delayed or lost while still rejecting foreign frames; completion finalizes only a contiguous chunk sequence. An accepted contiguous terminal sequence seals that run against every later action, progress, or completion frame, and the modal exposes its terminal controls immediately even if the start response is still pending. Committed repoint publication performs a bounded/retried backend release acknowledgement before gated cover/status work. Mixed runs continue unrelated groups.

LibraryService decomposition (services/library/)

The library sync subsystem is a façade over nine sub-services that coordinate through a shared LibrarySyncStateBox:

Module Role
service.py LibraryService façade — public callable surface; wires the sub-services and delegates
fetcher.py LibraryFetcher — read-only RomM roundtrips: list platforms/collections, the incremental/full pagination loop, per-unit work-queue construction. The outward half of the read pair below
local_library_reader.py LocalLibraryReader — the inward pair of fetcher.py: what THIS DEVICE recorded about the library, read back out of SQLite (the Rom rows and their bindings, the completion stamps, the persisted sibling keys, the last finished run), shaped into the projections a run decides against and the reachable set the collections listing counts against. Declared read-only (scripts/check_read_only_module.py)
sync_orchestrator.py SyncOrchestrator — preview (read-only), the per-unit apply pipeline (fetch → collapse → delta, then hand the delta to the dispatcher), cancel, the heartbeat clock, progress emission, and which terminal status a stopped run earns. Holds no ArtworkManager at all — the run's artwork seam is cover_preparer.py's and reporter.py's
chunk_dispatcher.py ChunkDispatcher — one unit's apply as durable chunk round-trips: emit sync_apply_unit, wait out the heartbeat clock for the ack, commit the chunk through the reporter, and stash or discard a chunk whose ack never came. It opens no transaction, and the run's try_begin_run / finish_run stay with the orchestrator
cover_preparer.py CoverPreparer — one unit's covers, readied before its shortcuts are emitted: the refresh_changed_covers invalidation pass, the download for the ROMs getting a shortcut, and each emitted entry's cover_path. Delegated to ArtworkService through the ArtworkManager seam, with the run's progress and cancel signals bound in
sync_run_recorder.py SyncRunRecorder — the run's own SyncRun row: the running row opened with the planned counts, and the single terminal transition (completed / cancelled / interrupted / paused / errored) that closes it, each in its own short write UoW. It writes the outcome; which terminal a stopped run earns stays the orchestrator's, and admitting or releasing the in-flight run stays _state.py's
reporter.py SyncReporter — post-apply finalisation (artwork filenames, the per-unit roms upsert, the stale unbind and the Steam-collection map) and the roms-derived queries, plus the sync_runs history reads behind the last-sync line and the run list
session_budget.py SessionBudgetMonitor — Steam's per-session renderer-heap budget: the GC-settled RSS measurement, the chunk-boundary pause verdict, the headroom clip that trims a chunk's additive cover work to what that verdict left, the post-preview prognosis, the run-start baseline, and the get_session_budget_status payload
shortcut_launch_resolver.py ShortcutLaunchResolver — resolves each ROM's launch facts for the shortcut bake: the disc-resolved installed path and the active emulator, both handed to build_shortcuts_data
_state.py LibrarySyncStateBox — shared mutable in-flight sync state; single source of truth threaded through every sub-service, and the sole owner of the run-lifecycle pair (sync_state / current_sync_id) and of the staged preview snapshot (pending_delta, TTL and the run-in-flight withholding included) via its verb methods

Two modules hold the artwork seam, and only two. cover_preparer.py holds it for the apply path's covers (the invalidation pass, the download, each emitted entry's cover_path); reporter.py holds it for commit-time cover-path finalisation (finalize_cover_path). Those are different questions asked at different points of a unit, which is why the confinement names a pair rather than one owner — scripts/check_seam_owner.py enforces the pair, so a third holder fails CI instead of quietly reopening the spread the cut closed.

The pipeline is split fetch (read-only) / apply (owns persistence): the fetcher never mutates the roms registry or rom_metadata, and the reporter's per-unit commit upserts each acked ROM's roms row and stamps its cached rom_metadata in the same write Unit of Work (Rom row first, then metadata — FK-safe). So a preview never mutates state, and an interrupted apply leaves only the units it already committed — incremental, per-unit delivery.

Where the synced-ROM state lives. The registry of synced ROMs, the last-sync timestamp, and the sync stats are SQLite, not JSON. The reporter upserts each acked ROM into the roms table via Rom.synced(...) / update_cover_path / assign_sgdb_id (artwork and steamgrid patch cover_path / sgdb_id on the same aggregate during the per-unit commit) and, in the same write UoW, stamps the ROM's cached rom_metadata (build_rom_metadata maps the live RomM metadatum — Rom row saved first so the rom_id FK holds); SyncRunRecorder writes the SyncRun row (start at apply-dispatch, complete / mark_cancelled / mark_interrupted / mark_paused / mark_errored at finalize), with the orchestrator deciding which of those terminals a stopped run earns. The terminal sync_complete event (and its progress frame) is emitted LAST — reporter.emit_sync_complete, called by the orchestrator only AFTER that terminal SyncRun status is persisted — so a frontend stats refetch triggered by the event reads the fresh run status instead of racing the write (the #39 lag); reporter.finalize_per_unit_run keeps only the stale-unbind + sync_collections emit and returns the app-id maps. sync_stats.roms is a registry-derived bound-shortcut count computed at read time (the ROMs still bound to a shortcut in roms, i.e. shortcut_app_id not NULL), not a stored scalar. The old JSON shortcut_registry / last_sync / sync_stats are gone from this path; all writes go through the roms / sync_runs Repository Protocols behind a narrow Unit of Work (per ADR-0006 the UoW spans only the DB write, never the up-to-60s frontend ack). The platform slug → display_name map resolves live from RomM each sync and is cached in a kv_config row for offline reads. Removing a shortcut unbinds the ROM (Rom.unbind_shortcut() NULLs shortcut_app_id, keeping the row and its per-ROM children) rather than deleting it, per ADR-0007. The full schema and aggregate model are in Database Design.

Delta-restricted apply: emit only new + changed (ADR-0025). Between the sibling-group collapse and chunking, _sync_one_unit runs domain/sync_diff.py::classify_roms over the unit's collapsed entries against the bound-shortcut registry and emits only the delta — new + changed (plus rebind entries, which always move their binding). "Changed" is content, not identity-only: on top of the identity triple (name / fs_name / platform_slug) the classifier compares each item's freshly built target launch_options against the applied_launch_options recorded on its roms row (the launch command last written to that shortcut). A match on both is content-unchanged and skipped — it never reaches the frontend, so no Set* walk and no ~2 s AppDetails confirm poll — while an install/uninstall or core/disc pin change that leaves identity untouched still flips the item to changed. A NULL recorded value (a pre-migration-015 row, or a freshly created row not yet recorded) is unknown and never skipped, so the first post-upgrade sync re-applies exactly as before, records values, and only later syncs skip. Skipped rows are still committed (chunking routes their groups to chunk 0's leftover), so no DB work is dropped and the per-platform completion stamp still rides the final chunk only — a platform is stamped exactly when its whole delta is durable, and an empty-delta platform emits one empty chunk that commits every row and writes the stamp. The recorded value is refreshed by six writer sites (sync ack-commit, download-complete, adopt-complete, uninstall, RetroDECK-home migration re-resolve, version switch), each recording the value the frontend just wrote onto the shortcut; a missed writer only ever causes a harmless spurious re-touch, never a wrong skip (the benign-failure asymmetry). The per-unit sync_apply_unit's unit_total is therefore the delta size, so the progress counter shows net progress and a resume converges quickly.

Cover-cache invalidation: the cover_source fingerprint (#1386). The per-ROM cover cache is only valid while the server's cover is unchanged, and the delta apply never re-downloads a skipped ROM's cover — so each roms row records a cover_source fingerprint: the full RomM cover source string (path_cover_large else path_cover_small, the embedded ?ts=… cache-buster included) whose bytes the cache file holds, compared as an opaque string (catches both a ts bump and a path change). Two compare layers run per unit, both in ArtworkService:

  • Apply-path covers (_resolve_cached_cover): a stored fingerprint that differs from the fresh one blocks both the cache-hit reuse and the grid→cache seed, forcing a fresh download.
  • The invalidation pass (refresh_changed_covers, reached through CoverPreparer from _sync_one_unit before the cover download): for every BOUND fetched ROM whose stored fingerprint differs — delta-skipped ROMs included — it re-downloads the cache file (atomic tmp+rename), republishes the grid {app_id}p.png copy, persists the new fingerprint in a small write UoW (an observed server fact, deliberately NOT ack-coupled like applied_launch_options), and collects {rom_id, app_id}. The pass scans against the unit's bound-row registry projection (LocalLibraryReader.do_read_apply_registry, which carries cover_source) — the same read the group collapse diffs against, so it opens no per-ROM DB lookups. That list rides the unit's first sync_apply_unit chunk as cover_refreshes, clipped to the session-budget headroom left after the chunk's own projected cost (each refresh ≈ one transient cover, COVER_TRANSIENT_KB) — clipped, never pausing the run: the grid files are already updated, a Steam restart shows the rest. The frontend re-applies each entry via SetCustomArtworkForApp at the 50 ms cadence before the chunk ack — without that push the tile stays stale in-session, since rt_custom_image_mtime only bumps through SetCustomArtworkForApp or a client restart.
  • The preview-side count (sync_preview → domain/cover_refresh.py::count_cover_refreshes): classify_roms is deliberately cover-blind (ADR-0025), so a cover-only server change yields an empty shortcut delta — and the QAM's preview flow used to short-circuit on "no changes", meaning the apply (and with it the invalidation pass) never ran and the fingerprint stayed stale forever. The preview therefore counts the fingerprint mismatches itself, with the SAME pure kernel the apply pass scans by (scan_cover_refresh_candidates — shared via domain/cover_refresh.py, so preview count and apply set cannot diverge), as an in-memory compare of the fetched union (platform and collection units alike) against the registry projection the classify already read — no extra DB pass, no downloads, no writes. The result rides the summary as the additive cover_refresh_count (absent/0 tolerated by old consumers); the frontend treats a cover-only preview (all diffs zero, count > 0) as apply-able, showing "No shortcut changes — N cover updates" with the normal Apply/Cancel confirm. NULL-fingerprint rows are not counted: their adopt-vs-download fork needs a cache-file check the preview must not perform, and the adopt is invisible to the user. The wholesale platform-skip gate needs no cover awareness — a cover change bumps the ROM's updated_at, which already forces the fetch; a skipped platform's registry-reconstructed thin ROMs carry no cover fields and count 0 by construction.
  • The empty-unit chunk vehicle: a cover-only unit has an empty apply delta, but build_unit_chunks yields exactly one empty chunk for it (ADR-0023's empty-unit round-trip), so the cover_refreshes list still rides a sync_apply_unit frame with shortcuts: []; the frontend processes the refreshes and acks the empty chunk, and the commit advances the unit normally. No special-casing exists on either side.

A NULL fingerprint (every pre-migration-016 row) with an existing cache file is adopted: the fresh fingerprint is persisted without a download, so the upgrade never mass re-downloads a library; a NULL with no cache file downloads on the apply path as before. The fingerprint is only ever advanced when the cache is actually confirmed against the server — a fresh download, a cache reuse/seed (the reporter's commit merges the confirmed value from the unit's pending_cover_sources staging, else preserves the existing one), or the NULL-adopt — so a failed download keeps the old fingerprint and the change is retried next sync. A cover-only change never re-applies the shortcut itself (that apply path writes launch options under the ADR-0025 invariant — pure churn for a cover).

url_cover fallback on a 404 RomM asset (#1450). When the RomM-local cover asset returns an HTTP 404 (the server's own cover resources are missing while its web UI still renders from url_cover), the cover download retries once against the ROM's url_cover — an external metadata-provider CDN (SteamGridDB / IGDB / …) — before giving up. The retry lives in ArtworkService._download_cover_atomic: only a RommNotFoundError (404) with a non-empty url_cover triggers it, so a transient transport error keeps today's retry-ladder behaviour and never falls back, and an empty/absent url_cover is exactly today's failure (warning, gray tile). The cover routes are exactly why the entity-verdict rule in what makes a 404 an entity verdict carves them out: a static-mount miss answers with the generic route-404 body, so demanding an entity answer there would disarm this fallback. The fallback fetch goes through a separate, bearer-free adapter seam (RommRomReader.download_cover_from_url → RommHttpAdapter.download_external): the host-bound RomM bearer must never reach a third-party origin, so only the plugin User-Agent is attached (the CDN behind Cloudflare Bot Fight Mode also 403s the default Python-urllib UA). Because url_cover is untrusted server-supplied input, download_external scheme-allowlists it to http/https before any fetch — a file:/data:/ftp:/scheme-relative URL is rejected with a RommApiError so a file:///etc/passwd never reaches urlopen (nothing revalidates the origin on a redirect hop, here or on the RomM download paths — #1889 — and nothing blocks a private or link-local address after DNS). The fingerprint records the source actually applied — url_cover on a fallback, threaded back through the download_artwork applied_sources accumulator (sync path) or the direct persist (refresh_changed_covers / refresh_cover) — so the compare stays truthful: because the fresh RomM path_cover string never equals the stored url_cover, a later fixed RomM asset (or a changed url_cover) is always re-checked, at the cost of re-downloading the fallback ROM's cover each sync until the RomM asset is repaired.

Conditional-request revalidation on a ts-only change (#1454). A server-side rescan re-stamps every ROM's updated_at — and thus the ?ts= in every cover path — without touching the cover files, so the cover_source fingerprint changes library-wide and the #1386 compare would re-download every cover. Instead, when a fingerprint change is ts-only (the paths are equal once the ts query param is stripped — domain/cover_refresh.cover_ts_only_change) and a stored validator exists, the cached bytes are revalidated with a conditional request rather than re-downloaded: every regular download records the response's HTTP validator (ETag / Last-Modified) in a {rom_id}.cover-meta.json sidecar beside the cache file, and a ts-only change re-requests with If-None-Match (else If-Modified-Since) through the normal authenticated RomM adapter path (RommRomReader.download_cover → RommHttpAdapter.download_conditional, bearer + UA as today). A 304 Not Modified keeps the cached bytes, adopts the fresh fingerprint (so the next sync is clean — this is the whole point), and refreshes the validator if the 304 carried one, with no grid republish and no in-session tile re-apply (the tile is already current); a 200 replaces the bytes, validator, and fingerprint. Validators are an optional capability — a ROM with no stored validator (first sync, or a server/proxy that strips them) plain-downloads and seeds one for next time, and a conditional-request transport error is exactly today's failed-download path (retry ladder, old bytes + old fingerprint kept). The sidecar lives beside the cache file (pruned with it on shortcut removal and by the orphan sweep); the #1450 url_cover/CDN fallback path is out of scope for revalidation and keeps its once-per-sync re-fetch. No RomM minimum-version change — the floor is untouched.

Unstamped-platform re-run: the restamp_platform_count signal (#1416). A heartbeat-timeout run's late-ack recovery (ADR-0023) leaves a platform complete but unstamped: the timed-out apply cleared the stamp at its start and its late ack re-binds the chunk without re-writing it (the late-ack path never passes a platform_stamp). The wholesale-skip gate then full-fetches that platform on every future sync (no stamp = no skip authority), and — because its shortcut delta is otherwise empty — the QAM's preview would short-circuit on "no changes", so the re-walk that would re-stamp it never runs and the run's interrupted status lingers indefinitely (only an apply run records a fresh SyncRun). sync_preview therefore counts the enabled platforms lacking a PlatformSyncState stamp (LocalLibraryReader.do_count_unstamped_platforms, a side-effect-free read) and rides it as the additive summary field restamp_platform_count (absent/0 tolerated by old consumers). The frontend treats a restamp-only preview (all diffs zero, count > 0) as apply-able — "No changes — finishing an interrupted sync." with the normal Apply/Cancel confirm — mirroring the cover-only flow. The gated apply then runs the unstamped platform (the skip gate never skips it), and its 0-delta empty final chunk re-writes the stamp and records a fresh completed SyncRun, healing both symptoms. The stamp write stays pipeline-owned: the preview only counts unstamped platforms, it never stamps (check_sync_lifecycle_owner / the final-chunk-only rule).

Per-unit apply is chunked, durable per chunk (ADR-0023). The delta emit list is split into fixed-size chunks (_APPLY_CHUNK_SIZE, 200) by the pure domain/sync_chunking.py::build_unit_chunks, and ChunkDispatcher (services/library/chunk_dispatcher.py) processes them one at a time: emit sync_apply_unit → wait for the ack → commit that chunk's roms rows durably → next chunk. Each sync_apply_unit carries chunk_index / chunk_count / chunk_offset / unit_total, and shortcuts is the chunk slice; the reporter's per-chunk commit is the same group-aware two-pass write (every fetched sibling upserted, only representatives bound; Rom row then rom_metadata, FK-safe) over that chunk's row subset, so a committed chunk is crash-safe on its own and a mid-unit failure forfeits only the in-flight chunk. Chunks cut only at sibling-group boundaries (overflowing 200 to keep a game's dumps whole so they never straddle two commits); no-emit groups and unmatched leftovers ride chunk 0, and an empty unit is one empty chunk (the empty round-trip still commits its unbound rows). The whole-unit staging (pending_sync / pending_all_roms / pending_cover_sources) is set once by the dispatcher, before its first chunk goes out; only the per-chunk coordination is re-armed each chunk. The motivating field crash is #797: a 3084-shortcut unit emitted in one frame lost ~24 minutes of work when steamwebhelper OOM-crashed before any ack — a 200-chunk caps both the frame payload (~200 KB vs ~3 MB) and the crash blast radius (~2 min vs 24+).

Per-chunk wait: timeout vs. cancel. The dispatcher emits each chunk's sync_apply_unit, then waits on unit_complete_event (heartbeat-clocked) for the frontend's report_unit_results ack. When the wait returns None the teardown branches on the cause (#1052) — and either way, every chunk committed before this one stays committed:

  • User cancel (is_cancelling() already True at the return) — the in-flight chunk is intentionally discarded: clear the staging, null unit_complete_event, and clear active_unit_id + active_chunk_index. A stray late ack then no-ops.
  • Heartbeat timeout (still RUNNING) — the frontend has already created this chunk's Steam shortcuts and will fire a late report_unit_results, but in production after the run has already wound down. The dispatcher moves the abandoned chunk into an abandoned_chunk stash on the box (stash_abandoned_chunk): its run/unit/chunk identity plus this chunk's ROMs (only the abandoned chunk, never the whole unit, the metadatum source), while it keeps the whole-unit staging (pending_sync / pending_all_roms / pending_cover_sources) live for the recovery commit to read and clears the dispatch identity (unit_complete_event + active_unit_id + active_chunk_index). It then marks the run interrupted and flips CANCELLING so the loop stops. The stash lives outside the run-lifecycle state and deliberately survives finish_run (which nulls current_sync_id), so a late ack arriving after teardown can still recover it (#1367 — an earlier design kept active_unit_id live and lost the ack once the run wound down). SyncReporter.report_unit_results then matches the stash by identity (take_abandoned_chunk) and drives commit_unit_results itself over the stashed rows — every fetched sibling upserted (identity + metadata), only the acked representatives bound, and never a platform_stamp (a timed-out platform is incomplete) — persisting the delivered bindings instead of leaving orphan shortcuts that the next sync re-creates as duplicates. The stash has a bounded lifetime: the next run's try_begin_run clears it, so a frontend that crashes and never acks just leaves inert data. The committed binding is mapped by the next sync's existing-shortcut scan, so no active orphan deletion is needed (a Steam shortcut is the sole record of its tile).

Session-budget gate: GC-before-measure at each chunk boundary (ADR-0024). Steam's SharedJSContext renderer OOM-crashes at ~2.45–2.53 GB RSS and never self-recovers within a session; each created shortcut costs 0.7–1.5 MB permanently (measured on-device). Chunking (ADR-0023) makes hitting that cliff a cheap resume; this gate stops the run before it, at a chunk boundary, as a controlled pause. Every renderer reading and verdict below belongs to SessionBudgetMonitor (services/library/session_budget.py), which the chunk dispatcher holds directly — the fakeable seams are the two renderer adapters below, not the monitor. At each chunk boundary the monitor reads the renderer's RSS (RendererRssFn, adapters/renderer_rss.py — the max VmRSS across steamwebhelper processes from /proc) and, when that raw reading is at/above GC_SKIP_BELOW_KB (1.5 GB), forces a renderer GC (RendererGcFn, adapters/renderer_gc.py — an HeapProfiler.collectGarbage over the CEF debugger on localhost:8080, so the reading reflects settled heap not transient garbage that Steam's measured-unreliable natural GC hasn't reclaimed) and re-reads; below the floor the raw reading (which can only over-estimate) already clears every threshold, so the ~5 s GC is skipped and a small sync pays zero GC cost. It then runs the pure domain/session_budget.py::gate_decision over the chunk's composition-priced cost (chunk_worst_cost_kb): now that the emitted chunk is new + changed only (the delta-restricted apply above), a create is priced at the worst-case create rate plus its transient cover term while a changed/rebind item is priced at the lighter Set*-walk rate — the gate no longer prices every item as a cover-applying create. Pause iff rss + chunk_cost ≥ limit; the pause log line names the composition ("… + N creates + M updates projects …"). The frontend decides create-vs-update itself via its existing-shortcut scan, so a small backend/frontend mismatch only ever overprices (worst-case safe). Both modes are predictive and differ only in the limit line: every later chunk projects against cliff − margin (≈2.2 GB, keeping the anti-thrash safety margin), while the run's very first chunk projects against CLIFF_KB (≈2.45 GB) instead. Forward progress must be guaranteed (the run has to apply at least one chunk or it loops forever on a no-progress pause), so the first chunk is allowed to spend into the safety margin — but the predictive projection still stops it before the crash line. A resume's first chunk therefore proceeds only when its worst-case peak stays below the cliff (≈1.95 GB for a full 200-item chunk of cover-applying creates, each priced create + cover) and can never be projected past it; a resume attempted without a Steam restart re-pauses cleanly rather than driving a chunk into the cliff. On a pause it sets run_paused + a distinct interrupt_reason and requests cancel — the loop returns cleanly with prior chunks committed and the terminal write records the new terminal status paused (migration 014, its own status distinct from a crash's interrupted; both resumable, but the split lets the UI say "(paused)"). That reason is the one backend sentence the panel shows the reader verbatim (as the pause toast), so it names an action and never a control: nothing on this side can know what the Sync page's start button currently says, since that turns on the frontend's Skip-preview setting and on a resume question read from the stats. Completed platforms keep their PlatformSyncState stamps, so a resume redoes only the remainder. Every step is fail-open: an unavailable RSS reading (no steamwebhelper, unreadable /proc) skips the gate — and short-circuits further GC attempts — for the rest of the run (logged once), a seam error is caught locally, and a failed GC only makes the reading less precise; measurement never blocks a sync. The same seams feed the UI surfaces: sync_preview returns pause_likely (a predict_run_crosses prognosis pricing only new creates + changed updates, never fully-unchanged items, so an unchanged re-sync never warns), a clean run's sync_complete carries restart_recommended (post_run_advisory, RSS > ~1.8 GB, read GC-first), and the get_session_budget_status callable returns a live RSS reading (no GC) plus the three fixed threshold lines (warn_kb ≈1.8 GB, ceiling_kb ≈2.2 GB, cliff_kb ≈2.45 GB) for the persistent QAM banners (a blue "paused" banner, a yellow high-heap banner) and the always-on "Steam memory" status row. The row's value text is traffic-light coloured against those three thresholds — green / yellow (warn_kb) / red (ceiling_kb) — so the frontend holds no threshold magic numbers, and while a sync runs (or a paused banner is showing) the row polls the callable (~5 s during a sync, ~10 s while paused) so the number tracks the climbing RSS and the blue paused banner notices once a Steam restart frees memory. That notice is driven by resume_ready on the callable (domain.session_budget.resume_would_proceed: rss + RESUME_HEADROOM_CHUNKS × FULL_CHUNK_WORST_KB < ceiling — room for TWO worst-case chunks, ≈1.2 GB bar, because a one-chunk bar sits exactly on the pause point where Steam's own small frees flicker the verdict; None when RSS is unreadable) — when it flips true the blue banner announces memory is free and hides the restart button. The callable also carries the paused run's progress (run_done_items / run_total_items), so that banner can read "1200 of 2001 games done": the counters are run-scoped fields on LibrarySyncStateBox, stamped with the plan's ROM total and grown by the delta-restricted apply's SKIPPED entries (already correct on their shortcut), each wholesale-skipped unit's ROMs, and every COMMITTED chunk's acked items — an emitted-but-uncommitted chunk (cancel / heartbeat timeout / the pause itself) never counts, so the number can't over-report. They live in the backend deliberately: the plugin process survives the Steam restart the banner asks for, while the frontend reloads. In-memory only — a plugin reload wipes them and both come back None, which the banner renders by dropping the sentence rather than showing a zero. The banner names no button of its own (#1789): the panel hands it the sync button as one value — the label and whether pressing it resumes — so it quotes whatever the panel actually rendered ("Check for changes", "Resume Sync", "Sync Library", "Apply Sync"), rather than deciding from the paused status, which survives a Force Full Sync that the resume does not. When nothing can be resumed it also drops the progress sentence, because "1200 of 2001 games done" promises a head start the clear has just discarded. That row also shows the last run's signed RSS growth, appended inline after the value ("X.X GB · last run ±Y"), measured at EVERY terminal (completed / paused / cancelled / interrupted) so a paused run reads as its own consumption-so-far rather than a prior clean run's: a RAW read taken unconditionally at run start is the baseline (run_start_rss_kb — captured before any chunk, so even a fully-incremental-skip run still records one and reports ≈ +0.0 GB), the terminal RSS read is the end, and session_memory_delta differences them (an approximation for information only, which a raw start baseline is fine for). The value is retained in last_run_delta_kb so get_session_budget_status surfaces it on a QAM remount (in-memory only, lost on reload, no migration; None when either endpoint was unmeasurable, so a stale delta is never shown); the UI reads it from that callable, so it is deliberately NOT put on the sync_complete wire. Both banners also offer a Restart Steam now button that calls SteamClient.User.StartRestart directly from the frontend — a deterministic full client restart that resets the renderer's per-session budget to the ~430 MB baseline. The button is disabled while a game is running and hard-guarded on click (isAnyAppRunning) so a restart can never close a game. The RSS reader and GC trigger are wired through SessionBudgetMonitorConfig; the gate's per-item cost is a parameter, and because the apply now pushes each created shortcut's cover through Steam's artwork API (SetCustomArtworkForApp, transiently resident but GC-reclaimable — hence the GC-before-measure), the monitor prices each create at the worst-case create rate plus the transient cover term (COVER_TRANSIENT_KB) at both the chunk gate and the preview prognosis, while a changed item stays at the lighter update rate.

Run/unit/chunk identity on the ack (#1041). Every sync_apply_unit event carries the run_id (the run's current_sync_id UUID), the unit_id (the WorkUnit.id), and the chunk_index; the frontend echoes all three back on the report_unit_results ack. The orchestrator stamps the dispatched unit's id into active_unit_id and the current chunk into active_chunk_index just before it emits, and report_unit_results validates the ack against them: the run_id must match current_sync_id, the unit_id must match active_unit_id (both compared by string value, since a platform's id is numeric and a collection's is a string), and the chunk_index must match active_chunk_index (as an int). An ack that fails the active-unit check falls through to the abandoned-chunk stash (take_abandoned_chunk): a heartbeat-timed-out chunk's late ack — which in production arrives after the run wound down, so active_unit_id / active_chunk_index are already cleared and current_sync_id is null — matches the stash by identity and drives the recovery commit (the per-chunk timeout branch above, #1367). Only an ack matching neither the active chunk nor a stash — a late ack from a cancelled run arriving while a fresh run is in flight, a stray ack for a different unit, or a superseded chunk — is ignored (logged at debug, returns {success: True, count: 0, ignored: True}): it is neither recorded, signalled, nor committed, so it can never be credited to the wrong run/unit/chunk. active_unit_id + active_chunk_index are cleared once the chunk commits, is cancelled, or times out (its identity moves into the stash). On the frontend side, the unit handler does not send the ack at all once cancel has been requested — the first line of defence against a cancelled run's bindings landing in whatever run started next.

Run identity on the cancel (#1198). cancel_sync(run_id) is run-scoped, mirroring the ack-path _ack_matches_active_unit check above. A Cancel click meant for run N can land after run N finalized to IDLE (current_sync_id nulled) and run N+1 started fresh — an unscoped cancel would flip run N+1 to CANCELLING (a sync the user never cancelled), which would abort it and report cancelled. So when the requested run_id is truthy and does not match current_sync_id, the cancel is ignored (logged at INFO, returns {success: True, message: "Cancel ignored (stale run)"}) — sync_state is left RUNNING. A matching id (or no active run) flips RUNNING → CANCELLING as before. A falsy/None run_id cancels unconditionally — legacy callers and the "no active run id captured yet" safety case — so cancel is never made less reliable. Every cancel logs one INFO line recording the requested run_id, the active current_sync_id, and the sync_state at call time. The frontend sources the run id from the backend-fed sync_progress store — its additive runId field (see below) — and passes it on the Cancel click; in the pre-run "Fetching library…" window the store's runId is still "", which maps to the unconditional cancel path rather than a stale prior-run id the backend would reject as a cross-run mismatch. sync_cancel_preview is not run-scoped: it only discards the pending preview delta (pending_delta), never touches sync_state, and sync_apply_delta independently validates the preview_id, so a stale preview-cancel cannot abort a fresh sync.

Group-aware sync — one Steam shortcut per sibling group (ADR-0021). A game frequently exists in a RomM library as several dumps of the same title (region / language / revision variants). These share a sibling group (same roms.sibling_group_key, derived client-side by domain/sibling_group.py: a connected component over RomM's sibling_roms edges, keyed by the highest-priority metadata source the component agrees on — compute_component_group_keys, stamped onto the raw dicts before the shortcut build at both the preview and per-unit call sites, ADR-0022). The sync pipeline treats a group as one game = at most one new Steam shortcut; the sibling row holding shortcut_app_id is the group's active version.

  • Persist all siblings, emit one shortcut per group. The reporter's per-unit commit upserts a roms row (identity + version metadata, via the sync UPSERT) for every fetched ROM of the unit, but only the group's representative carries a Steam-shortcut binding — a non-representative sibling is a tracked, unbound row. Persisting the whole group is what the version picker (#1297) and the incremental-skip gate need; binding only the representative is what stops the same-named-collision and duplicate-shortcut bugs the ADR describes.
  • The collapse happens at the same build_shortcuts_data boundary in both paths (domain/sync_diff.py::collapse_sibling_groups). Preview and apply build the full per-ROM shortcut list, then collapse it to one entry per group against the bound-row registry (preview reads LocalLibraryReader.do_read_preview_baseline, apply reads LocalLibraryReader.do_read_apply_registry per unit). The collapse takes an explicit complete_group_view flag: a sibling group is per-platform, so the whole-library preview union and each platform apply unit see a group's whole membership (complete_group_view=True) and collapse identically — the preview counts can never diverge from what those units create (the #1292 counts-vs-reality bug class). A collection apply unit spans platforms and may fetch only one unbound sibling of a group whose bound member it never fetched, so it is a partial view (complete_group_view=False): a group that already holds a binding anywhere is grandfathered untouched — never rebound onto the partial member, never given a second shortcut — because the group's real representative rides its own platform unit in the same run. (Inferring "bound sibling absent ⇒ vanished" from a partial view would rebind a live installed game onto an uninstalled sibling — #1296.) classify_roms runs over the collapsed set, so its new / changed / unchanged / stale buckets count games, not dumps, and an unbound sibling stops reading as a perpetual "new".
  • Resolution chain (domain/sibling_resolution.py::resolve_group_representative, total + shuffle-stable): an installed sibling wins; else an existing binding; else RomM's per-user default (is_main_sibling); else the 1G1R ranking — prerelease demotion > region priority > revision (newest) > alphabetical fs_name_no_ext > rom_id. The first three legs are membership filters; the rest are the total order applied inside the surviving leg. Prerelease demotion ranks first: a member whose structured tags name a draft build (Alpha / Beta / Proto / Sample / Demo, case-insensitive, tolerant of a trailing number; Unl / Aftermarket / unknown tags are neutral) is demoted below every retail sibling across regions, so a finished (Japan) release beats a (USA) (Beta). Revision ranks just after region: within one region the newest revision wins (natural compare; base dump = lowest), but never lifts a lower-ranked region (a (USA) base still beats a (Europe) (Rev 9)). The alphabetical leg also keeps a base dump ahead of a filename-only re-dump ((Virtual Console), (Extended Edition)) RomM does not parse into a tag. Region priority ranks a version by its best regions entry against a fixed build-time order — World > USA > Europe > Japan, every other named region after these (alphabetically), a no-region version last (a fixed constant, not language/system detection) — which the user may re-head with a single preferred_region setting ("auto" = the fixed order; any other value lifts that region to the top). The leg is evaluated at resolution time, so the setting takes effect on the next sync and an existing binding shields its group (the installed/binding legs win before region priority is consulted). Used wherever one version must be chosen (a new group's shortcut, a rebind's target). The QAM dropdown that sets preferred_region is populated from fixed anchors (Default, World, USA, Europe, Japan) plus the distinct regions read from the local library (get_known_regions over roms.regions — no server call), and a confirmation modal states the apply-at-next-sync / no-rename semantics before persisting.
  • Canonical name at mint (domain/sibling_resolution.py::canonical_group_name): a NEW group's shortcut is named after the member ranked first by the pure order (prerelease demotion > region priority > revision > alphabetical > rom_id), ignoring the installed/binding/default filters — so a Japanese default still binds Japan but the shortcut carries the USA name (never majority voting: two Japan dumps + one USA yields the USA name). This is mint-time only: the name is sticky forever after (ADR-0021 §2), so a rebind/grandfathered entry carries the persisted bound name verbatim and a live shortcut is never renamed. The appId is minted from the canonical name by the frontend AddShortcut; the DB binding lands on the representative rom_id; the reporter is agnostic to the name↔rom mismatch (it binds on rom_id + the acked appId, and persists each sibling's own RomM name from pending_all_roms). preferred_region is read from settings by sync_orchestrator and threaded into every collapse call site (preview + apply) as the same value within a run.
  • Rebind, not remove. When every bound sibling of a group vanishes from the server but the group still has fetched members — and the collapse sees the group's whole membership (a platform unit or the preview, complete_group_view=True; a partial collection view never rebinds) — it emits one rebind entry keyed to the vanished sibling's rom_id (so the frontend reuses its existing shortcut through the rom_id → appId map — no delete + recreate churn), carrying the representative's launch bake and a bind_rom_id marker. The commit translates the ack through bind_rom_id, so the DB binding moves onto the representative; the collision-safe Rom save then unbinds the vanished row. The shortcut's appId, artwork, collections and playtime all survive — only the active version changes. A whole-group disappearance still routes through the normal stale path, and the #1036 committed-appId guard is unchanged (the reused appId is committed this run, so it is never emitted for removal).
  • Grandfathering. A group that already carries several bound shortcuts (a library synced before this model) keeps them all — each surviving bound sibling stays its own tracked shortcut; no new shortcut is minted for a group with a binding. Convergence to one-shortcut-per-group happens naturally as the user uninstalls duplicates.
  • Downstream group semantics. The incremental-skip gate compares RomM's platform rom_count against the persisted rows carrying the completion stamp's fetch generation — bound and unbound alike, so skip parity on platforms with sibling groups holds (a NULL sibling_group_key on such a row still forces a backfill full-fetch), while a row for a rom_id the server has since dropped is excluded rather than inflating the count forever (#1504; the row itself is retained per ADR-0007 — nothing is deleted). The backfill gate reads the same generation: only a row the last complete fetch returned may demand a full fetch, because a dropped row is never returned again and no fetch could ever fill its key in — counting it would wedge the platform into a full fetch on every sync, forever. A stamp with no generation predates the contract and falls back to counting — and backfilling for — every row. Artwork downloads only for the emitted representatives (+ grandfathered bound siblings), never eager sibling covers. Steam-collection membership resolves each RomM collection rom_id with a group fallback — an unbound sibling maps to its group's bound sibling's appId — so collecting or favouriting any version puts the game's single shortcut in the collection.

Same-named collections union into one Steam collection (#1503). RomM enforces name uniqueness only per-table and per-(name, user_id), so two enabled collections can share a display name — a standard collection and a smart/virtual collection, or (multi-user) another account's public collection the list endpoints return. Steam's collection namespace is by-name (RomM: [<name>] (host)), so both must resolve to the one Steam collection. The finalize accumulator (_state.py pending_collection_memberships) is therefore keyed by a collision-free (collection_kind, collection_id) identity — the same identity the #742 completion stamp uses — with the display name carried in the CollectionMembership value; the real-sync and preview write paths share one key builder so they cannot drift. The reporter (_resolve_collection_memberships) then groups the accumulator's entries by name and unions their resolved appId sets (order-preserving, de-duplicated across collections; each collection's own resolution already dedups within), emitting the unchanged by-name romm_collection_app_ids: {name → [appId]} contract — a single-collection name unions a set of one and is byte-for-byte the pre-#1503 output. The union runs after the owner scope (below): once a foreign collection is dropped from the work queue it never reaches the accumulator, so the own scope narrows what is unioned rather than changing how the union works.

Collection kinds — one virtual kind covers RomM's browsable virtual types (#1538). get_collections and build_work_queue group collections into three internal kinds: standard, smart, and virtual. RomM's own UI calls the ownership-carrying first kind Standard (the plugin's earlier name for it was user, a misnomer renamed internally by #1539 — display "My" → "Standard"; see the enabled-bucket migration note below). The virtual kind is RomM's ownerless VirtualCollection (base64 id, no user_id, no stable updated_at — never stamped), and it carries a virtual_type sub-field on each collection dict/setting so the Collections tab can list it under Franchises or IGDB collections. RomM's VirtualCollection has five type values, but the plugin syncs only the two RomM itself surfaces as browsable collections: IGDB franchise and the default IGDB collection (series). genre, company, and mode are intentionally excluded — RomM treats genre/company as ROM filter facets (not collections) and mode as neither, so they never appear in RomM's Collections view. The supported set is a single constant (services/library/fetcher._SUPPORTED_VIRTUAL_TYPES); the fetcher fetches each supported type (per-type fail-open) and merges them under the one virtual bucket. Because the type is baked into the base64 id, ids are globally unique across types, so one enabled-bucket keyed by id cannot collide. The owner scope, stamp-exclusion, and per-unit ROM-fetch dispatch all stay a single kind == "virtual" branch, not fanned out per type. On disk the enabled-collections bucket was renamed franchise → virtual by the lossless settings.json migration v10 → v11 (domain/state_migrations._migrate_v10_to_v11): it renames the bucket key while preserving every enabled id, so a previously-enabled franchise collection stays enabled and no re-login is required. (The historical v2 → v3 split still produces the franchise bucket; the v10 → v11 step renames it afterwards, so that frozen step is untouched.) The ownership-carrying bucket was likewise renamed user → standard by the lossless settings.json migration v12 → v13 (domain/state_migrations._migrate_v12_to_v13, same merge-into-existing semantics), and old collection_sync_state completion stamps keyed collection_kind = 'user' are rewritten to 'standard' by SQLite migration 022 so an unchanged standard collection still takes the incremental skip across the upgrade (#1539).

Batch collection enable — save_collections_sync (#1539). A single settings write that stamps every id it is given into one kind bucket and touches nothing else in it; the Collections tab's Enable all / Disable all send it the ids their table lists, so what is written is exactly what the page shows. An unknown kind or a non-list id argument is rejected with the canonical failure shape; an empty id list is a success no-op.

Collection owner scope — own is a sync scope, not a display filter (#1532). RomM's collection list endpoints return the signed-in user's own collections plus every other user's public collection. The QAM's Other users' collections toggle writes collection_owner_scope ("own" / "all", default "all"; how the QAM presents it is docs/architecture/qam-panel.md § Library). Ownership is a pure predicate (domain/collection_owner.is_own_collection): a collection is own when it is a virtual collection (RomM's VirtualCollection model carries no user_id column — these are global/derived and belong to no one, so they always survive), when the plugin's own identity is unknown (the unknown-identity fallback, which drops nothing), or when the collection's user_id equals the stored romm_user_id; only standard and smart collections carry a user_id to compare. get_collections tags each row with is_own from the same rule with one difference (domain/collection_owner.listing_is_own): while the identity is unknown a standard or smart row carries is_own: null rather than true, because nothing established that it is the user's; how the page reads it is docs/architecture/qam-panel.md § Library. Virtual rows are always true. build_work_queue applies the same predicate: under "own" with a known identity it drops foreign standard/smart units from the queue — so a foreign collection enabled earlier is never synced — while virtual units and every unit under "all" pass through unchanged. The scope applies over the per-kind enable state without mutating it, so switching back to "all" restores the prior enables. Because an unknown identity drops nothing, the feature is non-breaking: it silently no-ops until romm_user_id is stamped (see the ConnectionService lazy-identity note), then takes effect — no re-login required.

Each collection row states how many of its members are already in Steam, and who owns it (#1833). get_collections adds two fields to its rows. in_steam_count, on all three kinds, counts the collection's member ROM ids (RomM's listings carry rom_ids on each) that the sync's collection filing resolves to a shortcut (CONTEXT.md → Reachable, which says how it differs from the platform count). It counts members, not shortcuts: two versions of one game both count although they share one. It is absent, never 0, when it is unknown — the local read failed, which does not fail the listing and is logged as a warning. The fetcher asks its inward pair (LocalLibraryReader.do_read_reachable_rom_ids) in one short read UoW once all the listing requests have returned, never across one; the per-row count is domain/collection_listing.py's. owner_username is RomM's own field on the standard and smart listings (CollectionSchema / SmartCollectionSchema in every RomM version the plugin accepts), forwarded as null where a listing lacks it; virtual rows carry no such key, since they have no owner.

Incremental skip — the per-platform completion stamp is the sole authority. A platform unit skips only when its PlatformSyncState stamp exists (ADR-0023); the stamp's completed_at is the "unchanged since" reference for the updated_after server-delta probe. A completed SyncRun's last_sync is deliberately not a fallback: a run-scoped timestamp cannot see a platform whose shortcuts were locally removed and only partially re-applied afterwards, so trusting it can silently skip a platform with missing shortcuts. No stamp means a full fetch — including, once, every platform's first sync after this contract shipped. The reporter writes the stamp when a platform work unit's last apply chunk commits — atomically in the same write UoW as the chunk's roms upserts — so a platform that fully synced inside a run the user later cancelled/crashed still skips on the next run instead of re-walking every already-applied game through CEF. All the existing guards still gate the skip (zero-bound-rows, the sibling_group_key backfill for rows carrying the stamp's generation, the updated_after server-delta check, the persisted-row count match); additionally the stamp's rom_count must still equal the server's platform rom_count — a server-side count change invalidates it.

The stamp's contract is stamp exists ⟺ the platform's most recent apply attempt ran to completion, so a stale stamp can never skip a half-mirrored platform. Because unbinding keeps the roms row (ADR-0007), a platform's persisted-row count survives a partial re-apply or a local removal, so a surviving stamp with a matching rom_count would otherwise let the skip drop the un-recreated games (the #1025 silent gap). Two rules keep the contract true: the orchestrator clears the stamp at a platform unit's apply start (once the fetch succeeded and the apply is about to emit its first chunk) and only the final chunk re-writes it, so an apply interrupted by a crash / cancel / heartbeat-timeout before the final chunk leaves none; and the local destructive flows — "Remove all shortcuts" and per-platform removals (via report_removal_results) plus the Steam-UI-deletion reconcile (reconcile_live_shortcuts) — delete the touched platforms' stamps in the same write UoW as the unbind. The reporter's server-side stale removal is the deliberate exception (it leaves the stamp, since a server-dropped ROM lowers RomM's rom_count and the count guard catches it). "Force Full Sync" (clear_sync_cache) clears every stamp (and resets the recorded applied_launch_options to NULL), which is the entire full-re-fetch + full-re-apply arm — the stamps are the fetcher's sole skip authority. The sync_runs history is deliberately preserved (#1318): it feeds no skip gate and is the source of the "Last sync" display, so deleting it forced nothing and only blanked the panel to "Never" right after a reset.

That preservation is also why the panel's resume offer is derived from the surviving skip authority, not from the run history (#1789). The history says only that a run ended without completing, and after a Force Full Sync it says that while everything it implied is gone — so a history-derived offer promised to continue progress that had just been discarded. The condition it replaced measured the wrong thing in the same way: it paired the incomplete attempt with "bound shortcuts exist", but Force Full Sync does not delete shortcuts, it deletes the stamps and the recorded launch options.

A resume rests on skip authority, and this plugin keeps two kinds, cleared together by that one reset: a completion stamp (whole platform or collection skipped at fetch time, ADR-0023) or a recorded applied_launch_options (one game skipped at apply time, ADR-0025). Neither subsumes the other, so the offer reads both. A run cancelled inside its first platform unit reached no final chunk and holds no stamp, but its committed chunks wrote shortcuts and recorded their launch commands — the next run genuinely does less work, so that is a resume. A row predating migration 015 carries a NULL recorded value while its platform's stamp survives, so an upgraded install can hold stamps and zero recorded games and still skip those platforms wholesale.

get_sync_stats therefore carries two additive fields alongside the display ones: resumable_games (bound rows with a non-NULL applied_launch_options) and has_completion_stamp (PlatformSyncStateRepository.has_any() or CollectionSyncStateRepository.has_any()). Both ride in the read UoW that already scans roms — the game count is one more condition inside that loop, since iter_all already selects the column, and each stamp probe is a SELECT 1 … LIMIT 1 over a leaf table — so the panel mount pays nothing measurable.

The game count is bound-AND-recorded, never recorded alone. classify_roms (domain/sync_diff.py) sends an unbound row down the NEW branch — if not reg or not reg.get("app_id") — before it reads the recorded value, because the next run has to mint the shortcut regardless. Requiring the binding is what makes the count fall to zero after "Remove all shortcuts", where unbinding deliberately keeps the row and its recorded command (ADR-0007) — a count over every row would keep offering to resume shortcuts that no longer exist. The panel keeps roms > 0 as a conjunct on top, and it is load-bearing rather than a restatement: has_completion_stamp asks whether any stamp survives anywhere, while the removal path is surgical — it deletes only the platform slugs its removed rows name, and only the collection stamps whose member set intersects those rows. A stamp naming nothing the roms table still holds outlives a remove-all. Prune makes that reachable: it deletes roms rows and never touches platform_sync_state (PruneRegistry.delete_rows invalidates intersecting collection stamps only), so a platform whose games RomM dropped keeps its stamp with no rows left to name it, and the next remove-all cannot see that slug to invalidate it. Verified end to end: with such a stamp standing, a remove-all leaves roms == 0 and has_completion_stamp == true, so without the conjunct the panel would offer a resume over zero shortcuts. It carries the boundary rule too — a run stopped before a single shortcut was written begins from scratch. resumable_games is a count of what is done, never of what is left — naming the remainder needs the server's library — so the QAM line under the button states what a resume would skip, and is omitted when the count is zero.

Single-owner run lifecycle (#1202). The run-lifecycle pair — sync_state (idle/running/cancelling) and current_sync_id — is mutated only through four verb methods on LibrarySyncStateBox, never by direct field assignment from a sub-service. Confining those two writes to the box makes run admission, cancellation, and termination a single compare-and-swap on the plugin's one event loop, so a rapid Sync/Cancel can't interleave a stale terminal with a fresh run's start and leave a half-reset id:

  • try_begin_run(run_id) — the admission guard. Compare-and-swap: returns False (no state change) if a run is already in flight, else flips IDLE → RUNNING and stamps current_sync_id. Every entry point (start_sync, sync_preview, sync_apply_delta) goes through it, so an overlapping second Sync/Apply is rejected with {success: False, reason: "sync_in_progress"} instead of beginning a concurrent run that would double the terminal events. A rejected apply does not consume pending_delta, so the still-valid preview survives for the legitimate apply.
  • request_cancel(run_id) — the run-scoped cancel routing (the #1198 / #1200 logic, centralized): "no_sync" when idle, "stale" when a truthy run_id doesn't match the active run (cancel ignored, run left RUNNING), else flips to CANCELLING. A falsy run_id cancels unconditionally. cancel_sync and shutdown delegate here.
  • finish_run(run_id) — the terminal. Compare-and-reset: resets to IDLE + nulls current_sync_id only if run_id still owns the slot; a late, foreign, or doubled terminal is a no-op and can never null a fresher run. Every exit path (success, cancel, error, zero-unit) of the apply pipeline and the preview funnels into one finally: box.finish_run(run_id) per run — the scattered per-unit / reporter resets are gone.
  • is_in_flight() / is_cancelling() — read-only state probes for the per-unit loop and checkpoints.

sync_preview adds a final cancel checkpoint after the unit loop, immediately before it stages pending_delta, so a cancel landing in that last window routes into the cancelled branch ({success: False, reason: "cancelled"}, pending_delta left None) rather than staging a delta the user already cancelled. Nothing awaits between that checkpoint and the staging — the session-budget prognosis is taken before it, and the done frame emitted after it — so the window the checkpoint closes cannot reopen. The confinement is enforced by scripts/check_sync_lifecycle_owner.py (an AST gate that fails CI if sync_state / current_sync_id is assigned anywhere outside _state.py).

The preview outlives the panel that asked for it. PreviewDelta carries the exact answer dict sync_preview returned, and get_pending_preview hands it back verbatim — the QAM card dies with its render, while the backend holds the snapshot for its full 30-minute TTL, so a user who steps into a submenu used to lose an Apply that was still perfectly valid. Storing the answer rather than re-assembling one is what guarantees the restored card is what the user was shown. The answer also carries expires_at — an absolute epoch — which the card counts down against; absolute rather than a remaining-seconds count because plugin and panel share a clock and a deadline survives a suspend. get_pending_preview answers {"success": True, "preview": None} for all three of nothing staged, a snapshot past its TTL (dropped on the spot), and a run in flight — "no preview" is a normal answer, not a failure.

That third case is a backend guarantee rather than a frontend courtesy, because the panel cannot get it right on its own. It derives "a run is live" from the sync_progress store, and abortOptimisticSync retracts the optimistic running: true on every failed apply — including the sync_in_progress rejection, which is exactly the branch where the backend deliberately keeps the staged delta (#1202). The store therefore reads idle during a live run, and a panel that let the restored card decide the body would render it over the run's own progress rows, with the next Apply retracting the live run's UI a second time. A plugin reload mid-run reaches the same state through an empty store racing the get_sync_status seed. So the run-in-flight case is refused where the truth is: LibrarySyncStateBox.read_restorable_preview. The panel's own defence is a rendering rule rather than a second check on the read: a run in flight owns the body, so a preview held while one is going is kept and not shown, and its card appears the moment the run ends.

Frontend side: the pending preview is a module store, not panel state. frontend/src/utils/pendingPreviewStore.ts owns it. The reason is the same lifetime mismatch from the other end — sync_preview is an awaited callable, so its answer is delivered to the closure of whichever MainPage pressed Sync, and leaving the main page unmounts that instance while the run carries on. The card then never appeared for the instance actually on screen: the panel showed the idle Sync buttons over a preview the backend was holding, and only the next mount's get_pending_preview recovered it. The store is not a second source of truth — it holds the backend's answer verbatim, is only ever filled from sync_preview / get_pending_preview, and never derives or repairs one. Its ordering rule (a write takes a ticket when its information was issued; an answer applies only if no later-issued write already has; a read only ever fills, because preview: None conflates three different situations) is stated in full in that module's docstring. Two triggers fill it — the panel's mount, and a preview run reaching its terminal stage with the store still empty — and, because both write to the same destination, the race between them is "who writes first", not "which copy wins".

Withheld, not discarded. A run in flight suppresses the snapshot for the duration and nothing else — the same payload is handed back on the next mount after the run ends, as long as it is still inside its TTL. And the withholding belongs to that reader alone: sync_apply_delta's age check goes through the run-blind read_fresh_preview, so an overlapping apply is still refused by the run-slot claim with sync_in_progress instead of being rewritten into a staleness verdict.

pending_delta is box-owned, like the run-lifecycle pair. The staged snapshot's whole lifetime runs through the LibrarySyncStateBox verbs — stage_preview / read_fresh_preview / read_restorable_preview / matches_preview / discard_preview — and no module outside _state.py assigns the field (the façade exposes a getter only). The TTL lives with them: PREVIEW_MAX_AGE_SECONDS and the preview_deadline the answer's expires_at comes from are the box's, so the number the card counts down to and the verdict read_fresh_preview reaches cannot drift apart. "Too old" has exactly one expression, PreviewDelta.is_expired, reached only from read_fresh_preview — which both the pending read and sync_apply_delta's age refusal go through, so the read can never offer an Apply the apply would reject. Unlike the run-lifecycle pair this has no CI gate; it is prose plus the box's own tests.

In sync_apply_delta the ordering is load-bearing and deliberately left as three visible steps — identity check, run-slot claim, then discard — rather than one atomic "take" verb: a rapid second apply, or one landing while a run is in flight, is refused by try_begin_run before discard_preview runs, so the still-valid preview survives for the legitimate apply (#1202).

sync_progress carries runId. Every sync_progress event (and the persisted get_sync_status snapshot) now includes an additive runId: str field — str(current_sync_id or "") — so the frontend reads the active run id from the authoritative backend store rather than minting or threading its own. The idle default snapshot carries runId: "".

The snapshot describes an in-flight run only. finish_run puts sync_progress back to the idle default as part of ending a run — after the terminal frame has gone out, since every path emits or schedules it before reaching the finally: finish_run(run_id) that lands there. (The per-unit error path schedules its emit through create_task, so that one may run after the reset; it is safe because the coroutine binds the ERROR frame by value and the reset rebinds the attribute rather than mutating the dict a pending task is holding.) It exists so a remounting QAM can pick up a live run; a finished run's ending already reached the panel as an event, and a panel that merely finds a terminal frame on a later mount is required to ignore it (#1019). Left in place, the frame answered every later mount with a run that had ended — message and run id included — and the panel took it for the state of the run it was actually watching: pressing Sync, then returning during the window before the new run's first frame, showed a green "Preview ready" over a preview that was still going, and dropped the panel back to the idle buttons in the meantime. The reset is not itself a frame: nothing emits it, so no running: false reaches the panel from it.

get_sync_status answers with the run lifecycle too. The payload carries an additive inFlight: bool — box.is_in_flight(), the lifecycle state itself, not a re-reading of the frame. It rides this answer only, never an emitted sync_progress event, and the two legitimately disagree: during a cancel drain the CANCELLED frame already reads running: False while the run still owns the slot.

The panel needs it because a frame alone cannot separate two questions it must answer differently. When the module store says a run is live and the answer says nothing is running, retracting is right if the run really ended (its terminal frame was lost) and wrong if the backend simply has not heard of it yet — which is exactly the start window, where the panel has written an optimistic frame and sync_preview has not yet claimed the run slot. So the seed retracts only where the answer is evidence about that run: either it names it, or inFlight is explicitly false and the store's run carries a backend-stamped id. An idle answer against the panel's own unstamped optimistic frame is refused, and that cannot wedge, because the handler that wrote that frame always retracts it itself — from a dead instance too.

Inferring the lifecycle from the frame instead would wedge the panel: with the snapshot reset above, a finished run and a run that never started are the same answer, so a lost terminal frame would leave every later mount believing a run is live. Nothing recovers from that state — the start controls all live in the idle branch, "Cancel Sync" for an unknown run is answered {"success": True, "message": "No sync in progress"} with no terminal to follow, and Data Management and Library › Platforms gate removals on the same flag.

The same sync_plan capture point also clears the frontend's per-run cancel flag (_cancelRequested). The per-unit handler resets that flag at its own start, but an incrementally-skipped unit never runs that handler, so a skip-only run could otherwise carry a stale cancel from a prior cancelled run; resetting once per run on sync_plan is the reliable reset.

Sync time estimate and live ETA (frontend)

The QAM's time readout is a two-stage design layered on the sync_progress stream. It is pure frontend logic (frontend/src/utils/syncEstimate.ts and frontend/src/utils/syncEta.ts, both unit-tested); the backend supplies the plan (per-unit weights + planned totals, via sync_plan) and the applying frames.

  • The plan is skip-aware (#1382). At plan time the fetcher stamps every platform WorkUnit with two estimate-only riders read in one short UoW (_read_plan_estimates + the pure domain/skip_prediction.py): predicted_skip — whether the wholesale incremental-skip gate is expected to skip the platform, replaying the gate's local conditions only (completion stamp present, the stamped count and the count of rows carrying the stamp's fetch generation both match the server's rom_count, bound rows exist, no sibling-group-key backfill pending for a row of that generation; the gate's list_roms_updated_after server check is deliberately not replayed — no network at plan time) — and collapsed_count, the persisted post-collapse shortcut count, mirroring the collapse's lane selection (ADR-0021): max(1, bound rows) per sibling group — so a grandfathered legacy group with multiple independently-bound duplicates (§5) prices one shortcut per bound sibling, not one per group — plus one per keyless row. Both of those two ride the sync_plan payload conditionally-present (absent on collections, never-synced platforms, and failed reads); collapsed_count is additionally gated on the platform's completion stamp (#1412) — a never-synced platform holds only PARTIAL collection-sibling rows (ADR-0021), so an ungated count would weight the ETA below the true work, and without a stamp the frontend falls back to the raw rom_count. The payload's total_roms stays the raw pre-collapse total (backward compat); an additive total_estimated_items sums 0 for predicted skips, else collapsed_count ?? rom_count. A third rider, bound_count (#1511), counts the unit's known ROMs that already carry a shortcut_app_id, and is the one rider that rides both unit kinds. On a platform it counts the persisted rows, read in the same short UoW the skip prediction already needed that count for, and is not stamp-gated: a bound row genuinely has a Steam shortcut whether or not the mirror is complete, and zero persisted rows honestly means "every planned item is a create". On a collection it counts the bound members of the completion stamp's stored member_rom_ids (the same member set the skip replays), in one short read UoW covering every collection unit — no ROM fetch. The two sides are deliberately asymmetric on the empty case: a platform reports 0, an unstamped or virtual collection is omitted. A collection's membership exists only in its stamp, and virtual collections are never stampable (CollectionSyncState.stamp accepts only standard/smart), so 0 there would claim knowledge that does not exist. Absent and 0 price identically today; the distinction keeps the field honest for later consumers, so do not collapse it into consistency. A collection's stored member set may be stale if membership changed since the stamp — accepted and bounded, since this is estimate-only and a freshness probe would mean network I/O at plan time. A fourth rider, new_shortcut_count (#1517), is the create-side complement of collapsed_count: the shortcuts the next apply must mint rather than update — sibling groups with no binding anywhere, unbound keyless rows, and every server ROM the local mirror holds no row for (rom_count − persisted rows, clamped at zero). It is platform-only — a collection's rows belong to their platform's unit, so counting creates on a collection too would price the same shortcuts twice — and like bound_count it is not stamp-gated: an unbound group genuinely has no shortcut and a ROM with no local row genuinely has to be created, whether or not the mirror is complete. The frontend takes it as its create term directly instead of deriving creates by subtracting bound_count from the unit's weight (see the composition-priced seed below). Hard constraint (ADR-0023): the prediction never feeds the actual skip decision — _try_unit_incremental_skip at fetch time remains the sole skip authority, so a mis-prediction can only make the estimate read long or short, never mis-apply. A Force Full Sync needs no special case: clear_sync_cache deletes every stamp before the run, so a forced plan predicts no skips and drops every collapsed_count — the unit is then weighed at the full pre-collapse rom_count, but bound_count and new_shortcut_count both survive the clear (they read the rows, not the stamp) and keep the forced re-apply priced by composition (#1517). get_platforms carried the same count as a per-platform garnish for the old toggle label; the Library page's list shows RomM's own rom_count, so the garnish had no reader and the whole-table scan it cost — inside a BEGIN IMMEDIATE, on the read that gates the page's first paint — went with it (#1815).
  • Static walk-cost ceiling (pre-run seed). Before a run — in the preview, and again as the initial "up to X min" the instant a skip-preview run starts — the estimate is a pure cost model over three independent terms (#1511), because a run is three independent phases and one blended per-item rate cannot describe a mix of them: created × NEW_ITEM_SEC + changed × UPDATED_ITEM_SEC + (created + cover_refresh_count) × COVER_DOWNLOAD_SEC + a flat fixed-overhead allowance. Each constant is calibrated to its own measured mean with a ~15% ceiling margin (0.36 s per created shortcut against a measured 0.314; 0.13 s per update against 0.109; 0.15 s per cover download against 0.132 cold), and the allowance (45 s) covers the run's genuinely fixed cost — the one-time shortcut scan, the multi-page ROM/save fetches, the inter-chunk gaps, finalize (measured 17–24 s). So the seed still reads long, never short, but it now reads long by a margin rather than by a factor. Two consequences worth stating: an update carries no artwork cost (the apply loop gates cover application on created), and a cached cover is not a term — warm covers cost 0.0018 s, 73x cheaper than cold, and pricing them would need a backend cache probe in the preview path that is deliberately not added, so every create is priced as needing a cold download.
  • Composition-priced plan seed. The skip-preview seed prices the plan per unit by composition, not as one blended item count (#1511): a predicted-skip unit costs nothing, the unit's already-bound ROMs (bound_count) take the update rate, and the shortcuts that genuinely have to be minted (new_shortcut_count) take the create rate. Pricing every planned item as a create over-read by ~4x on the common case — any re-sync, and every Force Full Sync, which clears the completion stamps but unbinds nothing and is therefore an all-updates run.
  • Why the create term is read, not subtracted (#1517). The two rates are priced from independent counts; the create term is new_shortcut_count directly, never items − bound_count. The subtraction over-read a Force Full Sync of a sibling-heavy platform by ~2.5x: a forced run drops collapsed_count, so the unit is weighed at its pre-collapse rom_count, and on a platform carrying sibling groups (ADR-0021) that raw count exceeds the real shortcut count — one row per group is bound, its duplicates are not. Subtracting the bound rows priced every unbound duplicate as a phantom new shortcut, at the dear create rate plus a cover download it would never perform, even though a sibling duplicate never produces a shortcut. new_shortcut_count reports no creates there instead, because the clear takes the stamps and not the bindings, so no group is left without one. The symmetric hazard is a never- synced platform holding only PARTIAL collection-sibling rows (ADR-0021): counting creates from the known rows alone would price a handful of items for a whole platform — a short read, the one direction this estimate may not err in. The rom_count − persisted rows term prices those unmirrored ROMs as creates, so the seed stays at (or above) the full platform. A unit whose new_shortcut_count is absent (collections, older backends) falls back to the items − bound_count subtraction, the pre-#1517 behaviour.
  • Collections still price by bound members. Platform and collection units are priced the same way on an ordinary re-sync, so a collection-heavy library does not reinstate the over-read through its collections. A Force Full Sync is the exception for collections: clear_sync_cache clears collection_sync_state wholesale, and a collection's member set lives only in that stamp, so every collection unit is unstamped for that run and reverts to create pricing. Platforms are unaffected — their bound_count and new_shortcut_count are deliberately not stamp-gated. This follows directly from the None-not-0 rule above and errs long, the safe direction; correcting it would mean asserting membership the plan does not have. A unit whose bound_count is absent (older backend, unstamped or virtual collection) prices as all creates, the pre-#1511 behaviour. The live-countdown weights are unaffected and stay predicted_skip ? 0 : (collapsed_count ?? rom_count).
  • Measured live countdown (takes over within seconds). Once the apply is underway, syncEta.ts measures the real rate from the applying frames — one throttled sample per second over a ~30 s sliding window — and projects remaining = (planned_total − processed) / rate, rendered rounded up ("9 min left") so it never promises less time than it expects. It replaces the static seed as soon as the window spans enough real time to trust the slope (a couple of samples across a few seconds). Because it reflects the actual mix of cheap-update vs. full-create work, it is far closer to reality than the ceiling and ticks down as the run proceeds.
  • Estimator-owned sticky deadline. The estimator, not the UI, owns the displayed value: each fresh measurement re-anchors an absolute wall-clock deadline, and the countdown renders max(0, deadline − now). This is what keeps the readout smooth. The raw measurement re-arms to "not ready" between measurement segments, and a run's tail of small units each finishes inside the readiness window — so a raw snapshot would blink back to the static "up to X" seed for the whole tail; holding the last good deadline through those gaps keeps the countdown honestly ticking down instead.
  • Segment break across fetch gaps. Applying frames arrive roughly every second during real apply work, so a silence longer than ~10 s is always a unit or fetch boundary, never apply progress. Crossing it starts a fresh measurement segment (the prior samples are discarded), so the slope is never measured across the gap — pairing a pre-gap sample with a post-gap one would be a tiny item delta over a long span, an absurd rate that would briefly spike the countdown.
  • Stalled-prefix trim (shorter stalls). A stall too short to break the segment used to poison the window head instead (#1511). The applying stage is entered — and its first frame emitted — before the one-time shortcut scan, which then blocks for ~9–14 s (it scales with the existing library, not the delta). That frame became sample #1 and paired with the first post-scan sample: a couple of items across the whole scan, a rate ~10x below the real one, which the sticky deadline then held until the 30 s window slid past — the reported "6 min → 2 min" collapse. Leading samples are therefore dropped when the head gap is both wide (> 3 s, several sampling cadences) and unproductive (fewer than one item per sampling interval, below anything the 50 ms-paced apply loop produces). Both conditions are required: a wide gap that carried real throughput is slow work, not a stall, and discarding it would measure only the fastest stretch. Trimming can legitimately leave a single sample — "no measurement yet" is the honest answer while a stall is all the window has seen, so the static seed keeps standing until clean samples accumulate.

The coarse QAM progress bar reuses the same run weights (weightedCoarseFraction, syncEta.ts): each unit's bar width is its item-weight share — skip-aware at plan time, corrected to the real delta as units dispatch — with the within-unit fill scaled by the running unit's share, so a predicted-skip platform occupies no width and a huge platform fills the bar in proportion to its real work. It falls back to the old equal-per-unit index weighting when no plan is measured (QAM opened mid-run before any sync_plan, an older backend) or the plan can't apportion (unit-count mismatch, all-zero weights).

Two run-scoped latches keep the coarse fraction monotonic so the bar only ever moves forward. A run's leading zero-weight units still refresh covers, so rather than pinning the bar to zero they each claim an equal 1/totalUnits slice as a floor, with the weighted shares compressed into the band above it (#1506); that floor is held at its high-water mark, so raising a mispredicted skip off zero can't shorten the prefix and retract it. On top of that, the returned fraction itself is latched at the run's high-water mark by latchedCoarseFraction — the wrapper MainPage actually calls, keeping weightedCoarseFraction a pure reader (#1509): when observeUnitTotal corrects a mispredicted trailing skip up (0 → its real delta) it correctly grows the countdown's totalRoms (the run really is longer), but that shrinks the bar's completed/total ratio and would retract width already shown, so the bar holds while the countdown lengthens. Both latches touch only the bar output; the live ETA still reads the corrected total. The null fallbacks (no run, unit-count mismatch, zero total weight) pass through the wrapper un-latched.

The within-unit fill is itself split into three monotonic sub-slices (#1407, withinUnitFraction in frontend/src/utils/syncProgress.ts): fetch (FETCH_SHARE 15%) → covers (COVERS_SHARE 25%) → apply (APPLY_SHARE 60%). A unit is worked in that order — paginate the ROM list, download/refresh cover art, then create the shortcuts — and each phase fills its own slice by its own current/total, with a later phase's floor sitting at the sum of the earlier phases' shares. So the bar advances continuously through a unit's fetch and cover phases instead of resting frozen at the unit floor until applying, and it never jumps backwards at a phase boundary even though each phase restarts current/total from zero (each phase's frames land in a strictly-higher band than the phase before). The phase is tagged on the sync_progress payload's additive subStage field (camelCase, matching the sibling totalSteps / runId keys — the emit_progress Python kwarg is sub_stage, the emitted key is subStage): the fetcher's per-page frames carry subStage: "fetch", the artwork cover-refresh and download frames carry subStage: "covers" (both share the covers slice), and the frontend-driven apply frames are keyed on the applying stage alone — so a merged apply frame still carrying a stale subStage is unaffected. Within the covers slice the fill is not strictly monotonic: the covers phase runs two sequential passes — the cover-cache refresh first, then the download for the apply set — each counting current/total from zero, so a mixed unit (changed covers to refresh and new covers to download) can dip backwards within the covers band at the refresh→download handover, bounded to the covers share of that unit's slice. The fetch→covers→apply phase handovers themselves stay monotonic (the common cases don't dip either: a first full sync refreshes nothing, and a steady incremental sync with no cover changes runs neither pass). A frame with no sub-stage (the per-unit fetch anchor, or a pre-#1407 backend) rests at the unit floor, the old behaviour. emit_progress writes subStage into both the emitted event and the persisted get_sync_status snapshot, so it rides the QAM-remount re-seed path too.

The whole thing is an approximation by design — though a narrow one since the plan went skip-aware: seed weights and the applying frames usually both count post-collapse shortcuts now, the raw pre-collapse rom_count survives only as the fallback where the backend doesn't know better, and a unit the plan mis-predicted re-corrects on its first sync_apply_unit (observeUnitTotal). Estimate degradation is strictly one-directional: a wrong prediction or a stale collapsed count makes the readout run long or short, never changes what the sync applies. The readout wording ("up to", rounded-up countdown) keeps it an estimate, never a guarantee.

Save-sync serialization (device gate)

Every save-sync run on the device passes through a single device-level serialization gate (SaveSyncGate, in services/saves/sync_engine/_gate.py). Only one save-sync run is in flight at a time per device: when a second trigger fires while one is running, it queues — it waits for the in-flight run to finish rather than running alongside it. An asyncio.Lock is the serializer/queue; the gate owns only the bounded-acquire discipline around it (no run-lifecycle state, no run ids, no cancellation — those are out of scope here).

The wait is bounded so a stuck run never traps the launch path. Each of the four trigger methods on SyncEngine wraps its run body in bounded_run(timeout=…) with a per-trigger budget — pre_launch_sync (30 s), post_exit_sync (60 s), sync_rom_saves (15 s), sync_all_saves (60 s). If the gate can't be acquired within the budget the call returns its own fallthrough instead of blocking, and all four fallthroughs carry the same busy shape: reason: "sync_busy", no additive offline flag. A busy gate is a local wait — nothing on that path observed the server — so it never borrows a server-reachability slug and never sets the offline flag (#1625); the session-end toast would otherwise announce "Server offline" about a reachable server. Each caller routes the skip on success: False alone: the launch gate maps it to sync_failed (fallback-launch confirm, so Play is never trapped), the session-end toast keys on the reason. The lock is never leaked on timeout — a timed-out acquire releases any photo-finish hold before raising.

The gate sits outside the per-ROM lock (SyncEngine.rom_lock(rom_id)): the device gate admits one run at a time across the whole device, then each ROM still takes its own rom_lock for the read-mutate-write of its RomSaveSyncState. The cheap stateless early-out (is_save_sync_enabled()) runs before the gate, so a disabled-feature call never queues behind an in-flight run just to report it's off.

FirmwareService decomposition (services/firmware/)

The BIOS subsystem is a façade over four sub-services plus the demand reader they share. What the resolver says and what the RomM library holds are two different questions, and the split is drawn where they meet: demand.py answers the first and listing.py the second, and every surface below joins them rather than asking either source twice.

Module Role
service.py FirmwareService façade — public callable surface; wires the sub-services and delegates
listing.py FirmwareListing — the subsystem's only list_firmware caller and the owner of its TTL cache, in memory and in the SQLite firmware_cache table. A download and a delete both invalidate it through this one instance
demand.py FirmwareDemand — the machine's side: the per-platform verified reading, the whole-machine one the two callers with no platform to name fall back to, each file's destination under the BIOS root, the presence answer, and the rows for files the emulators want that the library never held. Sole owner of both resolver seams and of the BIOS root
status.py FirmwareStatusReader — get_firmware_status (which platforms the Library page can speak for, and nothing about any of them), get_platform_firmware_status (one platform's whole entry, the call that pays the reading) and check_platform_bios (the per-game check), the per-row build and the aggregates they answer from, and the per-platform delete count
downloads.py FirmwareDownloader — the three download entry points, the batch and its eligibility skips, and the downloaded_bios record each successful fetch leaves behind
deletion.py PlatformBiosDeleter — Delete BIOS, driven by those records and nothing else

The two resolver seams, the scope that tells them apart, and the reasoning behind every answer they produce are in the FirmwareService notes — the live firmware resolver section below.

DownloadService notes

RomM exposes three mutually exclusive file-layout flags on every ROM detail. They control how the server stores files and how the API serves them. The plugin maps each layout to a local on-disk path:

RomM flag RomM server layout What fs_name is Plugin local layout
has_simple_single_file roms/<platform>/<file> — one file, flat the filename flat in platform folder: roms/<platform>/<file>
has_nested_single_file roms/<platform>/<folder>/<file> — one file in a per-game folder the folder name flat in platform folder: roms/<platform>/<file>
has_multiple_files per-game folder with multiple files (multi-disc, BIN+CUE, etc.) the ZIP/folder name extracted into per-game subfolder named after the launch file: roms/<platform>/<launch-file-name>/...

has_nested_single_file quirk: fs_name is the parent folder name, not the filename. The actual filename with extension lives in files[0].file_name. The plugin reads from files[0].file_name so the downloaded ROM lands with the correct extension (e.g. Game.chd, not the extension-less folder name Game). A defensive helper falls back to fs_name and warns if files is empty or missing.

Why nested-single is flattened locally: a nested-single-file ROM has no sidecars by definition — RomM would mark it has_multiple_files if any companion files existed. The parent folder adds no value at the local layer, so the plugin drops it and stores the ROM directly in the platform folder, matching the simple-single-file layout. Multi-file ROMs keep their per-game subfolder because they contain multiple related files that belong together.

Extract-vs-flat gate keys on len(files) > 1, not on has_multiple_files: the plugin decides ZIP-extract vs single-file download with the is_multi_file_download helper (domain/rom_files.py), which returns len(files) > 1 OR has_multiple_files. This mirrors RomM's own download gate, which zips whenever the total file count is not exactly 1. RomM computes has_multiple_files from top-level files only, so the two counts disagree for a nested layout: a canonical Switch game (base file at the root plus update/ and dlc/ in subfolders) has exactly one top-level file (has_multiple_files=False, has_nested_single_file=True) yet more than one total file, so RomM serves a ZIP. Keying on has_multiple_files alone would take the single-file path and write the ZIP bytes verbatim into one unreadable .nsp. The boolean is kept as a defensive fallback for payloads that omit files; a genuine nested-single ROM has len(files) == 1 and correctly stays on the flat single-file path.

ES-DE directory-collapse rename: a multi-file ROM is extracted into a staging folder named after the ROM's identity — fs_name_no_ext, falling back to splitext(fs_name) (resolve_extract_dir_name in domain/rom_files.py) — never after files[0]. That distinction matters for a has_nested_single_file folder game served as a ZIP (a PS3 title whose first listed file is an arbitrary inner asset like a music file): resolve_local_file_name returns that asset's name for the local filename, but the extract directory must carry the game's identity or the whole install — launch file_path, rom_dir, and the folder-boot launch bake — inherits the wrong name. The staging folder is then renamed after the detected launch file including its extension (e.g. Example Quest - Second Journey (USA).m3u/ containing Example Quest - Second Journey (USA).m3u). ES-DE only collapses a directory into a single game entry when the folder name matches the launch file's full name with extension; without the rename a multi-disc game shows in ES-DE as a folder plus loose disc files. The launch file is only known after extraction (an .m3u may be auto-generated — see below), so the rename happens last, after launch-file detection, via es_de_collapse_rename (domain/rom_files.py) + the DownloadFileStore.move_dir whole-directory move. On a name collision (target already exists) the rename is skipped and the staging folder is kept — never clobbered or merged. Existing installs from before this feature keep their old folder layout until re-downloaded.

Folder-boot launch target (PS3): for a PS3 folder game detect_launch_file picks the nested …/PS3_GAME/USRDIR/EBOOT.BIN as the install's file_path — the correct launch file identity (save path, core, and displayed filename all derive from it). But RPCS3's directory-boot wants the game folder, not the EBOOT, so the baked launch_options carries the game directory instead. This is a bake-time path override (folder_boot_root, domain/rom_files.py) applied in the DiscLaunchResolver seam, never a file_path rewrite: file_path stays the EBOOT anchor while only the argument baked into the shortcut becomes the folder (ADR-0019, see Core and Emulator Selection).

Folder-boot launch target (PS3): for a PS3 folder game detect_launch_file picks the nested …/PS3_GAME/USRDIR/EBOOT.BIN as the install's file_path — the correct launch file identity (save path, core, and displayed filename all derive from it). But RPCS3's directory-boot wants the game folder, not the EBOOT, and RetroDECK's run_game.sh reinterprets a directory %ROM% as a "directory as a file" (run_game.sh:63-67) so it can never launch a bare folder. Two overrides fire together, both keyed on the same folder_boot_root fact and neither a file_path rewrite: the baked path becomes the game folder (folder_boot_root, domain/rom_files.py, in the DiscLaunchResolver seam), and the baked invocation becomes a direct sandbox command that bypasses run_game.sh — flatpak run --command=<launcher> net.retrodeck.retrodeck <args> "<folder>", resolved in ActiveCoreResolver.active_emulator_for_rom (standalone + folder-boot install → the SandboxLauncherFn seam, EsFindRulesAdapter.resolve_sandbox_launcher, gives the /app/…/component_launcher.sh sandbox path → EmulatorInvocation.direct). file_path stays the EBOOT anchor (ADR-0019, see Core and Emulator Selection). The multi-file download path also suppresses M3U generation and heals a .txt-suffixed PS3_DISC.SFB for a folder-boot dump (domain folder_boot_layout_root + DownloadService._maybe_heal_ps3_sfb_io).

M3U generation rule (needs_m3u in domain/rom_files.py): a game-named <fs_name_no_ext>.m3u is auto-generated (when no .m3u already exists) for multi-disc ROMs — two or more disc files of any kind (.cue/.chd/.iso) — so the emulator can switch discs, and for single-disc bin/cue ROMs — exactly one .cue — so the extract dir is renamed after a game-named playlist rather than a generically-named cue (disc1.cue/). Single-disc .chd/.iso arrive as single-file downloads that never reach the extraction path, so they get no playlist.

M3U is platform-gated on ES-DE's own extension list (ADR-0013). The file-count rule above only runs when the ROM's system actually supports .m3u. RomM bundles a platform-blind .m3u into the ZIP for every multi-file game, including cartridge systems (Switch .nsp, Xbox 360 .iso) whose emulators have no playlist concept — so an extension-only heuristic wrongly produced a <Game>.m3u/ folder that never collapsed. The plugin now asks whether ES-DE lists .m3u as a supported extension for that system, read from the same es_systems.xml ES-DE uses to decide directory-collapse, via AtlasCatalogueAdapter.system_supports_m3u(system) exposed through the SystemM3uSupportFn Protocol (services/protocols/) and threaded into DownloadService from bootstrap. When the answer is False, no .m3u is generated and the bundled one is never chosen as the launch file (detect_launch_file skips its .m3u preference), so selection falls through to the real game file and the folder is named <Game>.nsp/ / <Game>.iso/ instead. The capability crosses the service/domain seam as a plain bool — the domain functions (needs_m3u, detect_launch_file) take m3u_supported, never a system name or an adapter. The bundled .m3u is left inert on disk, never deleted. When es_systems.xml cannot be found the answer defaults to False (safe: a missing playlist only degrades disc-switching, a wrong one breaks the launch).

Launch-target validation

detect_launch_file ends in "largest file by size". When none of its format-specific rules match, whatever happens to be biggest becomes the launch target, is written into RomInstall.file_path, and becomes the shortcut's launch command — with nothing checking whether the system can act on it. A PS3 title distributed as .pkg + .rap is the reported case (#1582): a PKG is an installer, the game is still sealed inside it, no EBOOT.BIN exists anywhere in the download, so every rule misses and the multi-gigabyte package is baked. The download reports success and the failure only surfaces when the user presses play.

The verdict is decided once, at record time. _record_install_io — the single seam both the single-file and multi-file download paths pass through — calls is_launchable_target (domain/rom_files.py) and records the answer on RomInstall.launchable. The check reads the target system's live ES-DE accept-list through the SystemSupportedExtensionsFn Protocol, the same seam DiscLaunchResolver intersects the disc set with. Two cases pass without consulting it:

  • An empty accept-list — the source could not answer (unknown system, no ES-DE installation). A missing answer must never turn a working install into an unlaunchable one, so it accepts.
  • A folder-boot layout (FOLDER_BOOT_MARKERS) — the baked target is the game directory, not the nested EBOOT.BIN that file_path records, and ES-DE spells the directory case .ps3dir. The marker match is positive evidence that the plugin recognised the layout, so no extension is examined. Without this carve-out every working PS3 dump would be rejected: file_path ends in .bin, which ps3's accept-list (.desktop .iso .ps3 .ps3dir) does not carry.

Everything else is decided by the recorded launch file's extension. .pkg is absent from ps3's list; a bare track .bin is absent from dreamcast's (.cdi .chd .cue .dat .elf .gdi .iso .lst .m3u .7z .zip), which is why a multi-file GDI rip without a .cue is caught by the same rule.

The files are always kept. An unlaunchable download is a real install: the row is written, the ROM is uninstallable through the normal path, and nothing is deleted. Refusing the install would discard a package whose remaining use is exactly to be installed by hand in the emulator — the documented RetroDECK procedure for PSN titles — leaving the user worse off than the silent failure this replaces. What is withheld is only the launch command.

One seam withholds it everywhere. DiscLaunchResolver.resolve_bake_path returns "" for an install with launchable is False, before any disc work, and build_launch_options renders an empty path as the empty launch command. Every launch-bake site already draws its path from that resolver — library sync's installed_paths map, download-complete's re-bake, the core-change re-bake, the startup relaunch-options heal, the disc picker — so none of them can compose a command for content nothing can boot, and none needed a guard of its own. An unlaunchable install stays in the installed_paths map (it is downloaded, and collapse_sibling_groups reads the key set to choose a sibling group's representative) and maps to the empty path.

Empty is the established uninstalled state, not one invented here: addShortcut leaves a new shortcut's options untouched when the command is "" (frontend/src/utils/steamShortcuts.ts), the sync update/adoption path writes "" explicitly (rewriteShortcutIdentity, frontend/src/utils/syncManager.ts), and an uninstall records "" as the ROM's applied_launch_options (services/rom_removal.py). An empty launch command therefore means two things, and nothing may infer which: "not downloaded" and "downloaded but not launchable" are indistinguishable at the Steam-shortcut layer — the shortcut holds "" either way. Only the data layer separates them (no rom_installs row versus a row with launchable = 0), so a consumer that needs to tell them apart must ask the install record, never read emptiness as "not downloaded". The rule is restated on build_launch_options itself, where a later reader will hit it.

The frontend closes the loop in two places: the shared launch gate blocks with no_launch_target before any save-sync work — for both the Play button and the global launch watcher — and the game-detail page's ROM File section states that the download has no launchable format and that the files are on disk. Re-checking a recorded verdict once the user has installed the package in the emulator is separate work (#1654).

Where the knowledge comes from. The accept-list is the frontend's own per-system declaration, read live by the vendored emu-atlas resolver and handed over by AtlasCatalogueAdapter.get_supported_extensions, wired in as DownloadServiceConfig.system_extensions. The consuming call site is a single self._system_extensions(system) behind one Protocol-typed config field, which is what let the plugin's own duplicate parser leave in #1840 as one wiring change.

Filesystem writes go through DownloadFileAdapter. ZIP extraction is ZIP-slip protected and streamed: extract_zip copies each member in chunks and reports byte progress through an optional callback, so a multi-file ROM emits download_progress frames with status: "extracting" (bytes_downloaded/total_bytes over the uncompressed total, resumable: false) after the transfer hits 100%. The frontend reuses the same event — no new event name — to switch the download button and QAM queue into the non-cancellable Extracting… phase. Single-file downloads never emit it. No download_progress frame follows a download's terminal frame (download_complete, download_failed, or the cancelled frame): a progress tick or resumability verdict the worker queued before the end can reach the loop after it, and DownloadService._live_entry drops it once the entry has been evicted or its status is terminal.

Bounded concurrency + reserved-bytes pre-flight: at most two ROMs transfer at once, gated by an asyncio.Semaphore(2) around the transfer + post-IO critical section. start_download enters the queue with status queued and reserves the download's required bytes in _reserved_bytes[rom_id]; _do_download flips the status to downloading only once it acquires the semaphore (emitting a download_progress status: "queued" frame first if it has to wait), and releases the reservation in its finally. The disk pre-flight accounts for siblings' outstanding reservations (free_space - sum(reserved) < required) so two concurrent downloads that each fit alone but not together can't both pass — the second is rejected with an insufficient_space failure.

Cancel reaches the UI and never destroys a live install: cancelling a download emits a terminal download_progress status: "cancelled" frame so the frontend resets the button out of its downloading state (the cancel path used to be silent). Because executor threads run to completion regardless of cancellation, a cancel that loses the race to a just-committed install is reconciled (_reconcile_post_io awaits the in-flight post-IO future): if the install committed, the download is surfaced as completed (launch options baked, download_complete emitted) rather than torn down. _cleanup_partial_download removes only the transient transfer artifacts (.zip.tmp / .tmp) and, for a multi-file ROM that did not commit, the extract dir(s) this download created — it never deletes the bare target_path, so a re-download that fails mid-stream (or a cancel that lost the race) cannot destroy a pre-existing or just-committed install.

Sibling supersede — one downloaded version per group (ADR-0021 §4 amendment, #1298): before a download begins, start_download (and resume_download) strips any other installed member of the ROM's sibling group so a game keeps at most one copy on disk. _conflicting_sibling_install_ids reads the group in one short UoW (indexed iter_by_group_key) and returns members that are installed and either unbound or bound to the same shortcut — a member bound to a different shortcut (a grandfathered duplicate, ADR-0021 §5) is exempt and never removed. Each superseded install goes through the canonical RomRemovalService.remove_rom (files + rom_installs row; saves untouched per ADR-0007), the membership read UoW closed first so the removal's own UoW does not nest (ADR-0006); a not_installed result raced clean and is skipped, any other removal failure aborts the download so the invariant stays honest. A superseded sibling's paused queue entry is evicted (_evict_if_paused) so a stale .tmp can't later resume into a second install. resume_download additionally refuses (superseded, entry dropped) when a switch has moved the group's binding to a different member since the pause — it won't re-download a version the picker has already left. The in-progress claim is taken before the supersede await and released on every early exit so a concurrent start_download for the same rom is rejected rather than racing past it.

ConnectionService notes

The minimum-version gate is SemVer-aware on SYSTEM.VERSION. domain.version.meets_min_version compares the numeric core against _MIN_REQUIRED_VERSION; when the core equals the floor, a -alpha / -beta suffix (case-insensitive, optional .N build number) ranks below the release and is rejected — so 5.3.0-beta.1 fails at floor 5.3.0 while 5.3.1-beta passes. development and a missing version bypass the gate. Why the floor sits where it does is recorded in ADR-0040.

A Client API Token is bound to the server it was minted against. When the token is minted, the canonical origin of romm_url (full scheme://host[:port], default ports folded out, path/query dropped — lib/url_host.normalize_origin) is stored alongside it as romm_api_token_origin. RommHttpAdapter.auth_header() attaches the bearer only when that stored origin matches the current romm_url origin; on a mismatch it raises TokenHostMismatchError instead of sending the credential to a host the token was not minted for. The error is non-retryable and maps to a config_error failure (Your saved RomM login is for a different server. Sign in again to continue.), so every data flow fails fast until the user re-signs-in. https://h and http://h are deliberately different origins — a plaintext downgrade is a different destination, not the same one. A legacy token minted before origin stamping carries romm_api_token_origin = None and is treated as un-bound: it is still attached (never blocked) so existing installs keep working until their next sign-in stamps the origin.

Sign-in ordering: validate → probe → mint → persist (one atomic save). establish_token trims the entered URL and rejects a non-http(s) value before any network call. It then holds the candidate URL in memory only — clearing the stored token in memory first so the version probe never carries the old server's bearer to the candidate host — and persists nothing until the mint succeeds. On any failure (probe unreachable, version too old, forbidden/error mint, no usable token, or a disk error) the in-memory auth state is rolled back to the previous working URL + token, and because disk was never touched the prior working credentials survive a failed sign-in. Only a successful mint commits romm_url + SSL flag + token + id + origin to disk in a single save_settings() call.

The old-token DELETE is origin-guarded and provenance-guarded. RomM scopes a Client API Token to the account, and re-auth deletes the device's previous token. That DELETE is only fired when the old token's stored origin matches the new URL's origin (same-server re-auth) — replaying it against a different server would delete an unrelated token there, so the DELETE is skipped (and logged) when the origins differ or the old origin is unknown. It is also skipped whenever the previous token's romm_api_token_source is "user": a token the user pasted belongs to the user, not to this device, and must never be revoked by the plugin (a user token also carries no romm_api_token_id, so there is nothing to DELETE by anyway — the source check states the intent). The DELETE uses Basic auth from the one-time credentials, unaffected by the cleared bearer.

A token's provenance is recorded in romm_api_token_source. The value is "minted" for a token the plugin minted from a username/password (establish_token and the startup migrate_legacy_credentials) and "user" for a token the user pasted (establish_user_token). It is part of the snapshot/restore auth-state set, so a failed sign-in rolls it back with the rest of the auth state. Settings migration v9 → v10 seeds it — "minted" when a token already exists (a pre-v10 install could only hold minted tokens), else None.

Pasted-token sign-in (establish_user_token) — the OIDC path. OIDC / SSO accounts have no password to mint from, so the user creates a Client API Token in RomM's web UI and pastes it. establish_user_token mirrors establish_token's validate → probe → gate → validate → persist-on-success-only shape, but the credential is the pasted token rather than a fresh mint, so there is no mint and no server-side DELETE of any prior token. The entered URL is trimmed and rejected if non-http(s); a blank/whitespace token is rejected as config_error before any network call. The candidate URL and the pasted token are held in memory only, with romm_api_token_origin stamped to the candidate origin and romm_api_token_source = "user", so the auth-header guard attaches this token (not the old server's bearer) to the validation probe. Validation is an authenticated GET /api/users/me: a 401 means the token is invalid or revoked, a 403 means it authenticates but lacks a required scope (the plugin cannot introspect a token's granted scopes, since /api/users/me's oauth_scopes reflects the user's role, not the token's grants, so this connect-time authenticated probe is the only validation). Both map to the canonical auth_failed failure with a token-specific message. On any failure the in-memory auth state is rolled back and disk is never touched; only a successful validation commits URL + SSL flag + token + id = None + origin + source in a single save_settings(). The token value is never logged. The device-forget-on-origin-change and playtime-scope-notice clear run on the success path exactly as they do for establish_token. The /api/users/me validation + persist tail is shared with the paired-token path below via the private _validate_and_persist_user_token helper (both hand the plugin a token to validate and store with "user" provenance), so their observable behaviour stays identical.

Pairing-code sign-in (establish_paired_token) — the OIDC path without pasting. The same OIDC accounts can sign in by entering a short-lived RomM pairing code instead of pasting the token: the plugin exchanges the code for the token over the public, unauthenticated POST /api/client-tokens/exchange endpoint (the one-time code is itself the credential). It mirrors establish_user_token's validate URL → probe version → gate → obtain-credential → validate via /api/users/me → persist-on-success-only shape, but the credential is fetched by the exchange rather than pasted, and the exchange runs with the token trio cleared in memory (like establish_token) so no old bearer leaks to the candidate host during the unauthenticated call. The code is normalized the way RomM normalizes it — all whitespace and - stripped, then uppercased ("ab-cd ef23" → "ABCDEF23") — and a code that is blank after normalization is a config_error before any network call. The exchange is never auto-retried: a pairing code is single-use, and a replay would burn both the code and RomM's per-client rate limit, so the transport's unauthenticated_post_json skips with_retry. Each RomM rejection maps to a distinct, actionable message: an invalid/expired/used code, a token-no-longer-exists 404, and a disabled-owner 403 all carry auth_failed; a 429 carries the bespoke rate_limited reason. On success the freshly rotated raw_token (the exchange regenerates the token server-side, so any previously copied raw value stops working) is host-bound in memory and run through the shared _validate_and_persist_user_token tail, so it is persisted with "user" provenance and id = None exactly like a pasted token — no mint, no server-side DELETE. The pairing code and the returned token are never logged.

The registered device id is forgotten only on a genuine origin change. A device registered with RomM (POST /api/devices, its id stored in kv_config["device_id"]) is bound to the server it was registered against — RomM's negotiate save-sync transport hard-404s a foreign device id. So a successful sign-in that genuinely changes origin forgets the stored device id (SaveService.forget_device → DeviceRegistry.forget_device, wired as the DeviceForgetFn injected into ConnectionService), so the next save-sync run re-registers against the new server. The decision is is_origin_change(old_token_origin, new_url) in lib/url_host — both origins are normalized and the id is forgotten only when both are known (parseable) and differ. This is deliberately the opposite failure posture to the old-token DELETE guard (same_origin, which fails closed): a same-server re-sign-in keeps the device identity, including a token swap on the unchanged URL, URL-formatting variants, and a None/unstamped old origin (an unknown old origin is treated as unknown, not different — never a change). Without this, a same-server token swap dropped the identity and the next post-exit sync flagged a spurious conflict for a save this device itself synced (#1437). The old_token_origin comparison input is captured from the pre-clear auth-state snapshot, so the leak-safety in-memory token clear (which zeroes the origin) never poisons the comparison. The forget is local-only (the row on the old server is left for RomM's machine-id dedup to reconcile on a later re-registration) and best-effort on the success path — a failed local clear never turns a good sign-in into a failure, and a failed sign-in (in-memory snapshot restored) keeps the still-current server's id. ensure_device_registered never deletes the id on any failure, so a permission-degraded (403/timeout) session leaves the identity intact. This is Phase 0a of the RomM Device Sync negotiate adoption (ADR-0016).

The signed-in user's own id (romm_user_id) is bound to the token — stamped at sign-in, backfilled lazily, cleared on sign-out. It drives the collection owner scope (build_work_queue + get_collections, described in the LibraryService section above). Every sign-in path re-derives it from the freshly authenticated token so it can never linger for a different user or a different server: the mint path (establish_token) probes GET /api/users/me after host-binding the minted token (the mint response carries only the token id, not the user id); the pasted/paired paths reuse the id from the /api/users/me validation probe they already run (no second call). In every case the id is set in memory before the token-persist save, so it rides the same single atomic save_settings() — the sign-in write shape is unchanged. It is part of the snapshot/restore auth-state set, so a failed sign-in restores the previous id, and it is cleared in each path's in-memory auth clear so a probe failure or a malformed payload leaves it None rather than stale — so the own scope drops nothing until the next backfill. Existing installs (a valid token minted before this setting existed) carry no id; test_connection lazily backfills it — when a token is present and the id is missing, the successful connection check probes /api/users/me and persists the id (its own save), so the scope takes effect without a re-login. A known id needs no network on later checks. Every identity read is best-effort: a failure never fails the sign-in or the connection check.

The no-sign-in URL change path (SettingsService.save_server_url) deliberately does not touch the token, so pointing the URL at a different origin leaves the stored token's origin mismatched and the auth-header guard makes subsequent data flows fail fast with config_error until the user signs in again. Because every data flow (including device registration) is inert until that sign-in, a stale device id from this path cannot be used before the sign-in clears it.

Sign-out (sign_out) is a local forget — never a server-side delete. sign_out clears the token quad (romm_api_token / romm_api_token_id / romm_api_token_origin / romm_api_token_source) in the in-memory settings dict and persists them in a single save_settings() (the same atomic-write funnel _persist_token uses), keeping romm_url and the SSL flag so the user need not re-enter them. It mirrors the sign-in paths' persist discipline: the auth state is snapshotted first, and a failed save rolls the in-memory quad back to the snapshot and returns the canonical failure shape (via error_response), so a disk error never strands the user with a half-forgotten but still-valid token. Only on a successful save does it drop the cached RomM server version (set_version(None)) so no stale value lingers. It is synchronous — no run_in_executor, no network — and idempotent (signing out when already signed out still returns success). It never issues the server-side token DELETE: a minted token deliberately lacks me.write (deleting it would require re-entering the password, which sign-out does not have), and a user-supplied token belongs to the user. That is why the direct re-authentication path stays as Sign in again (not sign-out-then-in): establish_token's same-origin minted-token cleanup (#1038) only fires while the old token id is still stored, so signing out first would strand the old minted token on RomM (which caps tokens per user). After sign-out, test_connection returns the canonical "Not signed in" config_error because romm_url survives but the token is gone. Signing out (like re-signing-in) while a sync or download is in flight needs no guard: once the token is gone the in-flight operation simply fails authentication and surfaces its normal error — deliberate, not a race to defend.

Server-supplied paths are validated, fail-stop on traversal: every server-supplied path component — the firmware file_name, the ROM platform slug, and post-extraction URL-decoded ZIP member names — is checked through lib/path_safety (safe_join for realpath containment, safe_path_component for single-component names) before any write. A traversal attempt (../, an absolute path, or a %2e%2e%2f-encoded ZIP member that decodes to ../ after the pre-decode ZIP-slip check passes) aborts the whole download rather than skipping the offending entry: already-extracted members are cleaned up (no half-installed ROM), a canonical {"success": false, "reason": "path_traversal", "message": ...} failure is returned, and the download_failed event fires so the UI doesn't hang on "downloading". Firmware downloads surface the same canonical failure from download_firmware.

StartupHealingService notes

Beyond the disk-prune and orphaned-SyncRun reconciliation, this service owns the startup launch-options reconcile (#1043). launch_options (the full Steam-shortcut launch command) is written only event-driven — at sync, at download-complete, and on RetroDECK-home migration (ADR-0009) — so any path that misses its bake leaves an installed shortcut stuck on the "" placeholder, and bin/tender-rom-launcher then runs with no args and exits non-zero. There was no backstop short of a Force Full Sync or uninstall/reinstall.

get_installed_relaunch_options() is the read half of the fix: a 0-arg read whose items are one {app_id, launch_options} for every ROM that is both installed (has a rom_installs row) and bound (its roms.shortcut_app_id is set). It snapshots the install/ROM rows in one short read UoW, then re-bakes each command outside that UoW through the same active_core / disc_resolver seams every other bake site uses. Each seam is outside for its own reason: resolving active_core inside the iteration UoW would deadlock, since ActiveCoreResolver.active_core_for_rom opens its own UoW (the per-connection write lock is not re-entrant), while disc_resolver opens none and instead walks the install directory — and BEGIN IMMEDIATE takes the write lock even for a read, so a walk inside the UoW stalls every other writer for its duration (#1779). Uninstalled and unbound ROMs are skipped by construction. The endpoint is read-only and checks the prune rule but not the migration rule; a non-empty answer carries an installed_reconcile lease for the frontend's writes.

The frontend pulls this on mount, once the backend is proven reachable (it reuses the app-id/metadata init's retry/backoff, not a second loop), and confirm-sets each entry via the existing setLaunchOptionsConfirmed fire-then-poll. The pass is idempotent and appId-safe: re-confirming a correct command matches the read-back instantly, and a launch_options write does not change the shortcut's appId, so artwork, collections, and the shortcut identity survive. This heals drift from every cause at the next plugin load.

UpdateCheckService notes

The service answers one question — is a newer Tender release out — and stays silent whenever it cannot say. It asks through LatestReleaseFn, whose adapter reads GitHub's "latest release" route: drafts and pre-releases are excluded by that route, a tag other than tender-v<version> is not a Tender release, and the release counts as available only once its romm-tender-<V>.tar.gz asset is attached AND GitHub states a valid sha256 digest for it (64 hex digits after sha256:): a tarball that cannot be verified before an install makes the release not available, and a stored answer without a valid digest is dropped when it is read back. The assets job uploads the tarball minutes after the release is published, so for that window the release is read, answered and passed over: the last available release the checks found stays standing, and a check that reached nothing at all does the same.

  • At most once a day, when the panel loads. The last answer lives in kv_config under update_check_last_seen — its version, the tarball's browser_download_url and the lowercase sha256 hex of its digest, which name one release and stay paired, plus the time of the last attempt. The stamp records the ATTEMPT, so an offline start does not pay the ten-second timeout at every panel load; a stamp dated in the future is due at once.
  • The user's two keys are in settings.json and are written only through the SettingsPersister: update_check_enabled (absent means on) and update_notice_dismissed_version, which holds a version rather than a flag so the next release raises the card again. With the switch off nothing is requested — not by the daily check and not by Check now either.
  • Check now (check_for_update_now) skips the throttle and forgets the dismissal, and adds reached to the answer, so the Settings section can tell "nothing newer" from "nothing found out".
  • What the request carries is the program's User-Agent (romm-tender/<version>, from domain/identity.py) and GitHub's JSON Accept header; GitHub sees the machine's IP address like any other request, and nothing else about the user.

Two answers come from the environment rather than from the service, and the entry point resolves both once (domain/update_release.py::resolve_update_source) next to the directories. TENDER_RELEASE_API is the installer's own test seam, read under the same name with the same default, so a unit drop-in points both at one fake server. installed_program is True only where TENDER_CODE_DIR names the directory the running code actually sits in: only the installed unit sets that variable, and requiring it to name this process's own directory keeps a checkout from claiming to be the install because a shell exported the variable for an installer test. A run from a checkout checks and shows the card like any other, and the Settings section shows a line naming this a development build.

Adapters (backend/adapters/)

Adapters own all I/O, and implement the Protocols defined in services/protocols/ wherever a service puts the question through a seam. Selected adapters:

Module Role
romm/http.py RommHttpAdapter — HTTP transport: auth, SSL, User-Agent, configured proxy headers, platform map, error translation
romm/retry.py RetryLadder — the attempt policy behind that transport: retry ladder, backoff, known-unreachable state
romm/romm_api.py RommApiAdapter — RomM REST surface (saves, ROMs, platforms, firmware, devices, play-sessions) over the HTTP transport
steam_config.py SteamConfigAdapter — Steam VDF read/write, grid dir, shortcut icon write, Steam Input config
steamgriddb.py SteamGridDbAdapter — SteamGridDB REST client
github_releases.py GithubReleaseAdapter — LatestReleaseFn: GitHub's "latest release" route, the tender-v tag and the release's romm-tender-<V>.tar.gz asset with its digest; every failure answers None
sgdb_artwork_cache.py SgdbArtworkCacheAdapter — on-disk SGDB artwork cache
cover_art_file_store.py CoverArtFileStoreAdapter — RomM cover art I/O across the per-ROM cover cache and the Steam grid dir (download, copy_file publish/seed, read, prune)
persistence.py PersistenceAdapter + per-domain persister adapters — settings.json read/write plus the one-time legacy save_sync_state.json read that feeds the bootstrap settings fold
repositories/ SqliteUnitOfWork + per-aggregate repository adapters — SQLite I/O (the live persistence path; see Database Design)
sqlite_migrations.py apply_migrations — schema migration runner (db/migrations/NNN_*.sql, PRAGMA user_version)
download_file.py DownloadFileAdapter — download filesystem
firmware_file.py / migration_file.py / rom_files.py / save_file.py per-subtree filesystem adapters (BIOS, RetroDECK migration, ROM removal, local saves)
retrodeck_paths.py RetroDeckPathsAdapter — reads retrodeck.json for ROMs/saves/BIOS/home paths
recovery_bundle.py RecoveryBundleAdapter — safe recovery-root derivation, exact source measurement, verified copies, checksums/human manifests, fsync, staging cleanup, and atomic bundle sealing
prune_artifacts.py PruneArtifactAdapter — per-ROM cover/validator and SteamGridDB cache discovery/removal
steam_recovery.py SteamRecoveryAdapter — per-shortcut Steam grid, Steam Input, and controller-setting recovery inventory plus post-confirmation cleanup
atlas_catalogue.py AtlasCatalogueAdapter — the ES-DE emulator catalogue through the vendored emu-atlas resolver (the picker's list, the system-layer default, the libretro active core, the per-system accept-list). Sorts the resolver's effective order back to the declared one, so no gamelist selection moves the default (ADR-0030)
atlas_saves.py AtlasSaveLocationAdapter — where one ROM's save lives and what it consists of, asked of the catalogue entry the plugin resolved for it and always with the ROM's own content path, since the answer turns on the content file's extension (ADR-0034). Caches no answer; holds only the installation handle. Also describe_core_probe_interpreter(), the start-up log line naming the interpreter the resolver's core probe would run under
es_find_rules.py EsFindRulesAdapter — ES-DE es_find_rules.xml: whether a standalone emulator's binary is installed, and the sandbox component launcher the folder-boot bake execs
gavel_native.py GavelNativeAdapter — loads the compiled romm-gavel core (backend/native/libgavel-x86_64-linux.so) via ctypes; is itself the ResolveUploadConflictFn seam and provides the ComputeSyncActionFn seam, the two save-sync decisions (no Python fallback)
system_clock.py / system_uuid_gen.py / asyncio_sleeper.py concrete Clock / UuidGen / Sleeper seams
hostname.py / path_probe.py / debug_logger.py hostname, the generic path seams (exists, symlink-resolve), settings-aware debug logger. The program's own name and version are not read from anywhere — they are constants in domain/identity.py
renderer_rss.py / renderer_gc.py RendererRssFn — max steamwebhelper VmRSS from /proc; RendererGcFn (HeapProfiler.collectGarbage) over the CEF debugger. The session-budget measure + settle seams (ADR-0024). The "free memory" action is a frontend SteamClient.User.StartRestart, not a backend adapter
game_process.py GameProcessControl — resolves a flatpak app's live instances via the per-user registry (info / bwrapinfo.json) plus the /proc child walk, reporting each tree's PIDs and argv separately, and signals them. Direct reads + os.kill, no subprocess; fail-soft on every read

RommHttpAdapter notes: one retry ladder per call stack

How often a request is attempted is not the transport's own question: RetryLadder (romm/retry.py) owns the ladder, its re-entrancy scope and the reachability state, and RommHttpAdapter holds one and delegates with_retry / is_retryable to it — the two methods that make the adapter a RetryStrategy.

with_retry runs up to 3 attempts with a base_delay * 3^attempt backoff (1s then 3s — a 3-attempt ladder has only two gaps) and only retries what is_retryable calls transient. It is reached from two levels: most of the adapter's own request methods wrap themselves in it, and services wrap adapter calls again through the RetryStrategy protocol (satisfied by RommHttpAdapter itself). The outermost ladder wins — a re-entrant call runs its function straight through, guarded by a thread-local flag on the ladder, because the blocking work runs one call per executor thread. Without that guard the two levels multiply into 9 HTTP attempts with both backoffs stacked, which is what made a single game-detail page take ~30 s to fill in against an unreachable server.

The guard lives in with_retry rather than in the ~20 service call sites so that a twenty-first site cannot re-introduce the nesting, and because the service-level wrap is the only ladder for the adapter methods that deliberately have none (upload_multipart, request_once, unauthenticated_post_json, basic_auth_request) and for the coarse wraps whose callee paginates or chains several requests — those keep being retried as the whole unit their caller meant.

The ladder also remembers one bit of transport state: the server is known unreachable. A ladder that gives up on an exception that means the server could not be reached sets it; while it is set every ladder shrinks to a single attempt with no backoff, and a ladder already sleeping cuts its backoff short (the lanes a game-detail page opens start simultaneously, so this is what makes the first load after an outage fast, not only the second). The degraded ladder still really performs its call — it is never skipped — which is what makes the state self-healing and what makes the UI's Retry button work with no path of its own. Any successful response clears it, including on the ladder-bypassing paths, because every request funnels through one _urlopen choke point; the reachability probe and the 30 s heartbeat therefore clear it for free. A new request method that opens its own connection would silently stop clearing it and leave the plugin degraded, so scripts/check_urlopen_choke_point.py confines urllib.request.urlopen to _urlopen. An HTTPError deliberately does not clear it either: a 5xx answered by a proxy in front of a dead origin is one of the shapes that counts as unreachable in the first place.

An error carrying a 4xx status code never sets it — the server answered, whatever it answered. That peel runs before classify_error, which cannot draw the line itself: it is a user-messaging classifier that folds every unbranched RommApiError onto server_unreachable so a display string always exists, and reusing it raw as a transport verdict would import that coarseness. The 409 is the case that makes this load-bearing rather than tidy — every automatic save upload POSTs overwrite=false precisely so RomM can reject a stale head with one, and that upload runs inside a ladder, so a routine conflict would otherwise mark the whole server unreachable. The 404 is peeled for the same reason and matters for a second one: it is deletion authority downstream (next section). An unproven 404 does still set it, because it degrades to a plain RommApiError carrying no status code — nothing proved RomM answered, which is the same fail-open reading that denies it deletion authority.

The state is about the configured RomM server only. download_external fetches a ROM's url_cover from a third-party metadata CDN, so its ladder is entered with romm_origin=False and takes no part in either direction: a dead cover CDN must not degrade every RomM call, and reaching the CDN is no evidence that RomM came back. It also keeps its full ladder while RomM is down, because it is a different host.

RommHttpAdapter notes: the headers every RomM-origin request carries

_apply_origin_headers is the one place that attaches what a request carries by virtue of its DESTINATION rather than its purpose: the plugin User-Agent and the user's configured proxy headers (settings.json romm_custom_headers,

1822 — for a RomM server behind Pangolin, Cloudflare Access, Authelia or Authentik forward-auth, which rejects every

request before RomM sees it). _apply_default_headers is that helper plus Authorization, and the two sign-in paths that deliberately omit the bearer — unauthenticated_post_json (the pairing-code exchange) and basic_auth_request (the Client API Token mint) — call the helper directly. All three attachment points are deliberate: an authenticating proxy sits in front of sign-in as well, so a plugin that added the headers only to authenticated calls would leave a user unable to authenticate in the first place.

The one deliberate exclusion is download_external, which fetches a ROM's url_cover from a third-party metadata CDN and keeps the bare User-Agent it always had. A proxy access token is a credential for the user's own front door and has no business reaching a foreign host — the same reasoning that keeps the RomM bearer off that request.

A configured header may never displace one the adapter sets itself, and two mechanisms hold that — unevenly.

The first is validation: the names are refused case-insensitively (domain/custom_headers.py, RESERVED_NAMES), with authorization carrying a message of its own, because a proxy's own documentation suggests that header and it is exactly the one the RomM bearer occupies. It is reached from two places — resolve_custom_headers for what arrives over the wire, and stored_custom_headers, the single reading of the persisted list, which re-runs it on every request so a hand-edited settings.json cannot put a reserved name on the wire either. Those two gates are one mechanism read twice, not two: both consult the same frozenset, so removing a name from it opens both in a single edit.

The second is ordering: the configured headers go on FIRST, so a later req.add_header wins the same name whatever came before it. This one is partial, and the gap is worth knowing. It covers only the names the adapter re-adds after _apply_origin_headers — User-Agent, Authorization, Content-Type, and conditionally Accept-Encoding (on the JSON GETs only), Range, If-None-Match and If-Modified-Since. It does nothing for Host or Content-Length, which the adapter never sets: http.client supplies both itself and skips its own when the caller already did. So for host — reserved precisely because http.client._send_request suppresses its derived Host when the caller supplied one, which would retarget every request's virtual host — the frozenset is the only defence that exists.

Nothing pins the ordering leg. test_a_stored_authorization_never_displaces_the_bearer and test_a_stored_host_never_retargets_the_request read as though they do, but both pass because stored_custom_headers has already dropped the entry before add_header is reached; reverse the order inside _apply_default_headers and both stay green.

One thing the attachment rule does not reach at all: what the plugin attaches is not the same as what arrives. _urlopen uses urllib's default opener, whose redirect handler copies every header except content-length / content-type onto the followed request with no same-origin test, so a 30x from the RomM origin to a foreign host delivers the configured header — and the RomM bearer with it. That is the transport's long-standing behaviour rather than anything this feature introduced, and it is tracked in #1889.

The values are read from the live settings dict at call time, the way romm_url is: a value captured in __init__ would go stale the moment the user edits it. They never reach a log line — CustomHeader.__repr__ prints the name alone, and no refusal message quotes a value.

RommHttpAdapter notes: what makes a 404 an entity verdict

RommNotFoundError means "RomM's entity layer says this entity does not exist", and downstream that reading is authority: the removed-game cleanup deletes on it, the version picker marks a version vanished on it, save-sync drops a stale device registration on it. A bare HTTP 404 does not carry that meaning on its own — FastAPI answers a misconfigured path prefix with the same status and a generic {"detail": "Not Found"} body, and a reverse proxy in front of RomM (Cloudflare Tunnel, Traefik) answers a misroute with an HTML or empty one. So on the API routes the adapter raises RommNotFoundError only for a response that proves it came from RomM's entity layer: a JSON content type whose body parses to an object carrying a detail string that is neither blank nor FastAPI's stock Not Found (matched case-insensitively — a real entity answer always names the entity, so it is never the bare phrase). The requested id is deliberately not parsed back out of that detail — its wording moves between RomM releases while the generic default stays put, so blocklisting the default is the robust test. Every other 404 shape degrades to a plain RommApiError, which classify_error maps to server_unreachable, so infrastructure can no longer authorize a deletion and each caller fails open on it instead.

The byte-stream routes — download, download_conditional, download_external — opt out via translate_http_error(..., asset_route=True) and keep the plain status mapping. Their 404 answers about a file, not an entity, so it is never deletion-authority grade; and it has to keep raising RommNotFoundError, because RomM serves its cover resources from a static mount where a genuinely missing cover answers with exactly that generic route-404 body, and the sole consumer reacts by refetching from the ROM's url_cover (the fallback above) — non-destructive and self-correcting rather than a deletion. Each of the three call sites carries that constraint as a comment: a future unification of the two classes would silently disarm the fallback, so both sides are pinned by tests.

Every deletion-authority probe reaches the network through the JSON-API entry points, never a byte-stream one: get_rom / get_rom_once (request / request_once) behind the version picker's liveness and vanished probes, list_saves (request) behind the save-status, copies and slot-setup reads, and update_device (put_json) behind the device re-registration. The one consumer that branches on a byte-stream 404 is the url_cover fallback.

The entity-answer shape is captured from RomM 5.1.0, and RomM 5.3.0 raises the same one — a FastAPI HTTPException(status_code=404, detail=…) naming the entity — for a missing ROM (exceptions/endpoint_exceptions.py), save (endpoints/saves.py) and device (endpoints/device.py). If a future release's entity-404 ever carries a different body shape (no JSON content type, no detail string), every entity-404 on that server degrades to server_unreachable: the cleanup never confirms a ROM gone, no version is marked vanished, no stale device registration is dropped. That is the fail-open direction by design — if a user on such a server reports exactly that symptom set, this paragraph is the explanation, and the fix is a version-aware entity check, never a return to trusting the bare status.

PersistenceAdapter notes

  • File locking: write methods acquire an exclusive fcntl.flock before touching the file, preventing concurrent writes from corrupting state.
  • Schema versioning: every state file written includes a version field. On read, a mismatch causes the file to be treated as absent (cache discarded, state reset to defaults) rather than loading incompatible data.
  • Crash-safe atomic writes: settings.json is written with the durable write-tmp → fsync(tmp) → os.replace() → fsync(dir) recipe. The temp file's bytes are forced to disk before the rename, and the directory entry the rename creates is forced to disk after it. This closes the power-loss window on the Steam Deck's ext4: without the fsyncs, a crash after the rename but before the kernel flushed could leave a truncated or empty settings.json — which boot rewrites every run, so the window recurred. The directory fsync is best-effort: on the rare filesystem that rejects it, the error is logged at debug and swallowed (the content is already durable via the temp-file fsync).
  • Corrupt-file quarantine (never silently factory-reset): a FileNotFoundError on read is a legitimate first run — defaults are returned silently, no backup, no flag. An unparseable settings.json (a JSONDecodeError, e.g. a truncated file from a prior crash) is the data-loss hazard: returning defaults silently would let the immediate bootstrap save overwrite the corrupt file, wiping the user's RomM URL, API token, SGDB key, and platform/collection selections with no trace. Instead the adapter logs the corruption loudly at error level, renames the unparseable file aside to settings.json.corrupt-<ts> (the <ts> is the injected Clock's epoch seconds — filesystem-safe), restricts the backup to 0600 like the live file — it still holds the token and the SGDB key, and a rename keeps whatever mode the corrupt file had — and sets a transient in-memory corrupt_reset flag before returning defaults. If the backup rename itself fails (e.g. permissions), the error is logged and defaults are still returned so boot never crashes; a failed chmod is logged as well and leaves the backup and the flag standing. Bootstrap reads that transient flag after migration and — before the immediate save — folds it into the settings dict as a persistent _settings_reset_notice marker ({"backed_up_to": <basename>}), so it survives a plugin reload. The frontend reads it via the non-consuming get_settings_reset_notice callable and surfaces a persistent notice — a QAM PanelSection banner (with a Dismiss button) plus a game-detail WarningCard (informational; its copy points the user to the QAM to dismiss) — not a toast — telling the user their settings were reset and where the backup landed so they can re-enter the server URL and sign in. The marker is cleared only by an explicit user acknowledgement: the QAM Dismiss button calls dismiss_settings_reset_notice, which pops _settings_reset_notice and persists; the frontend clears the shared store on success so the banner and every game-detail card disappear at once. Sign-in does not clear the notice — the user decides when they have read it.
  • Version never down-stamps: on write, the version field is stamped to max(stored_version, _SETTINGS_VERSION). A file written by a newer plugin (stored version > current) is preserved as-is rather than down-stamped, so a later re-upgrade does not re-run migrations against down-stamped data. An absent or older version is stamped up to the current _SETTINGS_VERSION.

GavelNativeAdapter notes — the compiled save-sync core

Both save-sync decisions run through the compiled romm-gavel core rather than in-tree Python:

  • the full per-(rom, filename, slot) sync action (Skip / Upload / Download / Conflict — the matrix in Save File Sync Architecture), and
  • the upload-409 resolution (the decision, on a 409 from RomM's add_save, to either download the server head or surface a conflict).

GavelNativeAdapter loads backend/native/libgavel-x86_64-linux.so via ctypes at construction and binds both symbols. It is itself the ResolveUploadConflictFn seam, and its compute_sync_action method is the ComputeSyncActionFn seam (both in services/protocols/infra.py); both are injected through SaveServiceConfig → SyncEngineConfig → MatrixExecutor, which calls the decision once per file in iter_matrix_outcomes and the 409 backstop in _handle_upload_409 (services/saves/sync_engine/matrix.py).

  • What crosses the boundary: the adapter marshals the caller's raw dict shapes onto the core's caller-owned structs. Two conversions are the adapter's, not the core's, and both exist so an "unknown" stays a flag instead of becoming a value: a server save's ISO-8601 updated_at becomes epoch seconds plus a has_updated_at flag (an unparseable timestamp arrives as a cleared flag, never as a substitute instant, so it cannot win head selection), and an absent size becomes a cleared has_size flag (0 is exactly what the corrupt-local guard reacts to, so it cannot double as "no size recorded"). Presence of the local file rides on the pointer alone — a file that exists but could not be measured is a real case, distinct from a missing one. The core answers with the chosen server_save_id, which the adapter resolves back to the caller's own save dict so Download / Conflict carry the full record their consumers read.
  • What ships: backend/native/libgavel-x86_64-linux.so (romm-gavel v1.0.1, a freestanding build with zero library dependencies — it loads on any x86_64 Linux), vendored verbatim from the upstream release with a pinned SHA-256 checksum. The C ABI has been part of upstream's promise since v1.0.0: struct layouts, signatures and enumerator values now cost a major bump to change, which is what makes pinning a compiled artifact meaningful. Provenance and the update procedure live in native/README.md. The checksum is re-verified by CI, so a swapped binary fails the pipeline; nothing asserts the .so reaches a built package, because no build here produces one.
  • No fallback: if the library cannot load, GavelNativeLoadError propagates so bootstrap() aborts and the plugin stays inert — the same "fatal until the environment is fixed" posture as the SQLite migration gate. There is no Python implementation of either decision to fall back to: domain/sync_action.py holds only the SyncAction vocabulary the core answers in. What holds the shipped binary to the contract is the vendored gavel vectors — the ladder family in tests/adapters/test_gavel_native.py, the decision table in tests/adapters/test_gavel_native_table_vectors.py — alongside the hand-enumerated cases in tests/adapters/test_gavel_native_decision_table.py, the property tier in tests/adapters/test_gavel_native_property.py, and the boundary cases that pin the two marshalling conversions above.

FirmwareService notes — the live firmware resolver

Which firmware files an emulator wants is read live off the machine through two seams — FirmwarePlatformResolver for one platform, FirmwareResolver for the whole machine — both implemented by adapters/atlas_firmware.py over the vendored emu-atlas copy in backend/_vendor/atlas/ (provenance in _vendor/README.md). It replaced defaults/bios_registry.json, a frozen snapshot that no longer exists upstream and could never be refreshed again.

  • The seam exists because of a contract, not a preference. domain/ may not import _vendor (the domain-stdlib-only import-linter contract), so the answer has to arrive through an adapter and domain/ keeps only the vocabulary — domain/firmware_wants.py, the same split domain/sync_action.py makes for the gavel core.
  • Never cached. A firmware answer is about files the user is actively downloading and deleting, so it is read per query.
  • Failure is "unknown", never "nothing needed". The resolver raises on its own invariant violations and promises nothing about not raising, so the adapter wraps every call and answers an unresolved catalogue. Downstream that classifies every file unknown — a BIOS warning is never cleared on ignorance (#1693).
  • Caveats are the only degradation channel — the resolver never logs. The adapter carries their stable code values (never the human message, which may change freely) and traces them through the injected DebugLogger.

The question is asked per platform, and the scope is what makes the two seams different. firmware_for_system(<system>, verify=True) answers for every emulator ES-DE offers for one system — libretro and standalone alike — and it is what every platform-scoped answer reads. firmware_inventory() sweeps the machine instead — the installed .so files, one entry each, plus (since the resolver's 0.19.0) the standalone emulators a packaged card covers — and it is asked unverified, so a card that identifies its image by content names no file there at all. It is the right question only where the caller has no platform to name, which is two places: the RetroDECK-home migration's untracked-BIOS sweep, and download_firmware(firmware_id). Both ask it for placement — where a file goes — never for readiness.

The per-platform reading is verified, and that is a contract rather than a tuning choice. Two answers cannot be had without reading bytes. A packaged rule card may identify its image by CONTENT, so unverified it names no file while still answering declaration="packaged" — DuckStation comes back with an empty requirement list and no system recording, which a reader that counts the list renders as a green "nothing required" over a console that does not boot without an image. And a folder declaration is satisfied by a file INSIDE the folder, which no stat can settle. Measured on the reference machine: 64–318 ms per system verified, against 248 ms for one unverified whole-machine sweep, because the per-system read sweeps that system's own scope rather than the whole BIOS root — it performs no unclaimed sweep at all. The overview therefore pays one reading per platform it renders; that is the cost of answering for a platform whose emulators are all standalone.

One emulator is one identity. The resolver states it under emulator, in the spelling the launch command uses — a libretro entry's core file basename (dolphin_libretro.so), a standalone entry's own command name (DUCKSTATION, PCSX2) — and both a catalogue entry and a firmware answer carry it under that one name, which is what lets a pick made in the emulator picker be matched against the firmware answer it should be judged by. adapters/atlas_identity.py is the one place that reads it. Three neighbouring fields look like they would answer and must not be used for it: label is presentation (ES-DE lists one pcsx2_libretro.so as both LRPS2 and PCSX2), core_so is None for every standalone emulator, and a caveat's token is the resolver's own packaged-card vocabulary, which upstream states is a different vocabulary again. An identity of None is a real state — the resolver could not identify the emulator behind a row — and nothing may be scoped to it.

A catalogue lists rows, and two rows can be one emulator. ES-DE declares pcsx2_libretro.so twice for ps2, and EmuDeck declares Cemu natively and under Proton; 15 of ES-DE's 172 systems carry at least one repeated identity. So a file's declaring set is deduplicated by identity, the first row winning — which is safe because the rows differ in the command that launches the emulator, never in the declaration they read.

Matching is per platform; completeness is per LAUNCH. They are different questions and collapsing them forces a false choice. A file is needed / optional when any emulator in the platform's answer declares it. Whether an absence may be read as "nothing wants it" is a property of the reading and is scoped to the one emulator the game will launch with: read → not_needed; unread, unidentified, or the resolver unable to answer → unknown. Collapsing not_needed and unknown is the defect the four-valued answer removes — a file nobody wants and a file nothing could be asked about are different answers, and the old boolean called both "not required".

Scoping the doubt to the launching emulator alone is deliberate. An emulator ES-DE also offers, that nothing could be read for, says nothing about a launch that does not use it — under the old platform-wide scope one unreadable core took the whole platform's answer grey whichever emulator the game ran. What it costs is that switching emulators can move a platform from a finished answer to a withheld one, which is the truth about the new launch rather than a regression in the old one.

A declaration that established nothing is unread, whatever its state says. read with an empty requirement list is the one pairing that means "this emulator needs no firmware". packaged with an empty list means a card exists and this query established nothing — the same silence a missing .info leaves — so it counts as unread; PCSX2 answers exactly that when its own ini names no BIOS file. unsupported is an emulator the resolver has no source for, and a refused declaration counts as unread too: the emulator does want something the resolver would not follow to a destination.

One platform, one emulator. Which emulator a platform's BIOS answers are ABOUT is resolved once, by domain/emulator_commands.py's resolve_platform_option — the per-platform override (settings.json platform_cores) when its label still names a bakeable emulator, else the es_systems default — and every platform-scoped answer takes a projection of that one pick: the pane's active_core_label is its .label, the BIOS filter's key is its .emulator, and download_required_firmware asks the same function so the button fetches the set the pane called required. It is the read-path precedence ActiveCoreResolver applies minus the per-game layer, which is what lets check_platform_bios answer launching_emulator=None from it — that is the platform-level callers' signal — and a second resolution behind it would put the two surfaces on two emulators. Which is what a device pass found: with the PlayStation's emulator set to PCSX ReARMed the pane displayed that name and judged the platform by the system default beside it, so one platform read not_demanded / ok on the game page and absent / missing on the pane. The per-game path passes the ROM's own resolved identity (ActiveCoreReader.active_emulator_for_rom(...).emulator), which names a standalone emulator as readily as a libretro core; active_core_for_rom cannot be used there, because it answers None for a standalone pick and would send the page back to the platform's.

Standalone emulators are in the scope, and the answer is only as good as upstream's cards. They are identified and scoped like any other emulator now. What they want is another matter: _vendor/atlas/data/standalone_firmware.json carries cards for five (CEMU, DUCKSTATION, MELONDS, PCSX2, XEMU) and those answer declaration="packaged"; every standalone emulator without a card answers declaration="unsupported", which upstream documents as meaning unknown rather than "needs nothing", and so reads as an unread emulator here. Measured against the RetroDECK release here, es_systems.xml offers 20 distinct standalone emulators, leaving 15 uncovered; the count moves with each RetroDECK release, the shape does not. A platform launching one of those 15 reads unknown — the honest answer, and the one ADR-0020's deferral was reaching for by a cruder route.

A platform's file list is the union of two sources. The RomM library holds files nothing wants; an emulator can want a file the library has never had. That third kind is shown like any other and marked on_server: False — the one field every consumer reads, the row's id: None being an honest absence nothing consults. It counts towards readiness (it is genuinely required and genuinely absent) and never towards a download button. It stays out of server_count / local_count too: that ratio is a progress bar over a set the user can complete, and a stock RetroDECK's SNES emulators declare 26 optional files, so folding them in would report 0 / 26 files, 26 missing for a system no core requires anything from. known_count / unknown_count are scoped to the server's rows for the same reason — they are weighed against server_count, and a beyond-server row always classifies needed/optional, so counting one would cancel the unknown verdict for a platform whose every server file went unanswered. Readiness therefore needs no server at all: ES-DE names the active emulator, the resolver says what it wants, and the resolver or the filesystem says what is there (the split is the next paragraph but one). An unreachable RomM costs the files only it knows about and the ability to fetch anything — not the answer.

Where a file goes comes from what the emulator spelled, not from where the path lands. A requirement carries both, and they answer different questions. _declared_location reads declared — the name the core will open — because the resolved path has already followed every symlink: RetroDECK points <bios>/pcsx2/bios at <bios>, so reconstructing LRPS2's location from the resolved path yields . and loses the folder the core opens. path is still read, for the question it does answer — whether the destination is inside the root the plugin owns — and None comes back for a destination outside it (a standalone emulator's own XDG tree), for a declaration that is absent or absolute, and for one normalising to . or climbing out. The consumer then falls back to its own flat layout. Joining a declared location under the BIOS root needs safe_join(..., allow_base=True): the default "strictly below the base" rule reads that legitimate link as an escape and would drop a real requirement. A server-supplied name never opts in — there the base is not a destination, and only "", . or a link back onto the root could reach it.

Presence follows the same boundary. FirmwareDemand.is_downloaded takes the resolver's present for a row it declared and placed under this root, because the resolver read that destination the way the emulator will reach it. Everything else is this service's own exists probe: a library file nothing declares, a placement whose location is outside the root (its reading is about somewhere else), and the download batch's re-check, which cannot re-read the whole machine per file. Two derivations of one fact is what the LRPS2 row cost — with the destination wrong, the resolver had the file and the service did not. present is three-valued and a None reads as absent: not a claim that anything is there, and the safe direction, since the row then shows work outstanding rather than a readiness nobody established. What the reading found travels per row beside it — supplied_by (the distribution whose own copy sits at the destination, named as the resolver's own display form for that distribution, never one mapped here) and caveats (the resolver's stable codes for whatever else it found at or in that destination, attributed as the paragraph on folder words below sets out, and collapsed on the code within one destination, because the row carries codes where the answer carries statements and requirements resolving to one place share them) — so both surfaces can say what a row IS instead of describing every one of them as a gap in the library. All of it goes silent with the location, for the same reason. declared_kind does not: it is what the emulator OPENS the destination at, a property of the declaration rather than of the destination, so it survives an empty one — a folder that is not there is still a folder to create, and the platform detail's download filter, _download_firmware_batch and download_platform_firmware_file all key off it so such a row is never offered as a fetch. The per-file entry point refuses with a declares_directory reason rather than passing the row over: it answers one file the user named, so a silent success would leave the row unchanged with nothing to explain it.

A folder requirement is answered by what is inside it. LRPS2 declares pcsx2/bios — a folder, required, and always present because RetroDECK links it onto the BIOS root. Reading presence as the verdict said "All 2 files LRPS2 requires are in place" over a PS2 install with no BIOS file at all; reading absence said red over a folder that is plainly there. So BiosFileEntry.satisfied is the row's verdict, and the required counts key off it and nothing else: count_required counts rows answered True, count_required_withheld counts rows answered None, and a row answered False is a requirement shown to be unmet — it reads red, and the play-row badge rises for it, because the game will not launch. All three are scoped to required_by_active first, so an unmet row the launching core does not require moves none of them; the library's own server_count / local_count ratio is a different axis again and keys off on_server / downloaded. Whose answer the verdict is depends on the declaration: a declared file is FirmwareDemand.is_downloaded, a declared folder is the resolver's own listing of that folder, and a declared file with a directory obstructing its destination is withheld (the resolver's firmware-path-obstructed, carried on the row).

A folder verdict costs a content read, and the per-platform reading pays for it once. There is no second, narrower question any more: firmware_for_system(<system>, verify=True) settles the folder rows in the same answer the rest of the platform comes from, so a row and the verdict over it can never come from two readings. Which rows a stat alone could have settled is the resolver's own three-valued answer and not a list of shapes held here — today an absent folder, a plain file sitting where the folder belongs, and a folder holding no file of a size the emulator would open are all settled without reading a byte, and only a folder with candidates in it is a question about bytes. The seam reads the disk, so _platform_firmware_resolver — the attribute every consumer binds it to — is registered in scripts/check_uow_seam_nesting.py's IO_SEAM_METHODS and must never be called with a Unit of Work open; its whole-machine sibling _firmware_resolver is listed beside it for the same reason.

A row hears what was found at its destination; only a folder row hears what a listing found inside it. _caveats_by_destination builds two indexes — one keyed by the path a caveat is about, one by the dir a listing was made in plus the parent of every path, which is how a verified read's per-image findings reach the folder row that asked for them. A file declaration never draws on the second, which matters because on a linked root the listed folder is the firmware root and so is the resolved destination of anything that collapses onto it.

A dir on a caveat is not always a folder declaration's own, and that is new with the per-system route. The resolver states one for a standalone emulator's SEARCH directory too — DuckStation ranks its images in the BIOS root, which is also where LRPS2's pcsx2/bios resolves — so keyed by place alone a folder row would word another emulator's search as its own verdict's cause. _speaks_for drops a statement attributed to a different emulator, reading the resolver's own attribution keys (core_so, token, core); a caveat naming none of them is a statement about the place with no owner — a listing that failed is the case that matters — and it stays on the row, which is the permissive direction: the row keeps a cause it might not own rather than losing one it does.

A required row nothing could judge declines the verdict instead of guessing it. _requirement_verdict_withheld takes both compute_bios_level and compute_bios_label to unknown while every file row keeps its own answer. The count travels as required_withheld because three surfaces need it: the platform detail tells its two unknowns apart with it, the shared utils/biosSummary.ts picks between "Nothing could be established about what X needs" and "One file X requires could not be checked", and the play-row badge subtracts it so a required file whose absence was established still warns. What a withheld row SAYS comes from its caveat codes and, for a declared FILE, from its checked — the resolver's own word for what became of the bytes, carried from the same entry declaration is, because a file the emulator read and does not recognise WAS checked while one whose bytes never came back was not, and one withheld verdict cannot say which. The verdict is the answer alone and carries none of its causes; the verdict decides only which family of codes can apply and what to say when none of them is recognised, which is the one sentence written off it — frontend/src/utils/biosFileNote.ts is the one place both surfaces derive what a row says from — a sentence, the lines under it, which is how a satisfied folder's images arrive as a list rather than folded into the row's own name, and the description on its own line under the row (biosFileDescription). A row is headed by the file it declares on both surfaces and never by its description: that is the packager's prose out of a core's .info, outside the resolver's contract, and it routinely spells the row's own name into its words — so the shared rule takes the name back out and answers null where nothing is left.

No unknown withdraws a download, because the two questions are independent. What the resolver could establish is the emulator's DEMAND; what is fetchable is what the RomM library HOLDS, and neither answers the other. A platform nothing could be read for still has a library behind it, and fetching from it is the one action that moves the platform along at all — so every download PlatformDetail offers is built off the fetchable set (on_server && !downloaded && declared_kind !== "directory", plus the offline server) and reads the verdict nowhere: not bios_level, not required_withheld, not system_image. The two further inputs those buttons do read are of the same two kinds — required_by_active is the launching emulator's own declaration and is what Download required counts, and the library's own finished ratio is what stops Download all offering a set already complete — so neither is a readiness gate either. Reading readiness there is what took the buttons off PS2, GameCube and PSP the moment a BIOS answer was scoped to the emulator that actually launches: those launch standalone emulators the resolver holds no card for, so the verdict declines over a library that still holds their files.

What the narrowest decline does decide is wording. Where nothing could be established the summary cannot name which files to place, so the pane adds the one route it can still state — a file put in the BIOS folder by hand works regardless — and nothingEstablished has that single consumer. A declined readiness verdict is not that state: its rows were answered, so the pane has a file list to point at instead. required_withheld is what separates the two on the wire, and system_image: "unsettled" (below) joins it on the side that keeps the file list. Neither condition is ever per file — a platform whose reading is complete for its launching emulator (reading_complete_for) may hold plenty of not_needed files, with no placement in the platform's catalogue, and trips neither, because "no emulator was found to ask for this" is an answer. Those files stay fetchable like every other, since fetchability never reads the answer.

The console's own firmware demand is a third axis — beside the launching core's required-file counts and the library's own held/offered ratio — and it is a value rather than a count. A libretro .info marks each file required or optional and can say nothing else: there is no way to say "one of these", and no way to say the console does not start without one. An author who knows a PlayStation needs a BIOS image has two lossy moves and the deployed catalogue takes both — SwanStation marks all five of its images optional, Beetle PSX marks three of its own required — so the file counts alone report a green "marks none of its BIOS files as required (0/20 RomM library files)" under the one core and three separate prerequisites under the other, over one PlayStation on which no game starts. The resolver answers the missing half from a packaged, source-cited table about the system (CoreFirmware.system_firmware), carried through the adapter per emulator as FirmwareCatalogue.emulator_verdicts and turned into domain/bios_status.py's classify_system_image. The narrower half of that per-emulator answer is read over every one of them at once (emulators_needing_one_of_their_files) and stamped on each row's own cores entry beside that core's required flag, as the NUMBER of files the core declares. The two keys are two speakers — the core's .info and the packaged table — so optional with needs_one_of: 5 is the informative pair rather than a contradiction, and neither is ever rewritten into the other.

A core is in that narrower half only where it marks nothing required, which is the one shape in which "one of these" is the whole of what the core says. Beetle PSX's console needs an image too, and what that core says about it is three required rows — already carried by every count and every surface — so annotating its optional rows as well stated one requirement twice, and put "the console will not start without one" under ps1_rom.bin, a file it marks optional while hard-requiring three others. BiosFileEntry.system_image_candidate is the same map read for the ACTIVE core onto the row, which is how the platform table marks the rows that can answer the demand; it is deliberately narrower than the set classify_system_image weighs, and the reason is stated at _active_core_answer, the module-private helper that reads the flag for the launching core.

  • It is not folded into required_count. The console asks for one of the images the core declares, so it is one requirement over the whole list rather than one requirement per file; put into that count it would read 0 of 5 files SwanStation requires are in place under the SwanStation this was observed on, five being what that core declares. The twenty in the page's own 0/20 RomM library files is a different set again — the RomM library's inventory for the platform, which this axis neither counts nor is scoped to. Every surface words it "at least one" and none states it as a ratio, and none of them points at the file list either — only the images the launching core declares can answer the demand, and the rows beside them cannot.
  • Whether an image is held is read off the rows, and requirements_met is not consulted at all. The demand comes from the system table, the presence from the file rows, and nothing weighs one against the other — which is the shape upstream intends for a consumer here. Reading that field as a second opinion would be the misreading it exists to prevent: at the resolver ignorance is always null, so a false is a demonstrated statement rather than a disagreement to resolve. It has exactly two causes — a different required file is absent, or one that is present has the wrong bytes — and each of them leaves one of this platform's own required rows unmet, so the ordinary counts already report it, by name, which this axis never could. The second cause needs a content check to arise at all, and the whole-machine inventory is asked unverified (the entry below), so on this path it cannot occur: on the reference device checked == "mismatch" appears on no requirement of any core, and hash_checked on that answer is false. The combination the removed branch was written for does not arise either — of the three cores there in the cannot-run-without-firmware state, none has a single row the resolver reports present. What a row's satisfied resolves to — presence, and the two shapes where it is null — is written out at classify_system_image, where the alternative is built. It is stated once, there, because an outside reader took the field for the resolver's usability answer and drew a false finding from it.
  • It only ever makes the verdict less green. absent lands on missing, tested ahead of the declines so that a demonstration outranks a platform nothing could be established for. Only one of the three declines can reach that comparison, so the order is a rule about that one and a guard for the rest: absent needs every row the launching core declares answered False, which a withheld required row contradicts by construction — it carries the active core, so it is one of those rows, with satisfied null — and which an empty file list cannot produce at all. What is left is _nothing_established's first shape, where the library's own rows all went unanswered under an incomplete reading while the core's declared images are rows the library does not hold and every one of them is absent. unsettled turns an ok grey and nothing else. held and not_demanded change nothing.
  • system_firmware: null is an unasked question. The table covers the systems somebody has looked at, so an absent entry is a system nobody has looked at. It reaches not_demanded, which makes no claim of its own — and it must never be spent as one, for the same reason unknown is not not_needed one axis over.
  • The answer is scoped to the launching emulator, so one unchanged PlayStation page reads absent under SwanStation, whose five declared images the table says the console cannot start without one of, and not_demanded under PCSX ReARMed, which carries its own HLE BIOS. Both come off the one builder, so a platform and its games cannot disagree about it.
  • Four frontend surfaces read it: the game page's BIOS headline, the platform detail's summary, the platform list's row tooltip, and the play row's red BIOS badge — where "absent" is a second established absence beside the required count, and "unsettled" raises nothing, exactly as a withheld required row does not. Where a withheld required row and an unsettled console demand hold together, the platform detail names the ROW. The two are not gaps over two different file sets: a required_by_active row always carries the launching emulator, so it is always one of the rows the disjunction is read over. It is always one of the unjudged rows that verdict is read over rather than a finding beside it — the decline needs at least one such row, and this is one — and need not be the only one, since another image the core declares can be unjudged too; it is the only half of the pair that can name a file.

No BIOS answer outlives the page that asked for it. get_cached_game_detail carries none and says so (bios_status_unknown), and the live get_bios_status fills it in a moment later; there is deliberately no cached twin of check_platform_bios for it to read. That is why BiosChecker has one method. The live answer's bios_level and bios_label are derived once, in _bios_aggregates, where the reading state that decides them is known; every consumer threads them through rather than re-deriving over the payload.

The BIOS delete's whole input is the download records. PlatformBiosDeleter._delete_recorded_io iterates downloaded_bios rows for the platform's firmware slugs, unlinks each row's own file_path, and prunes the row. Both halves matter. Authority comes from having placed the file, and the row is the only evidence of that, so a status listing cannot gate it — that gate hid the button entirely for a platform whose downloads had all left the RomM library. And the row's file_path is where the download actually wrote (kept current by the home migration's relocate), where a status row's local_path is recomputed from today's placement: for a file fetched before an emu-atlas bump moved it, the recomputed path names whatever now occupies the new destination. A row whose file is already gone is pruned without counting as a deletion or an error, which is also what makes two rows naming one path harmless. get_platform_firmware_status ships deletable_count — records still on disk, distinct paths — so the button's number is the same set the delete acts on; local_count is the library's progress ratio and is wrong for it in both directions.

Three buttons reach that one loop, and they differ only in which records they select: the platform's Delete BIOS (every record), a file row's Delete (delete_bios_file, the records naming it), and a declared folder's Delete (delete_bios_folder, the records written underneath it — a folder has no name a record could carry, and a folder being undownloadable says nothing about the files already in one). A predicate can only narrow that platform's own record set, so no caller can reach a file the plugin did not place; a second copy of the loop is what the register's BIOS-delete rule warns against, because the copies would drift in silence. The same answer is stamped per ROW as deletable_count, which is what each row's button is offered on — never downloaded, which is os.path.exists and equally true of firmware RetroDECK ships.

Domain (backend/domain/)

Domain modules contain pure logic with no I/O and no Decky imports. They take inputs and return outputs; anything stateless and I/O-free that would otherwise sit in a service lives here. Domain is stdlib + self only — it imports no other internal layer (lib and models included). Aggregate roots and the enforcement that keeps them honest are documented in Database Design. Selected modules:

Module Role
sync_action.py The SyncAction union (Skip / Upload / Download / Conflict) the whole vertical dispatches on and the compiled core answers in. The decisions themselves are made in the core, not here. See Save File Sync Architecture.
sync_diff.py ROM classification and platform/collection diff computation for the sync preview
cover_refresh.py cover-fingerprint compare kernel (#1386) — scan_cover_refresh_candidates / count_cover_refreshes, shared by the apply-path invalidation pass and the preview's cover-work count so the two can never diverge
preview_delta.py PreviewDelta shape for the sync preview (the staged answer included) plus is_expired / preview_expires_at, the one expression of its TTL
work_unit.py WorkUnit — the per-unit sync work item
rom_save_sync_state.py RomSaveSyncState aggregate + FileSyncState value object — per-ROM save-sync state, backed by rom_save_sync_states + rom_save_files
save_path.py / save_attribution.py / save_status*.py save path resolution, uploader attribution, status DTO building
save_answer.py SaveAnswer / SaveComponent / SaveGroup and the rule that classifies one resolver reading into exactly one of the five save states. Holds the progress-vs-configuration role rule, the download-name fold, and the benign-skip reason the four refusing states produce (ADR-0034)
firmware_paths.py / bios_status.py / firmware_wants.py BIOS path computation, the status shape, and the resolver's vocabulary. bios_status.py holds the status dataclasses (BiosFileEntry, BiosStatus) and owns the unknown/ok/partial/missing LEVEL (compute_bios_level / compute_bios_label) — the single source of truth all surfaces read; phrasing + color stay UI-layer. firmware_wants.py holds only the words a firmware answer arrives in (FirmwareCatalogue / FirmwarePlacement, the four wanted values) — the decisions are the resolver's, not this layer's
iso_time.py parse_iso / parse_iso_to_epoch — ISO-8601 timestamp parsing (stdlib only)
achievements.py achievement progress computation
shortcut_data.py shortcut data building (registry entries, shortcut dicts)
game_instance.py GameInstance (one live sandbox instance: its signal-target pids + argv) and match_instance_for_launch_path — which live instance is running a given ROM, so Stop Game signals that tree and no other
steam_categories.py Steam collection name computation
sgdb_artwork.py SGDB asset-type/endpoint maps and to_signed_app_id
installed_roms.py / rom_files.py installed-ROM detection, M3U generation, launch-file detection
state_migrations.py migrate_settings (settings.json) + fold_legacy_save_sync_settings (one-time legacy save_sync_state.json fold)
sync_state.py SyncState enum (idle, running, cancelling)
emulator_tag.py / version.py emulator-tag formatting, version parsing, core-change detection
identity.py DISPLAY_NAME — the plugin's name where a person reads it: the toast sender, the RomM token label, the registered-device client, and the headline of each of the two files opened by hand (a recovery bundle's README and the recovery root's). The identifier romm-tender is deliberately not here and not in one place; which name a new string takes is CONTEXT.md's "Display name vs identifier" entry; why its three homes stay apart is this module's docstring

Config-source parsers follow a dedicated domain+adapter template (pure parse in domain, I/O in adapter, callback Protocol into services). The full pattern, source catalog, and decisions log are on the Config Source Parsers page.

Models (backend/models/)

TypedDicts and dataclasses describing on-disk and in-flight data shapes (state.py, sync.py). Models import nothing from the other layers.

Other

File Role
main.py Plugin class — lifecycle (_main/_unload) and the endpoints (one method marked @route per endpoint)
bootstrap/ Composition root — adapters.bootstrap() builds adapters, services.wire_services() builds services
lib/errors.py Exception hierarchy (RommApiError, classify_error)
lib/list_result.py ErrorCode and the canonical callable failure shape

Where user data lives

Where this program's directories are is resolved once from the environment and handed in — it is never derived here (ADR-0036). domain/app_directories.py is the ladder, pure, with the environment passed as an argument so every rung is checkable against a table:

  1. TENDER_CONFIG_DIR, TENDER_DATA_DIR, TENDER_CACHE_DIR, TENDER_STATE_DIR, TENDER_CODE_DIR — what the installer resolved once and wrote into the service unit. This is the rung a service runs on. The installer and the service do not share an environment: the installer runs in a login shell, the service in the user manager's, and an XDG_DATA_HOME set in a shell profile never reaches the latter.
  2. XDG_CONFIG_HOME, XDG_DATA_HOME, XDG_CACHE_HOME, XDG_STATE_HOME, XDG_RUNTIME_DIR.
  3. The built-in defaults below.

Rungs 2 and 3 exist for a start by hand. A variable set to the empty string counts as unset — an empty path would resolve to whatever directory the process happened to be started in.

Root Holds
~/.config/romm-tender/ settings.json, plus its .tmp, .lock and .corrupt-<ts> siblings
~/.local/share/romm-tender/ romm_sync.db, the launcher, the single-instance lock; the installer's update-backup/ and rollback-backup/
~/.cache/romm-tender/ covers/, artwork/ — everything re-derivable from the server
~/.local/state/romm-tender/ the log file; the installer's update-failure.json
$XDG_RUNTIME_DIR/romm-tender/ the port file

The data/cache split is not filing tidiness. What lies under the cache root is re-derivable — covers and artwork are fetched again if they are gone — and what lies under the data root is not: the database is the only copy of what the user has installed, synced and chosen. A system that clears caches must be able to clear one and not the other. XDG names no default for the runtime directory, so a start that has none falls back to the state directory; the port file is a hint and connecting to it is the proof, so a stale one misleads nobody who checks. The entries marked as the installer's are written by install.sh when it updates and rolls back, never by the backend; what they hold is in Running an installed one.

One place reads them. WiringConfig.directories is the only path any service reaches through, and RuntimeBundle carries no directory at all. It used to carry two — the plugin folder and the loader's runtime directory — which is how a question about a plugin loader's own layout came to sit beside a question about the user's data.

There is no start-up migration. Earlier releases copied a user's data forward from a directory a plugin loader had named, because that name followed the packaging and moved every user's data at 0.31.0. Nothing derives a directory from a folder name any more, so the question has no asker: the ladder above is the whole answer, and 1.0 is a breaking change that is installed rather than migrated into.

The launcher's home

<bin root>/tender-rom-launcher — ~/.local/bin/tender-rom-launcher by default — is the file every Steam shortcut's exe names, and bootstrap() puts this release's copy there on every start (ADR-0038). adapters/launcher_install.py owns the write; domain/user_data_location.py owns where it goes, through two composers over one tuple — launcher_in_bin_dir for the installed copy and launcher_path for the one the release ships beside the program — so the components that make up /bin/tender-rom-launcher have one spelling, and the suffix ownership detection matches is derived from that same tuple.

The bin root is the one directory on AppDirectories not named after this program: it is shared with every other program the user installed for themselves, which is why the install creates it at the umask's mode rather than owner-only, and why an uninstaller leaves it alone. The reason the launcher is there rather than beside the code is that an update replaces the code directory, and a shortcut's exe must not name a file inside something that gets replaced. The reason it is not under the data root is that the data root holds the only copy of the user's library and nothing executable. The reason it is written on every start rather than once is that a launcher installed once would freeze at whatever version the day of the move brought. The reason it is written through a staging file that is renamed on — never in place — is that a game running right now is executing that file, and bash reads a script as it runs it.

It is unconditional. There was once an ordering to respect here — the install had to wait for a start-up migration's data half to land, because writing into an empty data root would have settled that migration's first rung for the life of the install. Nothing migrates now and the bin root is simply where this run was told it is, so the install runs on every start and the only question left is whether the write succeeded. A start whose write failed creates nothing, and ShortcutLauncher.path is then the copy that ships beside the program: a real file, so a sync in that state still produces shortcuts that launch.

ShortcutLauncher carries the two answers apart on purpose. path is what a newly built shortcut names, and follows the INSTALL rather than the intent: the home where this start actually got the launcher into it, the shipped copy otherwise — including the start whose write failed, whose home is empty. The one case where even the shipped copy is not a real file is a package shipped without its launcher, which is the same reason the install failed. at_home is the narrower question — path is the home in the bin root (launcher_in_bin_dir(directories.bin_dir)), with this release's launcher in it — and it is what repointing an EXISTING shortcut turns on: pointing one at a launcher nothing put there stops its game from starting, and no part of this plugin could put it back.

Repointing the shortcuts that already exist

services/shortcut_relocation.py answers which of Steam's non-Steam shortcuts are OURS and are not at the launcher's home. Ownership is one ending, /bin/tender-rom-launcher, and nothing else — a shortcut naming any other launcher, including one an earlier version of this program wrote, is foreign here and is left where it is. It reads them out of shortcuts.vdf through SteamConfigStore.read_shortcut_exes — one parse for all of them, where the frontend's own route to the same fact is a RegisterForAppDetails per shortcut — and hands the frontend a list of app IDs plus the exe and start_dir to write. Two shapes in that file are matched carefully because both fail quietly: the keys case-insensitively (Steam has written more than one case), and the app id converted out of the signed int32 form the file stores, since every SteamClient API takes the unsigned one.

It is a one-time transition with a recorded completion (kv_config, shortcut_launcher_relocated), and the reading is its only writer: a call that finds nothing of ours outside the launcher's home stamps it, and no later start reads the file again. The frontend writes and reports; it records nothing, so a completed rewrite is stamped on the FOLLOWING start — Steam writes its in-memory shortcuts to the file when it chooses, and the file is what the stamp rests on. Every uncertainty answers blocked instead — the launcher is not at its home, or the file could not be read — and a blocked answer is never stamped, so the next start asks again. The gap that leaves is named at get_shortcut_relocation: nothing clears the stamp, so a shortcut of ours that turns up later away from the home stays there while the backend reports done. The card that used to read the stamp is gone with the plugin loader; the stamp itself still decides whether a later start re-reads Steam's file at all.

Composition Root (bootstrap/)

The composition root is a package of two halves, one per phase, plus an __init__.py that is namespace and re-exports only — consumers write from bootstrap import … and never deep-import a submodule:

  1. adapters.py — owns bootstrap(), which is told where the directories are, and where releases are asked for, rather than deriving them, then builds every adapter, applies the SQLite schema migrations, and loads + migrates settings.json (folding in the one-time legacy save_sync_state.json settings) so the settings persister binds the live mutable settings dict at construction. Returns a typed BootstrapResult carrying four bundles (adapters, stores, callbacks, runtime_adapters), a small handles struct for Plugin-only outputs, directories — the AppDirectories this run was handed, seven fields — launcher, where the shortcut launcher lives in the bin root and whether this start got it there, and user_agent, <package name>/<version> from the one read of the manifest that the outgoing User-Agent comes from, which is also the identity the host answers under. The bundle dataclasses are defined here too — they are the vocabulary the second half consumes.

  2. services.py — owns WiringConfig and wire_services(), which takes the four bundles plus min_required_version, directories, launcher and update_source, and constructs every service, injecting each one's *ServiceConfig. Returns a frozen ServicesBundle holding every service and the prune conflict gate as typed fields.

The two-phase split exists because adapter instantiation and state loading happen first (bootstrap()), then main.py composes the runtime bundle (event loop, event funnel) and calls wire_services(). Services receive the settings dict (the only field on StateBundle) plus the SQLite Unit-of-Work factory / repository handles for all relational state — no plural in-memory state dicts remain. Some services are constructed before others to satisfy ordering constraints (e.g. MigrationService before SaveService so save sync can gate on a pending RetroDECK home migration). Forward references between peers are threaded via LateBinding.

bootstrap() also logs, at info rather than behind the debug toggle, which interpreter the vendored resolver would run its core probe under: describe_core_probe_interpreter() (adapters/atlas_saves.py). The resolver derives it from the running program — sys.executable is a real CPython here, the system one the service unit starts, or the venv's under mise run dev — and nothing registers one over it, because a registered path could only be a guess about the machine. Where it derives none it spawns nothing and answers unknown for every core it is asked about; the one question this program puts that reaches the probe is a libretro entry's save answer, which then usually establishes nothing at all. Nothing fails when that happens, which is why the answer is logged: the caveats the loss leaves (core-unqueryable, core-generation-unestablished) do reach the debug log and the wire, but that line is the only place their cause is named.

Per the process-boundary rule, adapter instantiation never happens in main.py, and no service wiring happens in bootstrap/'s caller other than via wire_services(). Both modules are governed by the ~1000-LOC decomposition threshold (scripts/check_module_size.py), neither is grandfathered.

Protocol Interfaces

Services depend on Protocols, never on concrete adapter implementations. The Protocols live in the services/protocols/ package, organised topically (consumers always deep-import from services.protocols import X):

  • transport — external system clients: RommApi (and its narrowed facets RommSaveApi, RommRomReader, RommDeviceApi, RommFirmwareApi, RommPlaytimeApi, RommLibraryApi, RommConnectionApi, RommPlatformReader, RommAchievementsApi, RommSyncApi, RommVersion), SteamConfigStore, SteamGridDbApi.
  • determinism — Clock / UuidGen / Sleeper test seams.
  • persistence — SettingsPersister.
  • paths — RetroDeckPaths, SystemResolver, CoreInfoProvider, SaveLocationReader (a game's save answer and its savestate directory), SandboxLauncherFn, CoreResolverFn, PlatformCoreReader.
  • infra — cross-cutting callable seams: EventEmitter, DebugLogger, PathExistsReader, ResolvedPathFn, HostnameReader, PendingSyncReader, DownloadQueueCleanup.
  • files — filesystem seams: CoverArtFileStore, DownloadFileStore, FirmwareFileStore, MigrationFileStore, RomFileStore, SaveFileStore, SgdbArtworkCache.
  • cross_service — narrowly-typed multi-method seams one service exposes to another so services stay independent: BiosChecker, AchievementsReader, ArtworkManager, ArtworkRemover, RetryStrategy, MigrationPendingFn, DeviceForgetFn, DeviceIdProvider (server device id, SaveService.get_device_id → PlaytimeService), PlaytimeScopeNoticeClearFn (PlaytimeService clears its re-sign-in notice on a fresh sign-in from ConnectionService), the LaunchGate* and Session* seams.
  • repositories — one repository Protocol per aggregate root, plus KvConfigRepository for the kv_config key-value surface.
  • uow — UnitOfWork, which exposes those repositories as typed properties, and UnitOfWorkFactory, the call-shaped seam a service holds to open one.

Protocol names carry a suffix that signals shape (…Reader, …Provider/…Fn, …Store, …Cache, …Persister; bare names for pervasive primitives like Clock).

RommApiAdapter implements RommApi over RommHttpAdapter.

Boundary Enforcement

Four CI-gated layers keep the dependency direction and the call-site rules from drifting. Aggregate-specific enforcement (the @cosmic_aggregate decorator and the field-assignment check) is documented in Database Design.

1. import-linter (CI-enforced)

.importlinter declares the layer contracts:

# Services must not import concrete adapter implementations (Protocols OK)
[importlinter:contract:no-adapter-impl-in-services]
type = forbidden
source_modules = services
forbidden_modules = adapters

# Adapters must not import services
[importlinter:contract:no-services-in-adapters]
type = forbidden
source_modules = adapters
forbidden_modules = services

# Utilities (lib/) must not import services, adapters, or domain
[importlinter:contract:utilities-independence]
type = forbidden
source_modules = lib
forbidden_modules = services, adapters, domain

# Domain is pure compute — no dependency on any other internal layer
[importlinter:contract:domain-independence]
type = forbidden
source_modules = domain
forbidden_modules = services, adapters, lib, models

# Domain is stdlib + self only — no vendored third-party packages
[importlinter:contract:domain-stdlib-only]
type = forbidden
source_modules = domain
forbidden_modules = _vendor

# Models must not import services, adapters, domain, or lib
[importlinter:contract:models-independence]
type = forbidden
source_modules = models
forbidden_modules = services, adapters, domain, lib

# The host is the process, not a layer: it reaches none of the code it hosts
[importlinter:contract:host-hosts-nothing-it-knows]
type = forbidden
source_modules = host
forbidden_modules = services, adapters, bootstrap, domain

# ...and only the entry point may reach the host. main.py is not a package,
# so nothing in this list can name it.
[importlinter:contract:nobody-imports-the-host]
type = forbidden
source_modules = services, adapters, bootstrap, domain, lib, models
forbidden_modules = host

# Services must not import stdlib I/O / non-deterministic primitives directly
[importlinter:contract:no-stdlib-io-in-services]
type = forbidden
source_modules = services
forbidden_modules = random, subprocess, threading, requests, time, uuid

# Services must be independent of each other
[importlinter:contract:service-independence]
type = independence
modules = services.library, services.saves, services.playtime, ...

Run with PYTHONPATH=backend lint-imports (or mise run lint). CI gates on this.

The service-independence modules list is hand-enumerated, so scripts/check_service_independence_contract.py (bundled into mise run lint and gated in CI) derives the expected services from backend/services/ and fails if the contract omits a service or carries a stale entry — keeping the list self-healing rather than silently rotting.

2. Cosmic Python call bans

scripts/check_cosmic_call_bans.sh (also bundled into mise run lint) complements the import-level guardrail at the call site: services may not call datetime.now() / asyncio.sleep() / time.time() / time.monotonic() / uuid.uuid4() / random.* directly — they inject the corresponding Clock / Sleeper / UuidGen Protocol instead.

3. Aggregate field-assignment check

scripts/check_aggregate_field_assignment.py (also bundled into mise run lint) is a small custom AST linter that enforces the mutation-only-via-methods rule for aggregates — a rule no type checker can express directly. It collects the class names decorated with @cosmic_aggregate in domain/ (the aggregate roots), then scans services/ for <aggregate>.<field> = ... assignments and fails CI on any it finds. The escape hatch is a trailing # pragma: no aggregate-check on the offending line. Full detail in Database Design.

4. Failure-shape dialect gate

scripts/check_failure_shape.py --check (also bundled into mise run lint) is a small custom AST linter that enforces the canonical failure shape for dict-returning callables — every success: False return in services/ must carry both reason and message and must not carry the legacy error_code key or a second error key. It collapses the three dialects that previously coexisted (error_code, error, and slug-less ad-hoc dicts) onto one vocabulary. The two documented carve-outs (discriminated-status unions — a status key with no success; and partial-success payloads carrying an additive server_query_failed / recommended_action flag) are pattern-exempt. Run without --check for the report-mode inventory grouped by classification. The routing slugs come from lib.list_result.ErrorCode (the Lean enum) plus bespoke plain-string reasons for non-server-reachability guards.

5. Enforced: underscore prefix

All internal methods use a _ prefix; public callables (exposed to the frontend via callable()) have none. main.py callable methods delegate directly to the corresponding service method. An endpoint is reachable because it is marked @route, not because it is async def; .claude/rules/callables.md owns when one is a def.

This is no longer just a convention — basedpyright enforces it with reportPrivateUsage = "error", so accessing a _-prefixed name from outside its owning class is a hard type error. Tests are exempt via an executionEnvironments override (white-box testing — inspecting and rebinding a system-under-test's private state — is an accepted pattern). One corollary: a method one sub-service calls on a peer is part of that peer's public surface and carries no underscore, which keeps reportPrivateUsage coherent with the saves-style peer-injection carve-out.

Service Dependency Summary

Every service receives its dependencies through a single *ServiceConfig dataclass. Cross-service dependencies are Protocol-typed (services never import each other's concrete classes). Selected wiring:

Service Key injected dependencies
LibraryService RommLibraryApi, SteamConfigStore, ArtworkManager, Clock/UuidGen/Sleeper, SettingsPersister, UnitOfWorkFactory (roms / sync_runs / kv_config / rom_metadata), ConflictRules (the rules its use cases check at their entry, and the leases sync_complete and sync_stale carry)
MetadataService UnitOfWorkFactory (reads rom_metadata / roms)
SaveService RommApi, RetryStrategy, SaveFileStore, UnitOfWorkFactory (rom_save_sync_states / rom_save_files / answered_save_directories), Clock, RetroDeckPaths, SaveLocationReader, SystemResolver, ActiveCoreReader, MigrationPendingFn (the sync engine's own RetroDECK migration backstop), ConflictRules (the rules its use cases check at their entry)
DownloadService RommApi, DownloadFileStore, RetroDeckPaths, Clock/Sleeper, RomInstallRecorder + DownloadTargetGateFn cross-service seams, ConflictRules (the rules a start and a resume check at their entry, the operation a download's task holds, and the lease a bound ROM's download_complete carries)
FirmwareService RommApi, FirmwareFileStore, FirmwarePlatformResolver, FirmwareResolver, CoreInfoProvider, RetroDeckPaths, UnitOfWorkFactory (firmware_cache), ConflictRules (the rule each download and delete checks at its entry)
SteamGridService SteamGridDbApi, RommApi, SteamConfigStore, SgdbArtworkCache, UnitOfWorkFactory (sgdb_id on roms), PendingSyncReader
MigrationService MigrationFileStore, RetroDeckPaths, RelaunchOptionsReader, FirmwareResolver (which files in a pending home are firmware), SaveDirectoriesRecorderProvider (records the answered save directories afresh once a home migration has moved the files)
GameDetailService BiosChecker, AchievementsReader (cross-service), Clock, UnitOfWorkFactory (one read UoW over roms / rom_installs / rom_save_sync_states / rom_metadata / kv_config), plus PathExistsReader + RetroDeckPaths + SystemResolver for the single target-path stat
AchievementsService RommAchievementsApi, Clock, DebugLogger, UnitOfWorkFactory (reads ra_id from roms)
SettingsService SteamConfigStore, SettingsPersister, UnitOfWorkFactory (reads bound shortcut_app_ids from roms)
PlaytimeService RommPlaytimeApi, RetryStrategy, DeviceIdProvider (server device id, satisfied by SaveService.get_device_id), Clock, UnitOfWorkFactory (reads/writes rom_playtime + rom_playtime_sessions, and the kv_config scope-notice flag). Exposes clear_scope_notice to ConnectionService (via PlaytimeScopeNoticeClearFn) so a fresh sign-in drops the re-sign-in notice, ConflictRules (the rules its use cases check at their entry, and the operation a session start's outbox flush holds)
RomAdoptionService RommRomReader, DownloadFileStore (the stat, the top-level listing, the directory scan, the hashing), AdoptionMoveStore (the rename across the roms/saves/states trees), RetroDeckPaths, SystemResolver, SystemM3uSupportFn, SystemSupportedExtensionsFn, SaveLocationReader (the save and the savestate directory of both launch paths of a rename, asked of the emulator ActiveCoreReader names), ActiveCoreReader, RomInstallRecorder peer, EventEmitter + Clock (throttled verify_progress frames), ConflictRules (the rules an adoption checks at its entry, and its lease)
RomInstallRecorder Clock, UnitOfWorkFactory (rom_installs upsert + the roms size / applied-launch-options writes), SystemSupportedExtensionsFn (the launchable verdict), ActiveCoreReader + DiscResolver (the launch bake)
RomRemovalService RomFileStore, RetroDeckPaths, DownloadQueueCleanup peer, UnitOfWorkFactory (reads/deletes rom_installs), Clock + EventEmitter (removal duration logging and uninstall_progress frames), ConflictRules (the rules its use cases check at their entry, and the leases they hand the frontend)
ShortcutRemovalService SteamConfigStore, ArtworkRemover peer, UnitOfWorkFactory (unbinds via roms, offline name via kv_config), ConflictRules (the rules its use cases check at their entry, and the removal lease)
SessionLifecycleService Session* cross-service seams (playtime / post-exit sync / achievement sync / migration reader), ConflictRules (the rule the finalize checks at its entry)
LaunchGateService LaunchGateRomLookup, LaunchGateInstalledChecker, LaunchGateSaveStatusReader cross-service seams, ConflictRules (the rule the evaluation checks at its entry)
GameProcessService GameProcessControl (instance discovery + signals), RomLaunchPathReader (the launch target each live instance is matched against, from the same resolver that bakes it), Sleeper (the grace window — the service owns no clock), and RetroDECK's flatpak app id from the single domain.shortcut_data.RETRODECK_APP_ID constant the launch command is built from
ConnectionService RommConnectionApi, SettingsPersister, min_required_version, DeviceForgetFn (cross-service — forgets the device id on an origin change), PlaytimeScopeNoticeClearFn (cross-service — clears the playtime re-sign-in notice on a fresh sign-in)
VersionSwitchService RommRomReader (live sibling_roms view + per-sibling detail), Clock, UnitOfWorkFactory (resolves a sibling group via iter_by_group_key and moves its binding on roms), settings (reads preferred_region for the default-badge ranking), SaveDriftProbeFn + ReachabilityProbeFn (the switch-away save-stranding soft block), RomRelaunchItemReader (re-bakes a switched-onto install's launch_options), ActiveDownloadRomIdsFn (refuse a switch while a group member is downloading), ConflictRules (the rules a switch checks at its entry, and its lease)
UpdateCheckService LatestReleaseFn, Clock, SettingsPersister, UnitOfWorkFactory (kv_config update_check_last_seen), the running VERSION, and whether this process is the installed program

Most services also receive the settings dict (StateBundle's only field), the runtime infrastructure (event loop, logger, the DebugLogger Protocol), and the UnitOfWorkFactory for relational state through their config. The old in-memory state / metadata_cache / save_sync_state / shortcut_registry dicts are gone — every relational read/write goes through the Unit of Work.