modernuo/dev-docs/threading-model.md
Kamron Batman a7e65aab01
perf(login): run password hashing on a parked worker thread (#2566)
## Why

An Argon2 verify is **~8.9 ms of frozen world per login attempt** — more than half a 16 ms frame. Failed attempts cost exactly the same as successful ones, by design, so a credential-stuffing flood is a full-cost stall per packet without needing valid credentials. `SetPassword` derives a hash too, so `[password`, the admin gump and account creation each pay the same.

## What the measurement says

Off-loading does not delete the cost, it relocates it. Three things stay on the loop:

| Component | Measured |
|---|---:|
| Inline verify (today) | **8.92 ms** |
| Dispatch to the worker | 210 ns |
| Drain the continuation off `LoopContext` | 13 ns |
| Loop's own work slowed by shared-L3 eviction | **0.05 – 5.44 ms** |

Net gain **3.5 – 8.9 ms** of on-loop time per login. Harness in `ModernUO-Benchmarks` (`Benchmarks/Argon2OffLoop/`): it models the loop as a dependent-load pointer chase swept across working-set sizes, which is an upper bound on cache-latency sensitivity, and copies `EventLoopContext` so the hand-off cost is the real one.

Two results shaped the design:

- **The contention tax peaks in the middle of the working-set range**, not at the top — 5.44 ms at 8 MiB (a quarter of this chip's L3), but 0.76 ms at 30 MiB and 0.10 ms at 256 KiB. A tiny hot set has nothing in L3 to lose; a huge one is already DRAM-bound.
- **Per-login tax falls as concurrency rises** (5.44 → 2.56 → 1.60 ms at 1/2/4 hashers) while *total* loop damage rises. Contention is shared, not additive, so a login rush is not the disaster case — a single login is.

## Why exactly one worker

It is load-bearing three times over, which is also why it must not quietly become a pool:

- **Cost bound.** Off-loop loses to inline only if a hash steals ~82% of the loop's throughput. One hasher contending for one core leaves the loop ~50%. **A single background hasher cannot cost the loop more than the inline verify under any scheduling regime**, which is what lets the measurement hold on hardware we cannot inspect — AMD, VPS, oversubscribed VM. Four hashers drop the loop to ~20% and break it.
- **Memory.** Exactly one hashing arena is live at a time whatever the login volume.
- **Ordering.** Writes apply in dispatch order *only* because a single thread drains FIFO. A second worker would need ordering reintroduced; `WritesApplyInDispatchOrder` fails if that happens.

Throughput is ~110 verifies/sec. Only loop time matters, not login latency, so head-of-line blocking during a rush costs nothing.

## Making every protection safe off-thread

The worker was initially Argon2-only. That was the right call for the wrong reason — it was blamed on Argon2's salt RNG, which is a stateless syscall wrapper and was never a problem. The real blockers were elsewhere, and both are fixed at the source:

| Protection | Was | Now |
|---|---|---|
| MD5/SHA1/SHA2 | shared `HashAlgorithm.ComputeHash`, which carries the running digest across `HashCore`/`HashFinal` through process-wide singletons | static `HashData` into a `stackalloc` span — no state, no allocation, identical bytes |
| PBKDF2 | `Utility.RandomMinMax` → shared `System.Random`, thread-unsafe *and* game state | `RandomNumberGenerator.GetInt32`, matching the salt beside it |
| Argon2 | already safe (`Verify` is static + stackalloc) | unchanged, singleton reused |

Literal digests are pinned in a test **before** the change and still pass after it. These are compared as strings against every account database, so any casing or encoding drift would lock out every SHA and MD5 account at once.

With all three safe, the worker no longer knows which algorithm it runs and the dispatch conditions collapse to "is off-loop available".

## Correctness

- **Phrase derivation** moves to `AccountSecurity.DerivePhrase`, so verification (stored algorithm's rule) and rehash (target algorithm's rule) cannot disagree. Deriving with the wrong one is the shape of the lockout fixed in #2562.
- **Liveness** is checked at dequeue *and* at apply — a connection can drop while queued or while the result sits in the loop queue. A job with no connection attached, such as an admin password change, runs regardless.
- **Queue overflow rejects** a login rather than verifying inline; steering work back onto the loop is what a flood wants. A password change instead falls back to hashing inline, because unlike a login it must not be dropped.
- **Shutdown and crash** both just stop the thread, and pending jobs are dropped. No save is initiated once shutdown begins — saving is the operator's choice up front, via the admin gump's save/no-save variants, and `WaitForWriteCompletion` honours one already in flight — so a write applied during teardown would reach no disk. The crash path needs its own subscription because `HandleClosed` skips `InvokeShutdown` when crashed.

## Bounding

`MaxPending` is 4096 — a backstop, not a flood defense. `SentFirstPacket` holds a connection to one pending verify and the engine caps connections at 4096, so the queue is already bounded by construction and this can only trip if that invariant breaks. A cap low enough to blunt an attack would reject real players first; during a mass reconnect they *are* the queue. Flood defense belongs at the connection layer.

The real DoS improvement is elsewhere: today every attempt stalls the world, and after this a flood occupies one core while the loop keeps ticking.

## Gate

Release builds on 4+ cores. Below that there is no spare core to move work to, so off-loading buys nothing by construction; `DEBUG` is excluded because dev boxes and test shards have few logins. Both modes call the same code — the gate only chooses where it runs.

## Engine change

One property, `AccountLoginEventArgs.Deferred`, so a subscriber can say "no verdict yet". `EventSink.AccountLogin` is `Action<...>` with no continuation, and the packet handler replies in the same call. Approved separately since it touches `Projects/Server/`.

## Docs

`dev-docs/threading-model.md` and the threading skill gain a vetted-workers section. The forbidden-patterns table bans `new Thread`, `ConcurrentQueue<T>`, `Interlocked` and `volatile` in `UOContent`, and its exceptions covered only `Projects/Server/` — the existing Advanced Search fan-out already sat outside it. The new section leads with proving the need (measure on-loop time, not wall-clock; gate on core count; record the measurement), keeps game logic on the loop via chunking, and documents the hand-off protocol in both directions.

## Testing

698 UOContent tests, 810 Server tests, Release build clean.

Covered: verify and rehash outcomes, phrase rules for SHA1/SHA2 vs Argon2, stored-format stability for MD5/SHA1/SHA2, jobs with no connection attached, and dispatch ordering through the real queue. The liveness and ordering guards are mutation-verified.
2026-08-09 00:13:34 -07:00

14 KiB

ModernUO Threading Model

This document covers ModernUO's single-threaded game loop architecture, the EventLoopContext synchronization context, memory pooling, and rules for safe concurrent code.

Core Principle: Single-Threaded Game Logic

All game logic in ModernUO runs on a single thread. There are no exceptions for code under Projects/UOContent/.

This means:

  • No locks, mutexes, or synchronization primitives needed
  • No concurrent collections needed
  • No volatile fields needed
  • No race conditions possible in game code
  • await is safe because continuations route through EventLoopContext

Game Loop

The game loop in Projects/Server/Main.cs runs continuously:

public static void RunEventLoop()
{
    while (!Closing)
    {
        _tickCount = GetTimestamp();
        _now = DateTime.UtcNow;

        Mobile.ProcessDeltaQueue();    // Send mobile state changes to clients
        Item.ProcessDeltaQueue();      // Send item state changes to clients
        Timer.Slice(_tickCount);       // Execute due timers
        NetState.Slice();              // Process network I/O
        LoopContext.ExecuteTasks();     // Run async continuations
        Timer.CheckTimerPool();        // Refill timer pool if needed

        // World save handling
        if (_performSnapshot)
        {
            World.Snapshot(_snapshotPath);
            _performSnapshot = false;
        }
    }
}

Each iteration:

  1. Updates timestamp
  2. Sends pending mobile/item updates to clients
  3. Fires due timers
  4. Processes incoming network packets
  5. Runs async continuations (from await)
  6. Checks timer pool health
  7. Handles world save snapshots if requested

EventLoopContext

EventLoopContext implements SynchronizationContext to ensure all await continuations run on the game thread.

Defined in Projects/Server/EventLoopTasks.cs:

public sealed class EventLoopContext : SynchronizationContext
{
    public enum Priority { Normal, High }

    private readonly ConcurrentQueue<Action> _queue;
    private readonly ConcurrentQueue<Action> _priorityQueue;
    private readonly Thread _mainThread;
    private readonly int _maxPerFrame;  // Default: 128

    // Post: queues action for next ExecuteTasks() call
    public void Post(Action d, Priority priority = Priority.Normal);

    // SynchronizationContext.Post: used by await
    public override void Post(SendOrPostCallback d, object state);

    // Send: immediate if on main thread, blocks if on other thread
    public override void Send(SendOrPostCallback d, object state);

    // Called once per game loop tick
    public void ExecuteTasks();
}

How await Works

// Safe in game code:
await Timer.Pause(TimeSpan.FromMilliseconds(100));
// After the pause, execution continues on the game thread

Flow:

  1. await captures EventLoopContext as the current SynchronizationContext
  2. When the awaited task completes, the continuation is posted to _queue
  3. LoopContext.ExecuteTasks() runs the continuation on the main thread
  4. Game state is safely accessible

Task Limits

  • Maximum 128 tasks per frame by default (configurable)
  • High-priority tasks (_priorityQueue) are always processed first
  • Normal tasks are processed up to the per-frame limit

Forbidden Patterns

In Game Code (Projects/UOContent/)

Pattern Problem Alternative
Task.Run(...) Runs on thread pool, races with game state Timer.StartTimer()
new Thread(...) Manual thread, races with game state Timer.StartTimer()
ThreadPool.QueueUserWorkItem(...) Thread pool, same issue Timer.StartTimer()
lock(obj) { ... } Unnecessary overhead, no contention Remove lock
Monitor.Enter(obj) Same as lock Remove
volatile int _field Memory barriers not needed Plain field
ConcurrentDictionary<K,V> Lock-free overhead, unnecessary Dictionary<K,V>
ConcurrentQueue<T> Same Queue<T> or List<T>
ConcurrentBag<T> Same List<T>
Interlocked.Increment(...) Atomic operations unnecessary _field++
Mutex / Semaphore OS-level sync, unnecessary Remove
ReaderWriterLockSlim Lock overhead, unnecessary Remove
Thread.Sleep(ms) Blocks entire game loop await Timer.Pause(ms)

Exceptions: Server Infrastructure

These files in Projects/Server/ MAY use threading because they handle I/O outside the game loop:

  • Main.cs -- Event loop setup, thread configuration
  • World/World.cs -- World save disk I/O (serialization on main thread, writes on background)
  • Network/ -- Network I/O processing
  • Timer/Timer.Pool.cs -- Async pool refill
  • EventLoopTasks.cs -- The synchronization context itself

Exceptions: Vetted Workers in Projects/UOContent/

Take great care here. A background thread is a last resort, not a tool of first choice.

The table above is about game logic, which is never threaded. A dedicated worker that touches no game state is the sanctioned way to move CPU-heavy or I/O work off the loop, and it necessarily uses primitives the table forbids -- new Thread, ConcurrentQueue<T>, Interlocked, AutoResetEvent, volatile. Those are legitimate at the thread boundary, and nowhere else.

First: prove the need

Do not add a worker because something "looks slow". Measure, and measure the right thing:

  • Measure on-loop time, not wall-clock. How long a player waits does not matter; how long the world is frozen does. A change that improves latency but not loop time buys nothing.
  • Off-loading does not create CPU. It converts "the loop is blocked for N ms" into "the loop competes for cores for N ms". On a 1--2 core host there is no spare core and it buys nothing at all -- gate on Environment.ProcessorCount.
  • Account for what stays behind. Dispatch, the continuation, and the loop's own work slowing down while the worker evicts shared L3. That last one is real and is usually the largest.
  • Write the benchmark down. A worker with no recorded measurement cannot be re-justified later, and will be removed by someone who cannot tell whether it earns its complexity.

Game logic stays on the loop -- chunk it instead

Work that needs game state cannot be threaded at any core count. If it is too slow for one tick, split it across ticks rather than across threads:

// Bound the work per tick, resume where it left off.
Timer.DelayCall(TimeSpan.Zero, TimeSpan.FromMilliseconds(50), () =>
{
    var budget = 0;
    while (_cursor < _items.Count && budget++ < 100)
    {
        Process(_items[_cursor++]);
    }
});

Bound by count or elapsed time, never by "until done". Threading game state is not a faster version of this -- it is a correctness bug.

Vetted workers

Worker Off-loop work Justification
Accounting/Security/PasswordWorker.cs Password verification and hashing docs/handoffs/2026-08-07-off-loop-argon2-hashing.md -- 8.9 ms/login on-loop at Argon2, measured 3.5--8.9 ms saved
Engines/Advanced Search/AdvancedSearchGump.cs Parallel entity search Admin-triggered full-world scan; saves disabled for its duration

Adding to this table needs the same bar: a measurement, and all five rules below.

The six rules

  1. No game state off-thread, read or written. Hand the worker immutable values (strings, structs) captured on the loop. Carrying a reference is fine only if the worker just passes it back untouched.
  2. Decide policy on the loop, compute on the worker. Anything rule-dependent -- which algorithm, which salt, which era branch -- is resolved at dispatch, so the worker holds no policy it could apply inconsistently.
  3. Park on a kernel wait; never spin. AutoResetEvent.WaitOne() costs nothing while idle. SerializationThreadWorker does spin, but only to await a producer mid-drain; absent that race, spinning is a bug that burns a core on shared hosts.
  4. Yield to world saves. Run only while WorldState is Running or WritingSave. World.Saving is not the right check -- it covers only the freeze and misses PendingSave, where the serialization threads are already awake and spinning on an empty queue.
  5. Bound the queue, or rely on a bound upstream and say which one in a comment.
  6. Everything the worker calls must itself be safe off-thread. A process-wide singleton is not automatically safe -- look for instance state. HashAlgorithm.ComputeHash carries the running digest across HashCore/HashFinal, so two threads sharing one corrupt each other. Utility's RNG is a shared System.Random, which is both thread-unsafe and game state. Prefer the static one-shot forms (SHA256.HashData, RandomNumberGenerator.Fill), and if a dependency cannot be made safe, fix it at the source rather than narrowing the worker around it.

Handing work across the boundary

Loop → worker (dispatch). Snapshot everything needed into immutable values. Capture any value you intend to overwrite later, so the continuation can tell whether it changed:

var job = new Job
{
    Target = state,              // carried, never dereferenced off-thread
    Expected = account.Password, // captured so the continuation can detect a change
    Input = DerivePhrase(...)    // policy resolved here, on the loop
};

if (!Worker.TryEnqueue(job))
{
    // Full. Reject -- do not fall back to running it inline, or a flood steers the work
    // straight back onto the loop.
}

Worker → loop (hand back). Two sanctioned routes, and no others:

// 1. Marshal the apply step. Preferred when a specific result belongs to a specific caller.
Core.LoopContext.Post(() => Apply(job, result));

// 2. Publish an immutable snapshot behind a single volatile reference, read lock-free by the loop.
//    Preferred for a shared lookup table rebuilt periodically.
Volatile.Write(ref _snapshot, newTable);

The continuation must re-validate. Time passed, and the loop kept running:

private static void Apply(Job job, Result result)
{
    // Gone? Never revive a dead NetState or a deleted entity.
    if (job.Target?.Running != true)
    {
        return;
    }

    // Changed? Do not overwrite a newer value with one derived from an older one.
    if (!string.Equals(account.Password, job.Expected, StringComparison.Ordinal))
    {
        return;
    }

    account.Apply(result);
}

Always post a result, including on failure. A worker that throws and posts nothing leaves whatever awaited it waiting forever. Catch, log, and post a failure verdict.

Never call into game state from the worker, and never await on the loop in a way that lets a continuation resume heavy work there -- ConfigureAwait(false) on every await inside off-loop work.

Memory Pooling

STArrayPool

Single-threaded array pool optimized for game code (no locks):

// Defined in Projects/Server/Buffers/STArrayPool.cs
public class STArrayPool<T> : ArrayPool<T>
{
    public static new STArrayPool<T> Shared { get; }

    public override T[] Rent(int minimumLength);
    public override void Return(T[]? array, bool clearArray = false);
}

Usage:

var buffer = STArrayPool<byte>.Shared.Rent(1024);
try
{
    // Use buffer (may be larger than requested)
}
finally
{
    STArrayPool<byte>.Shared.Return(buffer);
}

Architecture:

  • 27 buckets covering sizes 16 to 1GB+
  • Per-bucket cache (1 array) + stack storage (32 arrays)
  • Trim callbacks on Gen2 GC to reduce memory pressure
  • Formula: bucket index = Log2(size - 1 | 15) - 3

Use STArrayPool<T>.Shared in game code, not ArrayPool<T>.Shared (which uses locks).

PooledRefList

Stack-allocated list using pooled arrays:

// Defined in Projects/Server/Collections/PooledRefList.cs
public ref struct PooledRefList<T>
{
    public static PooledRefList<T> Create(int capacity = 32, bool mt = false);
    public static PooledRefList<T> CreateMT(int capacity = 32);  // Multi-threaded

    public void Add(T item);
    public bool Remove(T item);
    public void Clear();
    public int Count { get; }
    public T this[int index] { get; set; }
    public void Dispose();  // Returns array to pool
}

Usage:

using var list = PooledRefList<Mobile>.Create();
list.Add(mobile);
// list is stack-allocated, zero GC pressure
// Dispose() returns backing array to STArrayPool

Key properties:

  • ref struct -- stack-allocated, cannot escape to heap
  • Uses STArrayPool<T> by default, ArrayPool<T>.Shared with CreateMT()
  • Auto-grows when capacity exceeded
  • Must be disposed (use using pattern)

World Save Threading

World saves involve both threads:

  1. World.Save() -- Called on main thread, queues preserialize to thread pool
  2. Preserialize() -- Thread pool: allocates serialization heaps, wakes workers
  3. Snapshot() -- Main thread: serializes all game state (safe access), blocks game loop briefly
  4. WriteFiles() -- Thread pool: writes serialized data to disk (no game state access)
Main Thread:    Save() → ... → Snapshot() → ... → continue loop
Thread Pool:    Preserialize() → ... → WriteFiles()

The main thread blocks during Snapshot() to ensure consistent state, then the disk I/O happens asynchronously.

Best Practices

  1. Never use concurrency primitives in game code -- they add overhead for no benefit
  2. Use STArrayPool<T>.Shared instead of ArrayPool<T>.Shared
  3. Use PooledRefList<T> instead of new List<T>() in hot paths
  4. Use await Timer.Pause() instead of Thread.Sleep()
  5. Use Timer.StartTimer() instead of Task.Run() for delayed work
  6. Trust single-threaded invariants -- no need to protect shared state

Key File References

File Description
Projects/Server/Main.cs Game loop (RunEventLoop)
Projects/Server/EventLoopTasks.cs EventLoopContext
Projects/Server/Buffers/STArrayPool.cs Single-threaded array pool
Projects/Server/Collections/PooledRefList.cs Pooled ref list
Projects/Server/World/World.cs World save system