Nomad Cairn / Evidence

Cairn crash and duplication evidence

The crash and duplication scenarios Nomad Cairn runs before every release, what each one asserts, and where the evidence stops.

Cairn is built against a stress harness that reproduces failures buyers report against plugins that move items between owners, and asserts, for each one, that Cairn does not fail that way. It runs in full before every release, and every number on this page comes from the run that produced your build.

Read that as what it is. These are reproductions of specific, named failure modes, and passing them means Cairn does not fail in those ways, not that no failure is possible. A machine that dies hard dies in a state nobody designed, and the layer below every plugin has a hole of its own: Minecraft writes a connected player's inventory on quit and at a clean shutdown, not continuously, so items held by somebody standing still when the power goes may never have been on disk to recover. The rig measures that directly every release by SIGKILLing a server with no plugin loaded at all. What Cairn is built for is the part a plugin can control: items already in its custody, an interrupted hand-over resolving to one side rather than both, and anything undecidable being held, reported and reversible instead of quietly written off.

Specifically, the scenarios below reproduce and defend against:

It also states, in what this harness does not prove, exactly where the evidence stops, which is the part most reports of this kind leave out.

The one thing every scenario asserts

Not "the code took the branch I expected". Every scenario ends with a census:

Every item that entered is somewhere, exactly once. It is in a player's inventory, or it is in the custody ledger. Never in both. Never in neither.

The check is per item, by identity, with amounts, so a stack legitimately split between an inventory and a grave still passes, and a stack that quietly became two does not. Alongside it, Cairn's own ledger runs its independent conservation arithmetic over the append-only audit log (captured + received − delivered − sent = currently held) and that has to balance too.

How the crashes are produced, and why it matters

There are two honest ways to test a crash and only one of them is worth much.

The weak way is to hand-assemble the wreckage: write the half-finished state you believe a crash leaves, then assert the recovery handles it. That can only ever confirm what the test author already believed, and the bugs in this category live exactly in the gap between what an author believes and what the code does.

Cairn's harness does it the other way. It runs the real production sequencer forward (the same DeathCapture and GraveRelease classes the plugin ships) and then switches the machine off at a chosen instruction. Queued work is dropped, exactly as it would be if the server process died holding it. Whatever state is left behind is the state the production code produced. The harness then opens the ledger directory again, from disk, on a fresh server object, and lets the plugin's real startup recovery resolve whatever it finds.

The plugin's crash-safety argument leans on one property of Paper: a player's inventory and their persistent data container are written to the same playerdata/<uuid>.dat, so a crash rolls back both or neither. The harness models exactly that: a simulated save flushes both together and a simulated crash discards both together, which is why the two crash mid-delivery scenarios below can come out differently and both be correct.

The scenarios

Item counts below are the actual figures asserted in the run.

Death: the capture path

#What a server owner would seeWhat is assertedResult
1 The server dies the instant a player does, before anything has been written to disk. The drops are already out of the death event, so the items are in neither the world nor the ledger. This is a vanilla-style loss and it is not a duplication: nothing can be looted twice. No half-written grave, no ledger row. Pass: 64 items taken from the event, 0 rows written, 0 grave records
2 The server dies after the items are safely in the ledger but before the grave marker record is written. No partial grave record exists at all. All 60 items are still held, still owned by the player, and still listed by /cairn, because Cairn reads the ledger, not the marker. Pass: 0 grave records, 1 held row, 60 items
3 The server dies after the grave record is written but before the marker entity spawns. On restart the record is found still owing the world a marker, and the sweeper's own planner schedules a re-render rather than an expiry or a deletion. The 27 items never depended on the marker at any point. Pass: marker state PENDING, plan = RERENDER, 27 items held

Scenario 1 is worth reading twice, because it is the whole anti-dupe argument. Cairn removes the drops from the death event before it tells the ledger about them. That ordering means the worst case is a loss; the inverse ordering, write the grave then clear the inventory, is the elytra duplication.

Retrieval: the release path

#What a server owner would seeWhat is assertedResult
4 A player clicks their grave and the server dies before a single item reaches them. On restart the delivery witness reports the items never arrived, the row goes back to held, and the grave is genuinely lootable again on the real path. Pass: witness NO, 40 items held, re-loot delivered 40
5 The duplication window. The server dies after the items are in the player's inventory and their data file has been saved, but before the ledger has recorded the payout. The witness reports the items did arrive. The row closes as delivered. The player has all 64 items, the ledger has none, and three further clicks on that grave are refused. Pass: witness YES, 64 with the player, 0 in the ledger, 0 duplicated
6 The same crash, but the save had not happened yet, so the inventory rolls back. The witness reports the items did not arrive, because the token rolled back with them. The row goes back to held. The player has nothing and the grave still has all 64 items. Pass: witness NO, 64 items held, 0 lost
7 A player clicks their grave with a completely full inventory. Nothing is inserted, the row goes straight back to held under the same custody id, and the grave is untouched. Make room and the retry delivers all 60. Pass: outcome notDelivered, same id, 60 items held, retry delivered 60
8 A player clicks with room for only part of the grave. 40 items go to the player, the remaining 24 stay in the ledger under a fresh held row handed back by the ledger itself. 40 + 24 = 64, with nothing created and nothing lost. Pass: 40 delivered, 24 held, census exact
9 Two players click the same grave at the same instant (two real threads, released together). Exactly one gets a release ticket. The other is refused by the ledger's compare-and-swap, not by luck or a lock. No second licence to touch the items is ever issued. Pass: 1 ticket, 1 refusal, 64 items intact
10 A lagging client double-fires, or a player clicks a grave they already emptied, three more times. Every subsequent attempt is refused as contested. The player's inventory still holds exactly one payout of 40. Pass: 1 delivery, 3 refusals, 40 items

Retrieval that did not fit in one go: the remainder path

These three run against the real on-disk ledger, with its real idempotency store, and they exist because the rest of the retrieval suite does not: its fixture ledger has no idempotency store at all, so several hundred passing tests said nothing about the one thing core enforces on beginRelease: that a key presented twice for two different requests is a conflict, and is refused for ever. A partial delivery re-points the grave at a brand-new custody row whose sequence starts back at zero, which is exactly the shape that can rebuild an identical key against a different row. A fully-geared death takes this path every time: addItem fills only the 36 storage slots, so 36 + 4 armour + off-hand always overflows.

#What a server owner would seeWhat is assertedResult
16 A player loots with room for 30 of the 50 items in the grave, then comes back for the rest. The first click delivers 30 and core mints a remainder row for the other 20, HELD, at sequence zero. The second click, keyed on the custody row, not the grave id, delivers all 20. Keyed on the grave id instead, the second click is an idempotency conflict that every later click repeats for ever: the grave is bricked and the items are stranded with no operator remedy, because /cairn restore refuses a HELD row. Pass: 30 + 20, census exact, ledger holds nothing back
17 The same thing twice in a row: a remainder that is itself too big to fit. Every remainder gets its own key. Three clicks, 20 + 20 + 20, and the census still balances against the 60 minted. Pass: 60 delivered, 0 duplicated
18 A genuine retry of the same click: a lagging client, a re-sent packet. The second beginRelease with the same key is refused with IllegalTransitionException. Core will not hand out a second licence to mutate the world, and GraveRelease maps that refusal to "contested" rather than to "broken". Pass: 1 ticket, second attempt refused

The restart that is quick enough to miss its own recovery

The rest of the crash suite advances its clock past the release lease before every reboot, so the one restart timing that actually happens in production, a watchdog bringing the server back inside a minute, was the one nothing exercised. Under it, startup recovery correctly finds no stale row (the lease has minutes left), leaves the interrupted release alone, and without a catch-up pass never asks again: the items are safe and the grave is dead to the player until some later restart happens to fall past the lease.

#What a server owner would seeWhat is assertedResult
19 A crash mid-release, and the watchdog has the server back up thirty seconds later. Startup recovery resolves nothing, correctly: core will not touch a live lease, because another node might be mid-delivery on it. The real ReleaseCatchUp then keeps asking on a timer until the lease does expire, and the row returns to HELD. Bounded by lease + one period after the crash, not by the next restart. Pass: row RELEASING at startup, HELD within the bound, grave lootable
20 Nothing: this is the safety half, and it is about what the catch-up window must never do. The window closes at exactly the instant a release begun by this process could first look lease-expired. A pass running at or after that instant would ask the witness mid-delivery, hear NO, put the row back to HELD, and the delivery tick would hand the items over anyway. That is a dupe, and the boundary is pinned rather than assumed. Pass: closed for good, and genuinely used before it closed

When nobody can tell: the policy path

Scenario 11 is the one that decides whether Cairn's headline claim is honest. When the delivery witness genuinely cannot answer (the player's data file is unreadable), Cairn's operator default is mark-released: assume they got the items, never duplicate. That is only defensible if the resulting loss is rare, announced, and reversible. So this scenario asserts the whole chain, not just that a policy fired.

#What a server owner would seeWhat is assertedResult
11 A crash mid-delivery, and afterwards the player's data file cannot be read at all. One outcome: the row resolves once, to one state, with no second row appearing; running recovery again changes nothing. Audited: exactly one audit entry, flagged as decided by policy rather than by evidence, naming the owner and all 64 items. Announced: a notice is written to disk for that player and survives a restart, carrying the exact /cairn restore command (necessary, because the player is by definition offline when this happens) plus an admin summary produced without anyone asking for it. Reversible: the row really does restore, rebuilding all 64 items byte-for-byte, and the ledger's arithmetic still balances with that restore counted honestly as items entering. Protected: with the world configured to on-expiry: destroy, the sweeper's planner refuses to destroy this row and downgrades it to an archive, so the remedy cannot expire on a timer. Pass: all six
12 The same expiry, on an ordinary grave nobody ever lost. The protection is not blanket: a row with no policy resolution in its audit log is destroyed on expiry as configured. The exemption is driven by evidence, not by caution. Pass: plan = EXPIRE_DESTROY

The next two are the quietest dupe in the whole retrieval path, and the one no fake could ever catch. The witness is written as the last instruction of the delivery tick. An insert that throws half way through has therefore written no witness at all, and a witness reading the player's container answers a confident NO, "the data was readable and carried none", exactly as its contract requires. Taken at face value that is proof the items never arrived, the row goes back to HELD, and the grave is lootable again while the stacks the insert already placed are sitting in the player's inventory.

#What a server owner would seeWhat is assertedResult
21 A plugin conflict or a broken inventory throws part-way through handing the items over: 20 of 40 are already placed. The outcome is undecided and policy-resolved, not witness-resolved: the row must not, and does not, go back to HELD. Under the shipped never-dupe default it closes as RELEASED. The census then states the product's whole promise as arithmetic: no id is in an inventory and in the ledger at the same time. Pass: outcome undecided, row RELEASED, 0 duplicated
22 The same, except the rollback fails too, something else took the items out from under it. Also policy-resolved rather than decided by a witness that was never written. Items with the player, ledger holding nothing that is also with the player. Pass: outcome undecided, 0 duplicated

Two stores, one crash

A crash does not roll the server back to a tick boundary. It restores each store to its own last flush, and the two are scheduled independently: the custody ledger fsyncs its frames immediately, while playerdata/<uuid>.dat waits for the next player save. So recovery can reconstruct a state that never existed in memory: the ledger holding a grave while the player's file still holds the inventory it was made from. Both of these scenarios exist because the escrow invariant is necessary and not sufficient: it governs three in-memory locations and says nothing about which store reaches disk first.

Cairn's answer is to order the flushes rather than race them. The cheap-to-lose side (the player's file) goes to disk first, on both paths, and the ledger is only told afterwards. That makes the surviving failure a loss, which is recoverable from a backup and an apology, and makes the duplication ordering unreachable. It is the same choice made about the death event above, applied to the save boundary.

#What a server owner would seeWhat is assertedResult
23 The server is killed in the moment after a death empties the player but before their save file is written. The commit is chained behind that save, not racing it, so a crash here reaches neither. The player rolls back to the gear they died with and no grave exists, a clean recovery. A durable ledger row here would be a second copy of everything they are still holding. Pass: no row, no record, 0 duplicated
24 A retrieval where the player's save cannot be made durable at all: an offline player, the wrong thread, a failing disk. The release does not confirm. Telling the ledger "delivered" on the strength of evidence that may never reach disk is the exact failure the call exists to prevent, so the insertion is taken back out and the row stays HELD. Pass: notDelivered, 0 items with the player, row HELD

The items themselves

#What a server owner would seeWhat is assertedResult
13 A player dies carrying gear with plugin-written NBT tags, an item with a custom model id, a shulker box full of netherite, and a written book with accented and non-Latin characters. The stored blob is byte-for-byte the bytes handed to the ledger, its content hash matches, it survives a full server restart, and every stack comes back identical: custom data and all. 70 item entities across 5 stacks, counted as entities and not as stacks. Pass: bytes identical, 70 items, stacks equal
14 An ordinary 41-slot inventory that is mostly empty. Empty slots keep their positions through the ledger and a restart: the axe is still recorded in slot 0, the elytra in slot 36, the totem in the off-hand, nothing shifts up. 35 item entities from 4 occupied slots. Pass: all four slots exact, 41 slots wide

Read that scenario precisely, because the obvious reading is wrong. It proves the storage format is slot-aware, not that a death round-trips a player's inventory layout. In 1.0 it does not, and the reason is covered below, in what this harness does not prove: a death capture reads event.getDrops(), which is a dense list with no slot indices in it, so there are no empty slots for the codec to preserve. Items come back into the first free slots. Slot-restoring capture and delivery is planned for a later release and is deliberately not claimed on the product page.

What scenario 14 does speak to is the complaint that shows up against plugins in this category (all the items in the game cannot be recovered), which comes from storing the hotbar only. Cairn captures the entire drop list: armour, off-hand, every stack. That is the claim, and it is a different one from slot preservation.

Volume

#What a server owner would seeWhat is assertedResult
15 Months of ordinary use: thousands of deaths and retrievals, some graves left unlooted, some players too full to take everything. The arithmetic does not drift. Every item minted over the entire run is accounted for, once, at the end, by the harness's own census and independently by the ledger's audit-log arithmetic, globally and per player. Pass: see below

Actual figures from the run this document reports:

FigureValue
Cycles2,500
Players5
Graves created2,500
Graves looted2,142
Graves left buried358
Partial deliveries195
Items buried160,000
Items delivered130,848
Items still held29,152
Duplicated0
Lost0
Elapsed0.9 s

Every figure above except elapsed is deterministic and reproduces exactly on any machine, because the run is seeded. Elapsed time describes the machine the release was built on, not the plugin. It is quoted because every number here must come from the run that produced the shipped jar, and a figure carried over from different hardware would be the one number on this page that no longer described anything.

160,000 = 130,848 + 29,152, item for item, id for id.

What this harness does not prove

An honest limitation is worth more than an overclaim, so:

How this is run, and how often

The harness is re-run in full for every release, including one-line patches, and every number on this page is copied from that run rather than from any earlier one. A one-line patch is exactly the kind of change that has broken custody in this project before, so "it obviously cannot affect custody" is not accepted as a reason to skip it.

On top of the scenarios above, each release is also run against real Paper and Folia servers by a suite of headless clients that sweeps timing, repetition, inventory state, distance, world, permissions, config and concurrency, driving the actual plugin through the actual network protocol rather than calling its classes directly.

The live suite's last run, on Folia 26.2: 190 cases, 184 passed, 0 failed, 4 skipped, 2 unobservable. On Paper 26.2: 190 cases, 184 passed, 0 failed, 4 skipped, 2 unobservable. A skipped or unobservable case is never counted as a pass.

The run this document reports

Run on 2026-09-23 against the jar with SHA-256 3674af36bf3f9b159ec7c2e538bc496d45b48c5eba9cb3ce40ae34f0b83ac72b.

RunTestsFailuresErrorsSkippedTime
Whole suite 859 0 0 0 2.3 s
Stress scenarios 24 0 0.9 s
Scenario classTestsTime
DeathCrashStressTest40.01 s
ReleaseCrashStressTest80.01 s
PartialDeliveryStressTest30.01 s
FastRestartRecoveryStressTest20.01 s
UnwitnessedAbortStressTest20.01 s
PolicyLossStressTest20.01 s
ItemFidelityStressTest20.01 s
VolumeStressTest10.90 s

Read from build/test-results/test/*.xml after the build above. Every number on this page came from that run; nothing here is estimated. The times are per-class wall time from the release build. They are a rough figure for how long the evidence takes to regather, not a benchmark of anything.

Back to the product

See Nomad Cairn for the plugin itself, or the changelog for what shipped in each release.