Nomad Cairn / Evidence
Cairn crash and duplication evidence
The crash and duplication scenarios Nomad Cairn runs before every release, what each one asserts, and where the evidence stops.
Cairn is built against a stress harness that reproduces failures buyers report against plugins that move items between owners, and asserts, for each one, that Cairn does not fail that way. It runs in full before every release, and every number on this page comes from the run that produced your build.
Read that as what it is. These are reproductions of specific, named failure modes, and passing them means Cairn does not fail in those ways, not that no failure is possible. A machine that dies hard dies in a state nobody designed, and the layer below every plugin has a hole of its own: Minecraft writes a connected player's inventory on quit and at a clean shutdown, not continuously, so items held by somebody standing still when the power goes may never have been on disk to recover. The rig measures that directly every release by SIGKILLing a server with no plugin loaded at all. What Cairn is built for is the part a plugin can control: items already in its custody, an interrupted hand-over resolving to one side rather than both, and anything undecidable being held, reported and reversible instead of quietly written off.
Specifically, the scenarios below reproduce and defend against:
- The elytra dupe. A grave written before the death drops are cleared, so the items exist in both places at once. Cairn takes the drops out of the event first; the ordering is scenario 1.
- The crash-during-payout dupe. The server dies between the items reaching a player and the ledger recording it. Both sides of that window are tested (scenarios 5 and 6), including the narrower one where the player's save had not yet happened.
- Two players looting one grave on two threads in the same millisecond. Settled by the ledger's compare-and-swap, not by a lock and not by timing (scenario 9).
- The partial-delivery dupe. A player with room for only part of a grave, where a mishandled remainder either bricks the grave or hands the same stack out twice (scenario 8).
- A plugin conflict throwing mid-handover, with half the stacks already in the inventory and nothing written down: the quietest dupe in this category, because "never delivered" is the tempting answer and the wrong one.
- Double-clicking, stale menus, and clicking a grave that is being swept out from under you.
- Item-fidelity attacks. Shulkers full of gear, custom NBT, renamed and enchanted items, written books with non-Latin text, all checked to come back byte-for-byte rather than merely by item type and count.
- Deliberate mint and deliberate loss, scripted into the live suite as negative controls, which must both be caught or the whole run is treated as untrustworthy.
It also states, in what this harness does not prove, exactly where the evidence stops, which is the part most reports of this kind leave out.
The one thing every scenario asserts
Not "the code took the branch I expected". Every scenario ends with a census:
Every item that entered is somewhere, exactly once. It is in a player's inventory, or it is in the custody ledger. Never in both. Never in neither.
The check is per item, by identity, with amounts, so a stack legitimately split between
an inventory and a grave still passes, and a stack that quietly became two does not.
Alongside it, Cairn's own ledger runs its independent conservation arithmetic over the
append-only audit log (captured + received − delivered − sent = currently held)
and that has to balance too.
How the crashes are produced, and why it matters
There are two honest ways to test a crash and only one of them is worth much.
The weak way is to hand-assemble the wreckage: write the half-finished state you believe a crash leaves, then assert the recovery handles it. That can only ever confirm what the test author already believed, and the bugs in this category live exactly in the gap between what an author believes and what the code does.
Cairn's harness does it the other way. It runs the real production
sequencer forward (the same DeathCapture and GraveRelease
classes the plugin ships) and then switches the machine off at a chosen instruction. Queued
work is dropped, exactly as it would be if the server process died holding it. Whatever
state is left behind is the state the production code produced. The harness then opens the
ledger directory again, from disk, on a fresh server object, and lets the plugin's real
startup recovery resolve whatever it finds.
The plugin's crash-safety argument leans on one property of Paper: a player's inventory
and their persistent data container are written to the same
playerdata/<uuid>.dat, so a crash rolls back both or neither. The
harness models exactly that: a simulated save flushes both together and a simulated crash
discards both together, which is why the two crash mid-delivery scenarios below can come
out differently and both be correct.
The scenarios
Item counts below are the actual figures asserted in the run.
Death: the capture path
| # | What a server owner would see | What is asserted | Result |
|---|---|---|---|
| 1 | The server dies the instant a player does, before anything has been written to disk. | The drops are already out of the death event, so the items are in neither the world nor the ledger. This is a vanilla-style loss and it is not a duplication: nothing can be looted twice. No half-written grave, no ledger row. | Pass: 64 items taken from the event, 0 rows written, 0 grave records |
| 2 | The server dies after the items are safely in the ledger but before the grave marker record is written. | No partial grave record exists at all. All 60 items are still held, still owned
by the player, and still listed by /cairn, because Cairn reads the
ledger, not the marker. |
Pass: 0 grave records, 1 held row, 60 items |
| 3 | The server dies after the grave record is written but before the marker entity spawns. | On restart the record is found still owing the world a marker, and the sweeper's own planner schedules a re-render rather than an expiry or a deletion. The 27 items never depended on the marker at any point. | Pass: marker state PENDING, plan = RERENDER, 27 items
held |
Scenario 1 is worth reading twice, because it is the whole anti-dupe argument. Cairn removes the drops from the death event before it tells the ledger about them. That ordering means the worst case is a loss; the inverse ordering, write the grave then clear the inventory, is the elytra duplication.
Retrieval: the release path
| # | What a server owner would see | What is asserted | Result |
|---|---|---|---|
| 4 | A player clicks their grave and the server dies before a single item reaches them. | On restart the delivery witness reports the items never arrived, the row goes back to held, and the grave is genuinely lootable again on the real path. | Pass: witness NO, 40 items held, re-loot delivered 40 |
| 5 | The duplication window. The server dies after the items are in the player's inventory and their data file has been saved, but before the ledger has recorded the payout. | The witness reports the items did arrive. The row closes as delivered. The player has all 64 items, the ledger has none, and three further clicks on that grave are refused. | Pass: witness YES, 64 with the player, 0 in the ledger, 0
duplicated |
| 6 | The same crash, but the save had not happened yet, so the inventory rolls back. | The witness reports the items did not arrive, because the token rolled back with them. The row goes back to held. The player has nothing and the grave still has all 64 items. | Pass: witness NO, 64 items held, 0 lost |
| 7 | A player clicks their grave with a completely full inventory. | Nothing is inserted, the row goes straight back to held under the same custody id, and the grave is untouched. Make room and the retry delivers all 60. | Pass: outcome notDelivered, same id, 60 items held, retry delivered
60 |
| 8 | A player clicks with room for only part of the grave. | 40 items go to the player, the remaining 24 stay in the ledger under a fresh held row handed back by the ledger itself. 40 + 24 = 64, with nothing created and nothing lost. | Pass: 40 delivered, 24 held, census exact |
| 9 | Two players click the same grave at the same instant (two real threads, released together). | Exactly one gets a release ticket. The other is refused by the ledger's compare-and-swap, not by luck or a lock. No second licence to touch the items is ever issued. | Pass: 1 ticket, 1 refusal, 64 items intact |
| 10 | A lagging client double-fires, or a player clicks a grave they already emptied, three more times. | Every subsequent attempt is refused as contested. The player's inventory still holds exactly one payout of 40. | Pass: 1 delivery, 3 refusals, 40 items |
Retrieval that did not fit in one go: the remainder path
These three run against the real on-disk ledger, with its real idempotency
store, and they exist because the rest of the retrieval suite does not: its
fixture ledger has no idempotency store at all, so several hundred passing tests said
nothing about the one thing core enforces on beginRelease: that a key
presented twice for two different requests is a conflict, and is refused for ever. A
partial delivery re-points the grave at a brand-new custody row whose sequence starts back
at zero, which is exactly the shape that can rebuild an identical key against a different
row. A fully-geared death takes this path every time: addItem fills
only the 36 storage slots, so 36 + 4 armour + off-hand always overflows.
| # | What a server owner would see | What is asserted | Result |
|---|---|---|---|
| 16 | A player loots with room for 30 of the 50 items in the grave, then comes back for the rest. | The first click delivers 30 and core mints a remainder row for the other 20,
HELD, at sequence zero. The second click, keyed on the custody
row, not the grave id, delivers all 20. Keyed on the grave id instead, the
second click is an idempotency conflict that every later click repeats for ever:
the grave is bricked and the items are stranded with no operator remedy, because
/cairn restore refuses a HELD row. |
Pass: 30 + 20, census exact, ledger holds nothing back |
| 17 | The same thing twice in a row: a remainder that is itself too big to fit. | Every remainder gets its own key. Three clicks, 20 + 20 + 20, and the census still balances against the 60 minted. | Pass: 60 delivered, 0 duplicated |
| 18 | A genuine retry of the same click: a lagging client, a re-sent packet. | The second beginRelease with the same key is refused with
IllegalTransitionException. Core will not hand out a second licence
to mutate the world, and GraveRelease maps that refusal to
"contested" rather than to "broken". |
Pass: 1 ticket, second attempt refused |
The restart that is quick enough to miss its own recovery
The rest of the crash suite advances its clock past the release lease before every reboot, so the one restart timing that actually happens in production, a watchdog bringing the server back inside a minute, was the one nothing exercised. Under it, startup recovery correctly finds no stale row (the lease has minutes left), leaves the interrupted release alone, and without a catch-up pass never asks again: the items are safe and the grave is dead to the player until some later restart happens to fall past the lease.
| # | What a server owner would see | What is asserted | Result |
|---|---|---|---|
| 19 | A crash mid-release, and the watchdog has the server back up thirty seconds later. | Startup recovery resolves nothing, correctly: core will not touch a live lease,
because another node might be mid-delivery on it. The real
ReleaseCatchUp then keeps asking on a timer until the lease does
expire, and the row returns to HELD. Bounded by lease + one
period after the crash, not by the next restart. |
Pass: row RELEASING at startup, HELD within the bound,
grave lootable |
| 20 | Nothing: this is the safety half, and it is about what the catch-up window must never do. | The window closes at exactly the instant a release begun by this
process could first look lease-expired. A pass running at or after that instant
would ask the witness mid-delivery, hear NO, put the row back to
HELD, and the delivery tick would hand the items over anyway. That is
a dupe, and the boundary is pinned rather than assumed. |
Pass: closed for good, and genuinely used before it closed |
When nobody can tell: the policy path
Scenario 11 is the one that decides whether Cairn's headline claim is honest. When the
delivery witness genuinely cannot answer (the player's data file is unreadable), Cairn's
operator default is mark-released: assume they got the items, never
duplicate. That is only defensible if the resulting loss is rare, announced, and
reversible. So this scenario asserts the whole chain, not just that a policy fired.
| # | What a server owner would see | What is asserted | Result |
|---|---|---|---|
| 11 | A crash mid-delivery, and afterwards the player's data file cannot be read at all. | One outcome: the row resolves once, to one state, with no
second row appearing; running recovery again changes nothing. Audited:
exactly one audit entry, flagged as decided by policy rather than by evidence,
naming the owner and all 64 items. Announced: a notice is written
to disk for that player and survives a restart, carrying the exact
/cairn restore command (necessary, because the player is by
definition offline when this happens) plus an admin summary produced without
anyone asking for it. Reversible: the row really does restore,
rebuilding all 64 items byte-for-byte, and the ledger's arithmetic still balances
with that restore counted honestly as items entering. Protected:
with the world configured to on-expiry: destroy, the sweeper's
planner refuses to destroy this row and downgrades it to an archive, so the remedy
cannot expire on a timer. |
Pass: all six |
| 12 | The same expiry, on an ordinary grave nobody ever lost. | The protection is not blanket: a row with no policy resolution in its audit log is destroyed on expiry as configured. The exemption is driven by evidence, not by caution. | Pass: plan = EXPIRE_DESTROY |
The next two are the quietest dupe in the whole retrieval path, and the one no fake could
ever catch. The witness is written as the last instruction of the
delivery tick. An insert that throws half way through has therefore written no witness at
all, and a witness reading the player's container answers a confident NO,
"the data was readable and carried none", exactly as its contract requires. Taken at face
value that is proof the items never arrived, the row goes back to HELD, and
the grave is lootable again while the stacks the insert already placed are sitting in
the player's inventory.
| # | What a server owner would see | What is asserted | Result |
|---|---|---|---|
| 21 | A plugin conflict or a broken inventory throws part-way through handing the items over: 20 of 40 are already placed. | The outcome is undecided and policy-resolved, not
witness-resolved: the row must not, and does not, go back to HELD.
Under the shipped never-dupe default it closes as RELEASED. The
census then states the product's whole promise as arithmetic: no id is in an
inventory and in the ledger at the same time. |
Pass: outcome undecided, row RELEASED, 0 duplicated |
| 22 | The same, except the rollback fails too, something else took the items out from under it. | Also policy-resolved rather than decided by a witness that was never written. Items with the player, ledger holding nothing that is also with the player. | Pass: outcome undecided, 0 duplicated |
Two stores, one crash
A crash does not roll the server back to a tick boundary. It restores each store to
its own last flush, and the two are scheduled independently: the custody ledger
fsyncs its frames immediately, while playerdata/<uuid>.dat waits for
the next player save. So recovery can reconstruct a state that never existed in memory:
the ledger holding a grave while the player's file still holds the inventory it was made
from. Both of these scenarios exist because the escrow invariant is necessary and not
sufficient: it governs three in-memory locations and says nothing about which store
reaches disk first.
Cairn's answer is to order the flushes rather than race them. The cheap-to-lose side (the player's file) goes to disk first, on both paths, and the ledger is only told afterwards. That makes the surviving failure a loss, which is recoverable from a backup and an apology, and makes the duplication ordering unreachable. It is the same choice made about the death event above, applied to the save boundary.
| # | What a server owner would see | What is asserted | Result |
|---|---|---|---|
| 23 | The server is killed in the moment after a death empties the player but before their save file is written. | The commit is chained behind that save, not racing it, so a crash here reaches neither. The player rolls back to the gear they died with and no grave exists, a clean recovery. A durable ledger row here would be a second copy of everything they are still holding. | Pass: no row, no record, 0 duplicated |
| 24 | A retrieval where the player's save cannot be made durable at all: an offline player, the wrong thread, a failing disk. | The release does not confirm. Telling the ledger "delivered" on the strength of
evidence that may never reach disk is the exact failure the call exists to
prevent, so the insertion is taken back out and the row stays HELD. |
Pass: notDelivered, 0 items with the player, row HELD |
The items themselves
| # | What a server owner would see | What is asserted | Result |
|---|---|---|---|
| 13 | A player dies carrying gear with plugin-written NBT tags, an item with a custom model id, a shulker box full of netherite, and a written book with accented and non-Latin characters. | The stored blob is byte-for-byte the bytes handed to the ledger, its content hash matches, it survives a full server restart, and every stack comes back identical: custom data and all. 70 item entities across 5 stacks, counted as entities and not as stacks. | Pass: bytes identical, 70 items, stacks equal |
| 14 | An ordinary 41-slot inventory that is mostly empty. | Empty slots keep their positions through the ledger and a restart: the axe is still recorded in slot 0, the elytra in slot 36, the totem in the off-hand, nothing shifts up. 35 item entities from 4 occupied slots. | Pass: all four slots exact, 41 slots wide |
Read that scenario precisely, because the obvious reading is wrong. It
proves the storage format is slot-aware, not that a death round-trips a player's
inventory layout. In 1.0 it does not, and the reason is covered below, in what this
harness does not prove: a death capture reads event.getDrops(), which is a
dense list with no slot indices in it, so there are no empty slots for the codec to
preserve. Items come back into the first free slots. Slot-restoring capture and delivery
is planned for a later release and is deliberately not claimed on the product page.
What scenario 14 does speak to is the complaint that shows up against plugins in this category (all the items in the game cannot be recovered), which comes from storing the hotbar only. Cairn captures the entire drop list: armour, off-hand, every stack. That is the claim, and it is a different one from slot preservation.
Volume
| # | What a server owner would see | What is asserted | Result |
|---|---|---|---|
| 15 | Months of ordinary use: thousands of deaths and retrievals, some graves left unlooted, some players too full to take everything. | The arithmetic does not drift. Every item minted over the entire run is accounted for, once, at the end, by the harness's own census and independently by the ledger's audit-log arithmetic, globally and per player. | Pass: see below |
Actual figures from the run this document reports:
| Figure | Value |
|---|---|
| Cycles | 2,500 |
| Players | 5 |
| Graves created | 2,500 |
| Graves looted | 2,142 |
| Graves left buried | 358 |
| Partial deliveries | 195 |
| Items buried | 160,000 |
| Items delivered | 130,848 |
| Items still held | 29,152 |
| Duplicated | 0 |
| Lost | 0 |
| Elapsed | 0.9 s |
Every figure above except elapsed is deterministic and reproduces exactly on any machine, because the run is seeded. Elapsed time describes the machine the release was built on, not the plugin. It is quoted because every number here must come from the run that produced the shipped jar, and a figure carried over from different hardware would be the one number on this page that no longer described anything.
160,000 = 130,848 + 29,152, item for item, id for id.
What this harness does not prove
An honest limitation is worth more than an overclaim, so:
- Twenty of the twenty-four scenarios drive the real production sequencer to
the crash point. Scenarios 1–8, 10–13, 15–17, 19, 21 and 22 run
DeathCaptureandGraveRelease(the shipped classes) forward until the machine is switched off, and then let the shipped recovery resolve it. Nothing about the wreckage is assembled by hand. - Four are partly hand-driven, and here is exactly which parts.
Scenario 9 calls the ledger's
beginReleasedirectly from two threads, because the race being tested lives inside a single ledger call; the arguments it presents are the ones production presents, derived the same way. Scenario 14 writes to the ledger directly rather than through the death path, because a death cannot carry empty slots into the capture in the first place. Scenario 18 captures through the real death path but then presents the retry tobeginReleaseitself, because the property under test is a property of the idempotency key's format, which is the one thing a test is entitled to know. Scenario 20 drives onlyReleaseCatchUpand a clock: it is a statement about a boundary in time, and involves no items at all. - Scenarios 11 and 12 reproduce six lines rather than execute them.
They run the sweeper's decision code (
SweepSelector) against the real audit log, but the sweeper class itself needs a live Bukkit server, so its "was this row ever policy-resolved" lookup is reproduced in the test. - The last step of the death path is reproduced, not executed. Writing the grave record and then spawning the marker lives in the Paper listener, which cannot be loaded without a server. The harness performs those two steps in the same order the listener does. Everything before them (the decision, taking the drops, encoding, and the durable commit) is the shipped code.
- The player's inventory, the delivery witness and the two schedulers are simulated. paper-api is deliberately absent from Cairn's test classpath: nothing on the custody side of the boundary is allowed to see a Bukkit type, and a test that can build one has quietly acquired a server. The simulated pieces hold state and stop when told to; they decide nothing.
- A "crash" closes the ledger and reopens it from disk. That proves the row was genuinely durable and genuinely re-read, which is what these scenarios are about. It does not simulate a power cut during an fsync: recovery from a torn write is the custody engine's own concern and is covered by its separate test suite, not by this one.
- The volume run uses an in-memory ledger. The crash scenarios all use the real on-disk one; the volume run trades that for speed so it finishes in under half a minute on a laptop.
- Fake item objects, not Minecraft
ItemStacks. The framing codec that decides slot layout, item counts and blob structure is the shipped one; the platform's own item serialiser sits behind an interface and is stood in for. Whether Mojang's serialiser round-trips a stack is Mojang's problem, and every plugin in this category inherits it equally.
How this is run, and how often
The harness is re-run in full for every release, including one-line patches, and every number on this page is copied from that run rather than from any earlier one. A one-line patch is exactly the kind of change that has broken custody in this project before, so "it obviously cannot affect custody" is not accepted as a reason to skip it.
On top of the scenarios above, each release is also run against real Paper and Folia servers by a suite of headless clients that sweeps timing, repetition, inventory state, distance, world, permissions, config and concurrency, driving the actual plugin through the actual network protocol rather than calling its classes directly.
The live suite's last run, on Folia 26.2: 190 cases, 184 passed, 0 failed, 4 skipped, 2 unobservable. On Paper 26.2: 190 cases, 184 passed, 0 failed, 4 skipped, 2 unobservable. A skipped or unobservable case is never counted as a pass.
The run this document reports
Run on 2026-09-23 against the jar with SHA-256
3674af36bf3f9b159ec7c2e538bc496d45b48c5eba9cb3ce40ae34f0b83ac72b.
| Run | Tests | Failures | Errors | Skipped | Time |
|---|---|---|---|---|---|
| Whole suite | 859 | 0 | 0 | 0 | 2.3 s |
| Stress scenarios | 24 | 0 | 0.9 s |
| Scenario class | Tests | Time |
|---|---|---|
DeathCrashStressTest | 4 | 0.01 s |
ReleaseCrashStressTest | 8 | 0.01 s |
PartialDeliveryStressTest | 3 | 0.01 s |
FastRestartRecoveryStressTest | 2 | 0.01 s |
UnwitnessedAbortStressTest | 2 | 0.01 s |
PolicyLossStressTest | 2 | 0.01 s |
ItemFidelityStressTest | 2 | 0.01 s |
VolumeStressTest | 1 | 0.90 s |
Read from build/test-results/test/*.xml after the build above. Every number
on this page came from that run; nothing here is estimated. The times are per-class wall
time from the release build. They are a rough figure for how long the evidence takes to
regather, not a benchmark of anything.
Back to the product
See Nomad Cairn for the plugin itself, or the changelog for what shipped in each release.