An Atomic JSON Write Did Not Protect My Python Retry Budget.

← hexisteme · notes · 2026-10-10

My comment pipeline wrote each JSON file atomically, but when saving a completed model verdict failed or the process owning the call was killed, a queued recheck called the model again and the call log read ['first', 'second']. The fix reserves and counts the attempt durably before the API call, under a pipeline lock shared by both entrypoints, and 38 focused concurrency and recovery tests, including real-subprocess cases, passed. Writes are still individually atomic rather than one transaction, so the pipeline is not exactly-once.

My comment pipeline can attach a model verdict to a dev.to comment: a small structured JSON judgment stored next to the comment. Sometimes the verdict fails. The model returns malformed JSON, or the provider is overloaded. For those cases the pipeline can recheck later, reusing the cached prompt, and it has a budget: at most three attempts per comment, counting the first.

The pipeline is one canonical Python watcher. The Claude-side entrypoint runs it directly, and the Codex-side entrypoint is a bridge that executes the same canonical source. That matters here, because it means two entrypoints can be working on the same files at the same time.

By the time I looked at concurrency, every file write the watcher made was atomic: write a temporary file, then replace the original, so no reader ever sees half a JSON document. It turned out not to protect the thing I cared about.

Two ways to pay for one attempt twice

Local rechecks of cached failed verdicts had concurrency and crash holes. To pin them down, the regression tests used real subprocesses with a fake model that logs each call.

In the first case, the model answered, but saving the completed verdict failed. In the second, the process that owned the model call was terminated. Both had a second recheck queued behind the first. In both, the call log came back as ['first', 'second']. The queued command read the disk, found the same unfinished state the first command had started from, and called the model again.

Neither failure needed a torn file. Every file could be a complete, valid JSON document and the budget would still be wrong.

Why atomic writes did not help

An atomic replace protects one file at one moment. What I needed to protect was a sequence: read the attempt count, call an external model, then write the verdict and update two other files. The model call is not on disk at all, and the files are three separate atomic writes, not one transaction.

The attempt only became visible on disk when its result was saved. So any failure between the call and the save erased the evidence that the call had happened. Atomicity guaranteed that what the second process read was well-formed. It said nothing about whether it was current.

The repair

The repair moved the count to before the call and put the whole pipeline under one lock.

Scan and recheck now share a pipeline lock, taken from both the canonical and the bridge entrypoints. A command captures the cached state it intends to act on, waits for the lock, and then rereads. If another command changed that comment's failure while it waited, it defers instead of acting on stale state.

Before contacting the provider, a cached retry atomically reserves the next attempt in the comment's inbox file, keeping the prompt and the history of earlier errors. That reservation is the count. If the reservation cannot be written, there is no model call.

If the owner then dies, or its completion cannot be saved, the counted unfinished attempt stays on disk. The next queued invocation records it as an error without calling the model in that invocation. A later explicit recheck can try again within whatever budget remains. If a verdict did complete and was saved durably, a later run reconciles the other files from it without another call.

The shared "seen" state got its own short lock. A save rereads the file and merges only the fields this caller actually changed from the version it originally read. Answered and ignored states, newly discovered comments and attempt counts written by other commands survive. Attempt counts never decrease. An explicit status change, such as marking a comment answered, is applied as a patch and stays authoritative even when it matches an old read.

Conceptually, a cached retry now runs in this order. This is a sketch, not the watcher's code:

take the pipeline lock (wait at most 5 s, else exit with an error)
reread this comment's state
if attempts used >= the saved cap: stop and report it as exhausted
durably reserve and count the next attempt
if the reservation failed: stop, no model call
call the model once
save the verdict (if this fails, the attempt stays counted)

The bounds around it are deliberately boring. Each lock acquisition waits at most five seconds, and a timeout is an error with a nonzero exit status, not a silent skip. Each comment gets at most one model call per invocation. A batch defaults to five comments and is hard-capped at twenty. The first rate-limit response, HTTP 429, stops the batch. There is no retry loop around it.

Validation

A focused set of 38 concurrency and recovery tests passed, 19 of each. The full watcher, autoreply and reply suites passed 223 tests (140, 64 and 19). Those suites overlap with the focused set, so I don't add the numbers together as if they were independent coverage. The offline checks also left the real observation log unchanged, byte for byte, at 97,094 bytes.

After that, a live recovery of the three cached failures I actually had checked three and recovered three, with one model call each. Running the same recheck again made zero additional calls. That is a small, scoped piece of evidence, not a load test.

What it still does not do

The writes are individually atomic. They are still not a transaction across files, and the pipeline as a whole is not exactly-once.

The reservation is conservative. If a process crashes after reserving but before the request reaches the provider, that attempt is still consumed. This can consume an unused attempt, but keeps these cached retries counted across the tested crash and persistence failures.

The reservation only covers cached retries. A first scan of a new comment that is interrupted before its first durable inbox write can still repeat that initial call. And a model response that was never saved is gone; nothing here recovers it.

The locks are cooperative. They protect writers that go through the watcher's own functions. A process that overwrote these files directly would bypass them.

The underlying idea is the one from The Upload Succeeded, the Record Did Not, where an upload stage recorded its result only after verification and a rerun uploaded a duplicate. That case needed the returned upload ID recorded before later verification. Here the attempt reservation is recorded before the model call. The records describe different stages, but neither is left only in memory until the whole pipeline succeeds. Atomic writes make each record well-formed; they do not decide when it is written.

FAQ

Why doesn't an atomic file write protect a retry counter?

An atomic replace guarantees a reader never sees a half-written file. A retry is a sequence: read a count, call an external model, then write several files. Because the attempt only became visible on disk when its result was saved, any failure between the call and the save left no record that the call had happened.

What happened when the save failed or the process was killed?

Real-subprocess regression tests showed a queued recheck calling the model again, logging ['first', 'second'] instead of one call, both after a completion save failure and after the process owning the call was terminated.

What does reserving the attempt before the call change?

The next attempt is counted durably in the inbox before the provider is contacted, and if the reservation fails there is no call. If the owner dies afterwards, the next queued invocation records the unfinished attempt as an error without calling the model; a later explicit recheck can retry within the remaining cap of three attempts.

Is the pipeline now exactly-once?

No. Writes are individually atomic, not a transaction across files. A reservation can consume an attempt even if the request never reached the provider, a first scan interrupted before its first inbox write can repeat its initial call, and an unsaved model response cannot be recovered.

Related notes

Fixed your problem? Good — that's the whole point of this page. Every note on this site is free to read. I write one of these up whenever I hit a failure worth recording: agent-fleet operations, harness bugs, and the measurement that proved the fix. The email list for these notes — no issue has gone out yet, so you would be on it before the first one.
← hexisteme · notes · about · editorial policy · privacy · contact · CC-BY 4.0