I Rejected the Same dev.to Reliability Fix on One Path and Shipped It on Another

← hexisteme · notes · 2026-09-27

A delivery audit led me to add external verification to manual replies and defer a polling loop for articles. The decisions depended on distinct failure modes, retry behavior, and explicit conditions for reconsideration.

Readers made a reasonable suggestion about my publishing automation: verify delivery from outside the process that claims it succeeded. I had an article publisher and a comment replier. Both sent content to dev.to. It would have been easy to apply the suggestion to both and call the result consistent.

My audit reached different decisions. I added external observation to the manual comment path and deferred a new polling loop for article publication. The useful distinction was what each checker could get wrong, what would happen afterward, and which operational decision another observation would change.

This account is based on my delivery audit dated 2026-08-08. It describes that decision and its limits, rather than asserting the same configuration remains correct indefinitely. During preparation for publication on 2026-09-26 I also checked the surviving local records. That second pass narrowed a numerical claim in the original report.

Define what a failed check means

The comment incident was a false negative: the reply had been posted, but the checker said it had not landed. Its comparison did not adequately account for the difference between authored Markdown and rendered content.

That error had an operational consequence. A reply judged absent stayed pending. Someone following the pending list could answer again, producing a duplicate. The audit connects this mechanism to previously removed duplicate replies; it does not establish that every duplicate had the same cause.

The surviving record preserves the useful evidence: a note saying that an apparent failure caused by Markdown versus rendering was corrected after a live API check. Its current verification flag is true. That flag now describes the corrected state; the note preserves the earlier mistake.

The original report presented the comment result as 1/44 and attributed the false flag to this incident. A publication-time check could reconstruct 44 earlier autoreply records, but their remaining false flag belonged to a different entry. I therefore do not use 1/44 as a measured rate of this failure. A mutable operational ledger and a historical incident count answer different questions.

The recorded false negative is enough to motivate fixing its comparison path. An unreliable rate estimate would add apparent precision without improving that argument.

The article audit asked a different question

For articles, I compared success claims with external publication state. The historical audit reports 0 mismatches among 46 successful publication log entries. The denominator is successful log entries, not every scheduled run, every network attempt, or every possible delivery failure.

The retained log still supports that success-entry count through the morning run on the audit date. The external comparison result comes from the contemporaneous audit report; I have not reconstructed its historical API snapshot for this post. That is a boundary on reproducibility, not permission to silently turn the report into a fresh measurement.

The report also distinguishes known failures from false success. The publisher's DNS failures were reported as failures and left work for a later run. They were not examples of a success claim concealing an absent article.

Zero observed mismatches does not establish zero risk. These observations also share one deployment and its operating conditions, so treating them as independent trials would need justification. I do not infer a general reliability guarantee from the sample.

The severe counterexample remained possible: if the publisher falsely claimed success and moved the file out of the queue, an absent article could disappear from ordinary retry handling. That is why deferring the polling loop required an explicit reversal condition. Lack of an observed incident did not make silent loss harmless.

An extra observation should change an action

The comment verifier had an immediate decision to make: had this reply appeared, or should the work remain pending? Its retry policy also contained tunable choices: attempts=3 and a pause of 5 seconds. Measuring when a reply became visible could inform that budget.

The article path did not have a comparable short polling budget in the audited design. Reported failures waited for the next scheduled run. A new polling loop would introduce another timeout policy, another interpretation of temporary absence, and another failure path to maintain.

I deferred that addition because the audit had found a concrete comment-checking defect and no demonstrated article success-report mismatch. This was a local prioritization decision. I did not measure the cost of a lost article, calculate an expected loss, or establish that the polling loop could never repay its complexity.

That qualification matters. A system handling more consequential delivery could reasonably require external confirmation before any false success had been observed. My decision depended on this workflow and included a trigger to revisit it.

The comparison I would reuse is small:

QuestionArticle publication in the auditManual comment delivery in the audit
What was the observed defect?No mismatch reported in the checked success entriesPosted content could be judged absent
What happened after the adverse result?A reported failure waited for a later runA reply stayed pending and could be duplicated
What could new observations tune?No existing short polling budgetAttempt count and pause duration
What did I choose?Defer the added polling loopShare and use the external reply verifier

The table does not make the risks equivalent. It exposes why the same design suggestion led to different work.

Share the comparison, then observe the retries

The normalization fix already existed in the automatic reply path but had not reached the manual path. Copying it would have preserved the condition that allowed the implementations to diverge.

I moved the comparison and landing verification into the shared comment module. The automatic path delegated to it, and the manual path used the external observation. When that observation could not decide, the manual path retained an in-page fallback and announced that fallback. A weaker observation should remain visible to the operator interpreting the result.

The comparison also needed to handle links: the authored Markdown could contain a URL that was absent from the rendered visible text. Merely discarding punctuation did not make those strings equivalent. The fix addressed this mismatch in the shared implementation.

I have left the original report's rendering-exposure count out of this post because I did not recover the corresponding historical sample. Exposure to a string transformation would not, by itself, be a count of failed deliveries anyway. The mechanism and the recorded false negative support a narrower claim.

For retries, I kept the existing constants and started recording attempt_seen, attempts_budget, elapsed_s, and landed. That made later calibration possible without presenting a newly chosen constant as empirical. The observation record separates whether a reply was seen from how much budget the verifier consumed.

Make the deferral reversible

The article decision has a sharp falsifier: observe one publication success report that disagrees with external publication state, and reopen the rejected verification work. Evaluation began with the audit on 2026-08-08; the trigger applies whenever a later mismatch is found. A check is only useful if the workflow can surface that disagreement, so this condition is not a claim of continuous detection coverage.

The comment decision has a different review trigger: after at least 30 landing observations, if every observation succeeded on its first attempt, reconsider whether the three-attempt budget is excessive. That is a prompt for reviewing the budget, not proof that later attempts will never be needed. The observation mix and any failed checks still matter.

These conditions keep the decisions attached to evidence. They also prevent consistency from becoming a reason to build the same mechanism everywhere. In this audit, one path supplied a known comparison failure and a measurable retry budget; the other supplied a bounded clean sample and a serious unobserved failure mode. Different actions were defensible only while those premises stayed visible.

The reusable lesson is to specify the decision before adding a checker: which wrong claim are you trying to catch, what happens if you miss it, and what would a new observation make you do? That question gave me a smaller repair and a clear reason to reverse the part I deferred.

Prepared with AI assistance from my delivery audit and surviving operational records.

License: CC BY 4.0.

FAQ

Why apply external verification to comments but defer it for articles?

The audit found a recorded false negative in reply checking and a retry budget that observations could inform. It reported no mismatch in the checked article success entries, so I deferred a new polling loop with a condition to reopen it.

Does a clean audit establish that publication cannot silently fail?

No. The checked sample was limited, and a false success could remove an unpublished article from ordinary retry handling. The decision was local prioritization, not a reliability guarantee.

Why avoid presenting the original comment fraction as a failure rate?

The surviving mutable ledger no longer matched the report's attribution of the false flag. The preserved incident note supports a specific false negative, but not a reconstructed historical rate.

What would reverse or revise these decisions?

An article success report that disagrees with external publication state reopens verification work. If at least 30 comment landing observations all succeed on the first attempt, the retry budget deserves review; that does not prove later attempts will never be needed.

Related notes

Fixed your problem? Good — that's the whole point of this page. Every note on this site is free to read. I write one of these up whenever I hit a failure worth recording: agent-fleet operations, harness bugs, and the measurement that proved the fix. The email list for these notes — no issue has gone out yet, so you would be on it before the first one.
← hexisteme · notes · about · editorial policy · privacy · contact · CC-BY 4.0