<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:media="http://search.yahoo.com/mrss/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/"><channel><title>Silent Failure on Gruion</title><link>https://www.gruion.com/blog/tags/silent-failure/</link><description>Recent content in Silent Failure on Gruion</description><generator>Hugo</generator><language>en</language><lastBuildDate>Sun, 11 Oct 2026 06:01:24 +0000</lastBuildDate><atom:link href="https://www.gruion.com/blog/tags/silent-failure/index.xml" rel="self" type="application/rss+xml"/><item><title>A green checkmark proves the job ran, not that it did anything</title><link>https://www.gruion.com/blog/post/2026-10-11-when-deploys-break/</link><pubDate>Sun, 11 Oct 2026 06:01:24 +0000</pubDate><dc:creator>Gruion</dc:creator><guid>https://www.gruion.com/blog/post/2026-10-11-when-deploys-break/</guid><description>An AI agent that does nothing still reports success. Count real outcomes like orders, emails and records, and decide in advance who gets woken up.</description><content:encoded><![CDATA[<p>Here is a story we have seen in different forms, told as a composite so no one&rsquo;s real business is exposed.</p>
<p>A founder runs an automation that sends a welcome email and sets up the account for every new paying customer. It runs every night. Someone makes a small change on a Tuesday afternoon, and an access key quietly stops working. The nightly job still starts. It finds nothing it is allowed to touch, finishes in four seconds instead of its usual ten minutes, and reports <strong>Success</strong>.</p>
<p>Nobody looks at it for nine days. Then a customer writes: &ldquo;I paid a week and a half ago. Where&rsquo;s my account?&rdquo; The founder opens the dashboard. Every run for nine days is green. Eleven customers never got onboarded, and two have already asked for refunds.</p>
<p>Nothing crashed. No alarm could go off, because nothing happened.</p>
<h2 id="why-the-checkmark-lied">Why the checkmark lied</h2>
<p>A recent DevOps.com piece describes this exactly. An agent that does the wrong thing leaves evidence: a bad record, a wrong email, a complaint. An agent that does nothing and reports success leaves a green checkmark, &ldquo;the one signal your incident process is built to trust.&rdquo;</p>
<p>Almost every safety net we build assumes something <em>arrives</em>. An approval step reviews outputs. A review meeting looks at what exists. A feedback loop needs a result to learn from. A job that produces nothing never reaches any of them. It doesn&rsquo;t fail the review. It never gets there.</p>
<p>The common reaction is &ldquo;we need better monitoring.&rdquo; The same article argues this is the wrong reflex, and we agree. More logs and traces give you more detail about the <em>job</em>. This failure is only visible in the <em>result</em>. A perfectly detailed record of a job that did nothing is still a record of nothing.</p>
<p>Our position, which some tool vendors won&rsquo;t like: <strong>a &ldquo;done&rdquo; status is the agent&rsquo;s opinion of itself.</strong> You wouldn&rsquo;t accept that from an employee on payroll, and you shouldn&rsquo;t accept it from software.</p>
<h2 id="count-the-world-not-the-report">Count the world, not the report</h2>
<p>The fix is one move. Stop asking &ldquo;did the run error?&rdquo; and ask &ldquo;does the thing it was supposed to change show the change?&rdquo;</p>
<p>For our founder, that means a separate check that doesn&rsquo;t depend on the agent at all. Every morning it counts how many customers paid yesterday and how many welcome emails and accounts exist for them. If the numbers don&rsquo;t match, it raises an alarm.</p>
<p>Two details from the source are worth stealing:</p>
<ul>
<li><strong>Treat suspicious speed as a failure.</strong> A ten-minute job that finishes in four seconds didn&rsquo;t get efficient. It skipped the hard part. That should trigger the same alert as an error.</li>
<li><strong>The checker must fire because work was due, not because the agent reported in.</strong> A reviewer that only wakes up when the agent calls it inherits the agent&rsquo;s silence. No run, no review. The check runs on a clock the agent doesn&rsquo;t control.</li>
</ul>
<p>Make the checker boring, too. Plain, fixed steps that do the same thing every time. A separate piece we read this month makes the point well: use a simple scripted robot for stable, repeatable work, and save AI reasoning for the parts that genuinely change. Counting rows is not a job for AI. Don&rsquo;t ask the thing being checked to check itself.</p>
<h2 id="the-part-that-actually-broke-nobody-knew-who-to-wake">The part that actually broke: nobody knew who to wake</h2>
<p>Here is the uncomfortable truth behind our story. The tool wasn&rsquo;t the problem. Access keys expire, and agents misbehave. That&rsquo;s normal.</p>
<p>The real failure was that <strong>no one had been named as the person who gets woken up.</strong> The agent&rsquo;s builder assumed the founder would notice. The founder assumed the builder would. A dashboard existed, and it was green. Nine days is what happens when responsibility belongs to everyone.</p>
<p>Verification costs money, and the source is honest about that: you are paying to confirm work you already paid for. It&rsquo;s cheaper than eleven lost customers.</p>
<h2 id="the-one-habit">The one habit</h2>
<p>Before any agent goes live, write one sentence for it and put a name on it. This is the only artifact you need:</p>
<blockquote>
<p><strong>Every [day/hour], there should be at least [N] [orders / emails / records] in [place]. If there aren&rsquo;t, [one named person] gets a text by [time].</strong></p>
</blockquote>
<p>Rules for the sentence:</p>
<ol>
<li>It counts things in your actual product or inbox, not things the agent says it did.</li>
<li>&ldquo;Zero&rdquo; counts as an alarm when you expect more than zero.</li>
<li>The named person is a human with a phone, not a team channel, and they know it&rsquo;s them.</li>
<li>It runs on a schedule, independent of the agent.</li>
</ol>
<p>If you can&rsquo;t write the sentence, that tells you something. The source says a run with no declarable effect deserves hard questions. If you can&rsquo;t say what the agent should change in the world, you can&rsquo;t know when it has stopped changing it.</p>
<h2 id="what-to-do-this-week">What to do this week</h2>
<p>List every automation or agent that touches a paying customer. For each one, write the sentence above. You&rsquo;ll likely find that two or three have no obvious count to check and no obvious owner. Those are your nine-day-outage candidates.</p>
<p>Build the counters first and the fancy dashboards never. A number that says &ldquo;11 expected, 0 delivered&rdquo; is worth more than a month of green.</p>
<h2 id="sources">Sources</h2>
<ul>
<li><a href="https://devops.com/the-agent-failure-your-approval-gate-cant-catch/">https://devops.com/the-agent-failure-your-approval-gate-cant-catch/</a></li>
<li><a href="https://devops.com/quality-debt-is-the-new-technical-debt/">https://devops.com/quality-debt-is-the-new-technical-debt/</a></li>
</ul>
<hr>
<p><strong>Fractional CTO</strong> — from €4,000 / month, 4–6 days a month, 30 days notice. hands-on engineering and architecture every month: the technical co-founder you have not hired yet.</p>
<p>It starts with a free product teardown: two hours on what you are building, what already exists, and what is actually blocking launch. You leave with a written plan and a realistic number, whether or not you work with us. <a href="https://www.gruion.com/#contact">Book a teardown</a> · <a href="https://www.gruion.com/services-pricing.html">What it costs</a></p>
]]></content:encoded><enclosure url="https://www.gruion.com/blog/post/2026-10-11-when-deploys-break/cover.jpg" type="image/jpeg" length="0"/><media:content url="https://www.gruion.com/blog/post/2026-10-11-when-deploys-break/cover.jpg" medium="image" type="image/jpeg"/><media:thumbnail url="https://www.gruion.com/blog/post/2026-10-11-when-deploys-break/cover.jpg"/><category>Reliability</category></item></channel></rss>