Here is the usual version of this story. A founder has an automation that does one important job. A customer fills in a form, the automation checks the details, creates their account, sends a welcome email and tells the team. It has run for four months without trouble. Then, on a Tuesday morning, it stops.

Nothing crashes and no red banner appears. The form still says “Thanks, we’ll be in touch.” The customer sees a success message. On your side, nothing happens.

On Friday, someone asks why a customer never got their login. You open the automation and find a column of failed runs starting Tuesday at 09:12. In that time, say twelve people signed up. Four of them have already written to ask what’s going on. Two have asked for a refund. The rest never said anything. They just left.

The fix took ten minutes. Finding out took three days. That gap is the whole problem.

Why it went quiet

The cause is almost never exotic. Three boring things account for most of it.

A login expired. Your automation talks to other services (your email tool, your spreadsheet, your payment provider) using a permission, and those permissions expire. Someone changes a password, or a connected account gets re-approved, and the automation is suddenly locked out. It keeps trying and keeps failing.

Something upstream changed. A service renamed a field, tightened a limit, or returned an unexpected answer. The automation expected “email” and got “email_address”. It doesn’t improvise. It stops.

A limit was reached. Your plan allows a set number of runs a month, or the AI service you call ran out of credit. The automation doesn’t warn you. It simply declines to start.

None of these is a bug in the usual sense. They’re the normal wear of connecting five services you don’t control. Every automation will hit one of them eventually. The only variable is how long you spend not knowing.

The part nobody tells you about these platforms

Tools like n8n and Make are good at running things and fairly good at recording failures. They are not good at telling you, on their own, in a place you’ll actually look.

The failure log exists. Nobody reads it. A log you have to go and visit is not monitoring. It is an archive you consult after the damage.

There is a second, nastier case. Sometimes the automation doesn’t fail at all. It just doesn’t start, because the trigger stopped firing. A form connection dropped, or a schedule got paused. There’s no failed run to find, because there was no run. You can have a perfectly clean error log and a dead system.

That second case is why “turn on error emails” is not enough. Error emails tell you when something tried and failed. They say nothing about something that never tried.

What I’d tell you over coffee

Don’t build a dashboard. Don’t buy a monitoring product. Don’t hire someone to “own observability”. At your size, all of that is a distraction.

Do one thing: make silence the alarm.

Take your three most important automations, the ones where a stop costs you a customer or money. For each, ask one question: how often should this succeed, and what do I do if it doesn’t for that long?

Then set it up so the automation reports in, not out. At the end of every successful run, it sends a small “I’m alive” signal to a free watcher service. If the signal doesn’t arrive in the expected window, the watcher texts you. A failure, a paused schedule, an expired login and a dead trigger all look the same to it: the heartbeat didn’t come. You find out in hours, not days.

This is the only artifact you need. Copy it, fill it in, and put it somewhere you’ll see it:

Silence rule. For each critical automation, write down: (1) the longest it should ever go without a successful run, and (2) who gets a text if it exceeds that. Anything you can’t fill in is not being monitored; it is being hoped about.

For a lead form that normally gets a few submissions a day, “24 hours” is the right window. For something that runs hourly, it’s three hours. For something that should run on the first of the month, it’s the second. A few minutes per automation, once.

Why I’m being this blunt

Because the cost is lopsided. Setting this up takes an afternoon. Not setting it up costs you nothing for months, then a lot in one bad week. That shape, cheap to ignore until suddenly very expensive, is exactly why founders skip it. It always feels like next month’s job.

It’s also the kind of work that has no glamour. Nobody demos a heartbeat. But the difference between a product customers trust and one they quietly abandon is often just how quickly you notice you’re broken. Customers forgive an outage you announce on Tuesday. They don’t forgive one you discover on Friday because they told you.

Boring until it isn’t. The habit is boring. The Friday is not.

This week, pick your most valuable automation and do the silence rule on it. If you can’t say how long it could be dead before you’d notice, you already have your answer.


Fractional CTO — from €4,000 / month, 4–6 days a month, 30 days notice. hands-on engineering and architecture every month: the technical co-founder you have not hired yet.

It starts with a free product teardown: two hours on what you are building, what already exists, and what is actually blocking launch. You leave with a written plan and a realistic number, whether or not you work with us. Book a teardown · What it costs