HAMMAD YOUSUF

AI AGENTS

4 min read · 2026-08-09

n8n vs custom code for business automation: how I decide

TL;DR

Use n8n when the workflow is integration-heavy — many third-party APIs, webhooks, schedules — and non-technical stakeholders need to see and edit the flow. Use custom code when you need evaluation and observability hooks, complex branching, sub-second latency, or a real test suite. The pattern that works best in production is hybrid: n8n as the orchestration shell, calling custom services over HTTP or MCP for the reasoning-heavy core.

My decision rule after running both in production: n8n for integration-heavy workflows that non-technical people need to see and edit, custom code for anything that needs testing, evaluation hooks, or sub-second latency — and a hybrid of the two for most serious agent systems. The 'no-code vs code' framing misses the point; both tools work. The real question is where each one's failure modes land on your specific project, and that is decidable with a short checklist rather than a philosophy debate.

The real question is where the failure modes hurt

Every automation tool fails somewhere; you are choosing which failures you can live with. n8n's failures are about scale of logic: workflows that grow past a certain complexity become harder to reason about than the equivalent code, and testing them properly is awkward. Custom code's failures are about scale of maintenance: every integration, retry policy, and scheduler you hand-roll is code you now own forever, and it is invisible to anyone who cannot read it. Neither failure mode is theoretical — I have hit both — so the framework below is really a map of which one your project is more exposed to.

Where n8n wins

n8n earns its place in three situations. First, integration breadth: when a workflow touches many third-party APIs — a CRM, a spreadsheet, an email provider, a chat webhook — n8n's prebuilt nodes wire it up in hours, versus days of reading API docs and writing auth handling. Second, stakeholder visibility: the visual canvas is genuinely auditable by non-developers, which matters when a client or an ops lead needs to understand and occasionally tweak what the automation does without opening an editor. Third, infrastructure you would otherwise build: retries, error branches, cron scheduling, webhook endpoints, and execution logs come for free, and a solo builder should not underestimate how much time that saves.

Where n8n starts to hurt

The pain arrives on predictable fronts. Complex conditional logic turns the canvas into spaghetti — a decision tree that reads as fifteen clean lines of TypeScript becomes a tangle of IF nodes nobody wants to touch. The testing story is weak: there is no natural way to unit-test a node graph the way you test a function, so verification stays manual. Version control chafes: workflows export as large JSON blobs, and a git diff of two exports tells you almost nothing about what actually changed — a real problem when you want to review changes before they run against production systems. And cost structure matters at volume: cloud pricing is per-execution, so high-frequency workflows get expensive, while self-hosting is cheap in cash but makes you the person responsible for uptime, upgrades, and backups.

Where custom code wins

ONE TACTIC A WEEK

One tactic a week. No filler.

Code wins wherever the logic itself is the product. If the workflow needs evaluation hooks — scoring outputs, gating deploys on eval results — that harness wants to live in a codebase, not a node graph. Same for observability: emitting structured traces from every step is natural in code and clumsy in n8n. Sub-second latency requirements rule out a workflow engine's overhead entirely. Complex state machines with many branches read far better as code, sit in the same test suite as everything else, and go through the same pull-request review. And anything you expect to maintain for years benefits from the boring virtues: types, tests, diffs, and refactoring tools.

The hybrid pattern that actually works

The setup I keep landing on treats n8n as the shell and code as the core. n8n owns the edges: webhook triggers, schedules, third-party API calls, retry policies, and the visual overview a stakeholder can audit. The reasoning-heavy or latency-sensitive middle is a custom service the workflow calls through an HTTP node — or, increasingly, through n8n's MCP support, which lets a flow expose or consume Model Context Protocol tools and slots an agent's tool-calling loop cleanly into a visual workflow. You get n8n's integration speed without trapping your core logic somewhere it cannot be tested, and you can migrate either layer independently later.

A decision checklist — five questions

Before choosing, answer these five. Who maintains this in a year — a developer, or whoever is at the desk? Non-developer maintenance points to n8n. How many third-party APIs does it touch? Three or more favours n8n's node library. What is the latency budget? Anything user-facing and sub-second says code. Does the logic need to be evaluated or unit-tested? If quality gates matter, the core belongs in code. Who needs to see the flow? If a client or manager must audit it visually, that is n8n's strongest single argument. Two or more answers pointing the same direction is usually a clear call; a split verdict is the hybrid pattern's cue.

What this looked like in practice

My Google Ads automation is the worked example. The orchestration layer — schedules, fetching account data through APIs, routing alerts — runs through n8n with MCP in the stack, because that layer is integration plumbing and benefits from being visible and editable without a deploy. The optimisation logic that decides what to change sits in code, where it can be tested and traced, because a system adjusting live ad spend is exactly the place I want evaluation gates rather than vibes. That split took weekly optimisation work from nine hours to two and helped drive an 18% ROAS improvement quarter over quarter — and the division of labour, not either tool alone, is what made it maintainable. I wrote up the build in the Google Ads automation with n8n tutorial if you want the full wiring.

Hammad Yousuf

AI Marketing Automation Engineer · Dubai, UAE

FAQ

Common questions

When should you use n8n instead of writing custom code?

When the workflow is integration-heavy — several third-party APIs, webhooks, schedules — and non-technical stakeholders need to see or edit the flow. n8n's prebuilt nodes, retry handling, and visual canvas save real build time there. Keep logic that needs testing, evaluation, or low latency in code.

What are the limitations of n8n for AI agents?

Weak testability (no natural unit tests for a node graph), version-control friction (JSON exports diff poorly), canvas readability collapsing as conditional logic grows, and per-execution cost at high volume on cloud plans. It is strong as an orchestration shell, weaker as the home for an agent's core reasoning.

Can you combine n8n with custom code?

Yes, and it is usually the best answer: n8n handles triggers, schedules, integrations, and retries, then calls a custom service via its HTTP node or MCP support for the reasoning-heavy or latency-sensitive core. Each layer can then be maintained and migrated independently.

Is n8n good enough for production AI agents?

For orchestration and integration, yes — I run it in production for exactly that. As the home for evaluation-critical reasoning logic, no: quality gates, traces, and tests belong in a codebase, so put the agent's core in code and let n8n coordinate around it.