System integration connects separate software systems so data and processes flow between them automatically and stay consistent. The link can be an API call, an event, a file or shared middleware. The hard part is deciding which system owns each piece of data, and what happens when a message is lost, repeated or late.
This guide compares the main approaches (point-to-point, hub-and-spoke, enterprise service bus, iPaaS, event-driven and API-led) and the four classic integration styles, draws the line between application integration, data integration and process automation, then works through design rules, a worked orders-to-accounting example, failure modes and a starter plan. What is a webhook and dead letter queues go deeper on two of the mechanisms.
What counts as system integration?
Gregor Hohpe and Bobby Woolf, authors of the 2003 book Enterprise Integration Patterns, define enterprise integration as "the task of making separate applications work together to produce a unified set of functionality." Three kinds of work fit that description, and each calls for different tools.
| Kind | What it connects | Typical tools | Reach for it when |
|---|---|---|---|
| Application integration | Applications, at the level of functions and events | APIs, webhooks, middleware, iPaaS | One system must react to another, such as a paid order creating an invoice |
| Data integration | Many sources into one store or view | ETL, ELT, change data capture | You need one view for analysis, such as sales from three systems in a warehouse |
| Process automation (RPA) | Applications, through their user interface | RPA software | A system has no API and no budget for a deeper integration |
IBM draws the line on timing. Application integration "directly links multiple applications at a functional level" and can link data "in near real-time". Data integration is "batch-based", "isn't a real-time process" and is "commonly used after processes have been completed". API integration is a technique inside the first kind: IBM lists APIs, middleware and webhooks among its technologies.
Data integration has its own vocabulary. AWS describes ETL as moving data from source to destination "at periodic intervals": extract, transform, load. ELT loads the data into the target first and processes it there. Extraction can follow update notifications from the source, check periodically for changed records, or reload everything, which AWS recommends only for small tables. Change data capture (CDC) is the notification route for databases: Debezium is an open source platform that lets applications respond to the inserts, updates and deletes that other applications commit. For batch loads, CDC and streams into analytics, see building data analytics software; for landing the data in layers, see medallion architecture.
IBM describes robotic process automation as scripts that "emulate human processes", working "on the presentation layer of existing applications", which suits "situations where you don't have an application programming interface (API)". A bot depends on the screen rather than on a contract, so treat it as a bridge until an API exists.
How do the main integration approaches compare?
Start with the arithmetic. If every one of n systems must exchange data with every other, direct links number n(n-1)/2: 3 systems need 3 links, 6 need 15, 10 need 45 and 20 need 190. IBM calls point-to-point "the most straightforward integration strategy" and "relatively cheap and simple" to implement, but warns that "the larger the network of apps and processes, the more point-to-point integrations teams will have to configure and maintain." That makes it "best suited for small-scale integration projects." Each link also carries its own data-format translation. Enterprise Integration Patterns counts the translators needed if every application sends to and receives from every other: six applications need 30 when each pair is translated directly and 12 when each translates to and from one shared format, a canonical data model.

A middle layer removes the tangle and moves the risk. The table condenses the sources linked after it and our own practice; the last column is our judgement.
| Approach | Coupling | Timing | Who runs it | If it fails | Cost profile | Usually fits |
|---|---|---|---|---|---|---|
| Point-to-point | Tight, link by link | Immediate for calls, on a schedule for files | Whoever built each link | One pair stops, and each link needs its own fix | Cheap and simple at first, then every system adds links | A few systems with stable interfaces |
| Hub-and-spoke | Systems know only the hub | Near real time through the hub | A central team | The hub is a single point of failure | One platform to fund, "prohibitively costly and complicated to maintain" at scale | A mid-sized estate with a team to run the hub |
| Enterprise service bus | Systems know the bus | Near real time through the bus | A central team | One change can destabilize other integrations | A central platform with its own high availability to fund | Large estates with many legacy systems |
| iPaaS | Through the vendor's connectors | Set by each connector's trigger, often a schedule | A business or IT owner per flow | Depends on the platform's retries and alerts | A subscription that grows with runs or tasks | SaaS-to-SaaS flows of modest volume |
| Event-driven broker | Loose, through events | Asynchronous, as fast as the broker delivers | Producer and consumer teams, plus whoever runs the broker | Lag, duplicates and out-of-order events | Broker operations and skills | Several consumers of the same events, or spiky load |
| API-led layers | Through published API contracts | Synchronous API calls | Each system's API owner | A failed system API stops every chain that calls it | API design and management effort | Many consumers that need reusable access to core systems |
IBM reports that in many organizations the ESB "came to be seen as the bottleneck": it was centrally managed, so "application teams soon found themselves waiting in line for their integrations", and a change to one integration "could destabilize others". On its application integration overview, which also supplies the hub-and-spoke row above (a hub gives "a single point of monitoring and control" and is itself a single point of failure), IBM goes further: with the rise of cloud-native ecosystems, "ESB tools are becoming obsolete as integration tools". That is a vendor's position, but the lesson holds either way: do not put every change behind one team.
An event broker removes the central translator. In Microsoft's description of the style, "there are no point-to-point integrations" and new consumers "can be added without modifying producers or other consumers". The price is eventual consistency, ordering and duplicate handling, and Microsoft says the style might not suit simple request-response workloads, transactions that need strong consistency, or teams without experience operating distributed asynchronous systems. MuleSoft's API-led approach puts an experience layer, a process layer and a system layer in a row, and the system layer's APIs "provide a way of insulation" from the backends.
None of these has to be the only one. We pick the pattern flow by flow, on volume, latency and who will maintain it.
What are the four classic integration styles, and what do they look like today?
Hohpe and Woolf sum up the options in four styles: file transfer, shared database, remote procedure invocation and messaging. They warn against using one for everything: "The trick is not to choose the one style to use always, but to choose the best style for a particular integration opportunity." Two applications may integrate using multiple styles, each point of integration taking the style that suits it best.
| Style | Today it looks like | Trade-off |
|---|---|---|
| File transfer | A scheduled export dropped on SFTP or object storage as CSV, JSON or XML | Files are produced "at regular intervals", so data is only as fresh as the last file |
| Shared database | One schema that several applications read and write, or database links between systems | One copy keeps the data consistent, but a unified schema "is a very difficult exercise" with "severe political difficulties" |
| Remote procedure invocation | REST, GraphQL, gRPC or SOAP calls | The remote calls tie systems "into a growing knot", and sequencing makes independent change hard |
| Messaging | Queues, brokers and streams, including webhooks that land in a queue | "Sending a message does not require both systems to be up and ready at the same time" |
Timing is the difference you feel first. Hohpe and Woolf note that "Latency in data sharing has to be factored into the integration design; the longer sharing can take, the more opportunity for shared data to become stale, and the more complex integration becomes." A file drop moves everything at once, when the schedule says. A message moves one record when it changes. We ask what latency each flow needs, because many real-time requests turn out to mean within a few minutes, and the answer changes what the flow costs to build and run.

When is an iPaaS the right tool, and what are its limits?
Gartner defines an iPaaS as "a vendor-managed cloud service that enables users to implement integrations between applications, services and data sources, both internal and external to their organization." The label is wide: as of October 2026 the same Gartner Peer Insights category lists enterprise platforms such as MuleSoft Anypoint Platform, Boomi, Workato and SAP Integration Suite next to Power Automate, Azure Logic Apps and Zapier. Gartner's mandatory features make a good checklist; among them are role-based access control, tooling for versioning, testing and deployment, and tools for monitoring, alerting, reporting and auditing integrations in production.
Cost is where the types differ. As of October 2026, Zapier counts a task for every successful action step, and does not count triggers, filters or failed steps. Make bills in credits, and one operation in a non-AI app is one credit. Software tools for starting an online business works through the arithmetic for both, so price a month of your real volume before you commit to a plan.
n8n can be self-hosted on your own infrastructure, which helps when data must stay in your environment. Its Sustainable Use License is "fair-code" rather than open source: free for your own internal business purposes, but not for making n8n available to your customers so that they can connect their accounts and build workflows.
We use an iPaaS recipe when volume is low, the mapping is simple and a business owner wants to change the flow; when volume is high, the mapping is complex or reliability is strict, custom services fit better. The watch-fors are per-task pricing, thin retry logic and no code review. Before committing, ask whether you can see a failed run and replay it, whether flows can be exported into version control, whose name holds the licence and credentials, and what leaving would cost. The tooling itself has a price: Hohpe and Woolf note that special integration tools "can be expensive, can lead to vendor lock-in, and increase the burden on developers", and IBM notes that iPaaS deployment "can require a lot of time and forethought", especially in large environments.
Which system owns each piece of data?
Before choosing a pattern, write one table: entity, system of record, copies and direction. A system of record owns the authoritative version of an entity, and every other system holds a copy that it never edits. Ownership goes down to fields. We agree it field by field during mapping, and when two systems disagree the owner's value wins and the disagreement is logged for someone to review.
Connectors force the same decision. In HubSpot's Salesforce sync each mapped property gets a rule: prefer Salesforce unless blank, always use Salesforce, two-way, where "the most recent value will always overwrite any existing values", or don't sync. HubSpot also warns that duplicate mappings "can update multiple Salesforce fields, overwrite values, trigger an unintended mapping, or create sync loops." Use two-way sync only for fields that both sides may legitimately edit.
Four rules keep a design from decaying:
- Join on identifiers you control. Keep a cross-reference table with one row per entity and each system's ID for it. Email addresses and names change and collide.
- Go direct for three systems, canonical after that. In the translator counts above, a direct mapping and a canonical data model tie at three applications (6 translators either way), and the shared format wins from there (30 against 12 at six). Each new application then needs only a translation to and from it.
- Decide what an event carries. Microsoft describes two options. A payload with every attribute saves consumers a lookup but can cause "data consistency problems because of multiple systems of record". A payload with only keys gives "better data consistency because it has a single system of record" but "can have worse performance", because consumers must query the source.
- Prefer one-way flows from the owner to its copies, and add a reverse flow only for a field with a second legitimate editor.
How do you handle duplicates, retries and ordering?
Assume every message can arrive twice, late or out of order, because providers say it can. Shopify "doesn't guarantee ordering within a topic, or across different topics for the same resource", so an update can arrive before the create, and it recommends sorting by the X-Shopify-Triggered-At header or the payload's updated_at. It also warns that your app might receive the same webhook more than once. Amazon SQS standard queues deliver "at least once" and can deliver the same message more than once; when processing must be exactly once and in order, AWS points to SQS FIFO queues.
Idempotency makes the repeat harmless. Where a provider offers idempotency keys, send one with every write. Stripe saves the result of the first request for a key and returns it for repeats, "including 500 errors". Keys can be up to 255 characters; once a key is at least 24 hours old Stripe can remove it, and a reuse after that generates a new request. Xero keeps keys for 6 minutes and answers 400 to a key over 128 characters. Your own receiver needs the same discipline, keyed on the provider's event ID (for Shopify, the X-Shopify-Webhook-Id header):
-- One row per event; the primary key is the duplicate check.
CREATE TABLE event_log (
event_id text PRIMARY KEY,
status text NOT NULL DEFAULT 'processing',
attempts int NOT NULL DEFAULT 1,
updated_at timestamptz NOT NULL DEFAULT now()
);
-- Claim an event. A row comes back only if this worker should process it:
-- the ID is new, or an earlier attempt failed or stalled for ten minutes.
INSERT INTO event_log (event_id) VALUES ($1)
ON CONFLICT (event_id) DO UPDATE
SET status = 'processing', attempts = event_log.attempts + 1, updated_at = now()
WHERE event_log.status = 'failed'
OR (event_log.status = 'processing'
AND event_log.updated_at < now() - interval '10 minutes')
RETURNING attempts;
After the work succeeds, set status = 'done'; after it fails, set status = 'failed'. No row back means a duplicate: acknowledge it and stop. The writes the worker makes still need to be idempotent, because a worker that looked stalled can wake up and finish.
Tip
Provider idempotency windows are short (6 minutes at Xero, at least 24 hours at Stripe), and a nightly retry or a replay from a dead-letter queue falls outside both. Your own event table is what makes a late replay safe.
Retries need backoff and jitter. Capped exponential backoff spaces retries out, but clients that failed together still retry together. In AWS's simulation of 100 contending clients, "The solution isn't to remove backoff. It's to add jitter": randomizing the wait cut the number of calls by more than half and improved time to completion, compared with un-jittered backoff. When a provider says how long to wait, use that instead: when a call exceeds Xero's minute or daily limit, the 429 response carries a Retry-After header, and Xero says to pause requests to that tenant until then. In code:
import random
import time
def send_with_retries(send, attempts=6, base=1.0, cap=60.0):
for attempt in range(attempts):
try:
return send()
except RateLimited as error: # HTTP 429: the provider says how long
delay = error.retry_after
except TransientError: # timeout or 5xx, never a validation error
delay = random.uniform(0, min(cap, base * 2**attempt))
if attempt < attempts - 1:
time.sleep(delay)
raise DeadLetter("retries exhausted") # park it for a person
When your own application saves a record and must tell another system, that is two writes, and either can fail. AWS's transactional outbox pattern resolves this dual write: write the event to an outbox table in the same database transaction as the change, and let a separate process publish it. Consumers must still be idempotent, because the message can be delivered more than once, and AWS names change data capture as the alternative to an outbox table. Messages that still fail after their retries go to a dead-letter queue with an alert, and are replayed after the fix (see dead letter queues).
How do you secure, describe and monitor an integration?
Authentication comes first. Server-to-server calls usually use the OAuth 2.0 client credentials grant. RFC 6749 says it "MUST only be used by confidential clients" (clients that can keep their credentials secret) and that a refresh token "SHOULD NOT" be included. Ask for the narrowest scope: RFC 9700 says an access token's privileges "SHOULD be restricted to the minimum required". For partner links that need stronger proof of the caller, RFC 8705 defines mutual TLS client authentication and certificate-bound access tokens, and RFC 9700 says servers "SHOULD use mechanisms for sender-constraining access tokens", naming mutual TLS and DPoP.
Distrust partner data. OWASP observes that "Developers tend to trust data received from third-party APIs more than user input", even for well-known companies, so validate every inbound payload against its schema. Verify webhook signatures as well: Shopify says to "Always verify HMAC before trusting payload contents".
Write the contract down. OpenAPI is a "standard, programming language-agnostic interface description for HTTP APIs" (version 3.2.1 is dated 10 September 2026), and AsyncAPI does the same for message-driven APIs: it is "protocol-agnostic, so you can use it for APIs that work over any protocol", with AMQP, MQTT, WebSockets and Kafka among the examples. Expect partners to change theirs: Shopify releases a new API version every three months and supports each stable version for at least 12 months. For versioning, security and observability on the API side, see enterprise API development in 2026.
Make every record traceable. Put one correlation ID on a record at the first system and carry it through every hop. In MuleSoft's layered example, "For each layer, the Correlation ID doesn't change", which lets you correlate different log entries with one execution. For HTTP calls, the W3C Trace Context Recommendation (23 November 2021) standardizes the traceparent header for carrying trace context between services. Alert on the age of the oldest unprocessed message in each queue, not only on error counts, because a stalled consumer can leave a queue growing without raising an error.
A worked example: orders to invoices to the CRM
This example is illustrative, not a client. An online store on Shopify, accounting in Xero and a CRM, at about 150 paid orders a day. Swap Xero for an ERP and the shape is the same. Provider limits below are as published in October 2026.
| Entity | System of record | Copies | How it moves |
|---|---|---|---|
| Order, lines and payment status | Store | Accounting, CRM | Order webhook into a queue, two consumers |
| Invoice and its payment status | Accounting | CRM, status only | Nightly pull of changed invoices |
| Customer contact details and consent | CRM | Store checkout, accounting contact | Created from the first order, then edited only in the CRM |
The limits shape the design:
- The endpoint only receives. Shopify allows a one-second connection timeout and a five-second timeout for the entire request and treats any response outside the 200 range as an error. The handler verifies the HMAC signature, stores the payload with its
X-Shopify-Webhook-Idand returns 200 (see what is a webhook); workers do the rest. For bursts, Shopify recommends queueing payloads. - A broken endpoint loses its subscription. If Shopify gets no response or an error, it retries 8 times over the next 4 hours, and after 8 consecutive failures it deletes a subscription created through the Admin API. Shopify emails the app's emergency developer address, which is not a monitoring system. Alert on silence instead, an unusual gap with no events, and check that the subscription still exists.
- Accounting sets the pace. Per tenant, Xero allows 5 concurrent calls, 60 calls a minute and 1,000 calls a day on the starter tier (5,000 on higher tiers), and suggests a practical ceiling of about 50 nodes per request to stay under the 3.5 MB maximum size. A flash sale of 600 orders in ten minutes would use the whole minute allowance if each order were posted alone; batched at 50, it needs 12 calls. Xero is designed for volumes of up to 5,000 sales invoices a month, so 150 orders a day (about 4,500 a month) fits, but ten times that would need a different design, such as one summary invoice a day, which is a bookkeeping decision.
- Retries outlive the idempotency window. Xero keeps idempotency keys for 6 minutes, so a retry after a longer outage cannot rely on them. After repeated errors, Xero's own advice is to make a GET request to check whether the invoice already exists before creating it with a new key.
Warning
Do not rely on webhooks alone. Shopify says an app "shouldn't rely on receiving data from Shopify webhooks", because delivery "isn't always guaranteed", and recommends reconciliation jobs that periodically fetch data. A deleted subscription stops events without an error in your own logs.
A nightly job runs the reconciliation. It pulls the invoices changed since the last run (incremental extraction, in AWS's terms), compares them with yesterday's paid orders through the cross-reference table, and writes a report. The core query finds orders with no invoice:
SELECT o.order_id, o.total_paid, o.paid_at
FROM orders o
LEFT JOIN xref x ON x.entity = 'order' AND x.store_id = o.order_id
WHERE o.paid_at >= current_date - 1
AND o.paid_at < current_date
AND x.ledger_id IS NULL;
An example report, with illustrative numbers:
| Check | Store | Accounting | Result |
|---|---|---|---|
| Paid orders against invoices | 148 | 147 | 1 missing: replay it from the dead-letter queue |
| Value of paid orders against invoices | 18,420.50 | 18,265.50 | 155.00 apart, the missing order |
| Refunds against credit notes | 3 | 3 | Match |

The same shape holds beyond retail. In the booking and payment platform we built for a notary office, payment events are claimed once and reconciled on a schedule, and emails and calendar updates are sent from an outbox.
What breaks in production, and what prevents it?
| What goes wrong | How it shows up | What prevents it |
|---|---|---|
| Duplicate records | Two invoices for one order, or a customer created twice | Receivers keyed on the provider's event ID, idempotency keys on writes, a cross-reference table |
| Silent connector failure | Nothing arrives and nothing errors, because a subscription was deleted or a credential expired | Alerts on silence, scheduled checks of subscriptions and credentials, nightly reconciliation |
| Mapping drift | A field is renamed or reformatted and records land wrong | A pinned API version, a check of the version header, contract tests on real samples |
| Rate limits | HTTP 429 responses, stalled jobs, a nightly batch that runs into the morning | Batched writes, honouring Retry-After, staggered schedules, a queue in front |
| Out-of-order events | An older update overwrites a newer one | Comparing updated_at or a version number before applying, FIFO queues where order matters |
| Partner outage | Timeouts and a growing queue | Timeouts, backoff with jitter, a durable queue, a dead-letter queue, an alert on queue age |
Two of these leave no trace in a log. Shopify falls forward when your app targets a version that has been retired, and responds with the oldest accessible stable version. The X-Shopify-API-Version header shows which version answered, so compare it with the version you pinned. And a deleted webhook subscription looks like a quiet day, which is why the nightly reconciliation counts records instead of waiting for errors.
An integration plan you can start this month
- Inventory the flows. List every place a person copies data between systems: source, target, volume per day, how quickly it must arrive and who owns each end.
- Rank by pain and volume. Start with the flow that costs the most hours or causes the most errors, and that has a named owner.
- Name the system of record for each entity in that flow, and write the conflict rule for any shared field.
- Choose the pattern for the flow. Use a direct API call when one system needs an answer now, a webhook into a queue when events must not be lost, a scheduled file or database sync when a system has no API, and an iPaaS recipe for low-volume flows that a business owner will maintain.
- Pilot one flow end to end with idempotency, retries, a dead-letter queue, an alert on queue age and a nightly reconciliation report. Run it beside the manual process until the report is clean.
- Write the runbook: who gets each alert, how to replay a failed message, how to rotate credentials and what to do when a partner is down.
- Scale to the next flow with the same skeleton, and review the inventory every quarter, because partners change their APIs: Shopify ships a new version every three months.
Our integrations service follows the same order: map the data and who owns it, design the pattern and the failure handling, build against real samples, then run with monitoring and reconciliation. If a flow needs custom code, the software development process explains how that work is planned and delivered, including when to build and when to buy.


