Travel tips and city guides
Outbox Pattern, Idempotency and Retries: Reliable Integrations
Middleware and Integration Architecture ·
What is the outbox pattern and why does it save integrations?
The outbox pattern (transactional outbox) is a design pattern in which you write the message destined for other systems into an “outbox” table in the same transaction as the business record, and then publish that message safely from a separate process. Why bother? Because the “save the order, then put a message on the queue” approach leaves confirmed orders in the ERP with no invoice whenever the application crashes between the two steps. The reverse happens too: the message goes out, the transaction rolls back, and an e-invoice request is raised for an order that does not exist.
This guide is for software teams and technical leads who have to build reliable data flows between an ERP, an e-commerce platform, a payment provider and an e-invoicing service provider. Using PostgreSQL and TypeScript examples, it covers outbox table design, running the relay by polling or with change data capture (Debezium), at-least-once delivery, idempotency keys, retries with exponential backoff and jitter, timeouts, dead-letter queues, circuit breakers and message ordering. For the full architectural picture, we recommend reading it alongside our article on enterprise integration middleware architecture.
In short
- A dual write (writing to the database and to the broker separately) can leave the two systems inconsistent; the outbox reduces it to a single transaction.
- The relay publishes outbox rows by polling (for example with
FOR UPDATE SKIP LOCKED) or via CDC (Debezium). - The outbox gives you at-least-once delivery, so consumers must filter duplicates by message ID — in other words, they must be idempotent.
- Retries are designed together with timeouts, capped exponential backoff, jitter and a retry limit.
- Messages that can never be processed (poison messages) go to a DLQ, and the DLQ has an owner, an alert and a reprocessing procedure.
• • •
What is the dual-write problem?
As Chris Richardson defines it on microservices.io, the problem is this: a service command needs to update the database and send a message to a broker, and the two operations must be atomic. A distributed transaction (2PC) spanning the database and the broker is often not an option, or not a desirable one. Sending the message in the middle of the transaction is unreliable because there is no guarantee it will commit; sending it after the commit is unreliable too, because the service might crash at exactly that moment.
A concrete example: ERP order → e-invoice
When a customer order is approved, the ERP moves it to “confirmed” and an invoice request must go to
the e-invoicing provider. The application commits the order, then loses the connection while
sending the message. Result: the order is confirmed, the invoice never exists, and finance only
notices during month-end reconciliation. With an outbox, the order update and the
OrderConfirmed event are written in the same transaction: either both become durable or
neither does.
A concrete example: payment callbacks
A payment provider reports a successful payment through a webhook, and providers may send the same notification again when they do not receive a response. Your application has to treat that notification as “order paid” exactly once and trigger shipping and invoicing exactly once. Two patterns work together here: idempotent processing on the inbound side and an outbox on the outbound side.
• • •
How does the transactional outbox work?
According to the pattern description on microservices.io, the sending service inserts the message into an outbox table in its database as part of the transaction that updates the business entities; a separate message relay process then forwards those messages to the broker. The pattern promises three things: no 2PC, messages are sent if and only if the transaction commits, and messages are published in the order the service produced them. The same source also gives an important warning: if the relay publishes a message and crashes before recording that it did so, it will publish the message again when it restarts.
Designing the outbox table
Debezium’s Outbox Event Router documentation expects the columns id,
aggregatetype, aggregateid, type and payload in its
default configuration: aggregatetype determines the target topic,
aggregateid becomes the message key (and therefore drives ordering within a Kafka
partition), and id travels as a header that consumers can use to detect duplicates. If
you run a polling relay, a few extra columns for publication state and retry information make life
easier:
-- Outbox: same database as the business data, written in the same transaction
CREATE TABLE outbox (
id uuid PRIMARY KEY,
aggregatetype varchar(255) NOT NULL, -- e.g. 'order'
aggregateid varchar(255) NOT NULL, -- e.g. order number (message key)
type varchar(255) NOT NULL, -- e.g. 'OrderConfirmed'
payload jsonb NOT NULL,
created_at timestamptz NOT NULL DEFAULT now(),
published_at timestamptz, -- for the polling relay
attempts integer NOT NULL DEFAULT 0,
next_attempt_at timestamptz NOT NULL DEFAULT now(),
last_error text
);
CREATE INDEX outbox_pending_idx
ON outbox (next_attempt_at)
WHERE published_at IS NULL;
-- Application: order update and event in a single transaction
BEGIN;
UPDATE orders SET status = 'CONFIRMED' WHERE id = $1;
INSERT INTO outbox (id, aggregatetype, aggregateid, type, payload)
VALUES ($2, 'order', $1, 'OrderConfirmed', $3::jsonb);
COMMIT;
Relay: polling or CDC?
A polling relay reads pending rows at regular intervals and publishes them. The PostgreSQL
documentation explains that SKIP LOCKED skips any rows that cannot be locked
immediately; it is not suitable for general-purpose work, but it can be used to avoid lock
contention when multiple consumers read from a queue-like table. Polling is simple and needs no
extra infrastructure; the downsides are query load and latency equal to the polling interval.
-- Several relay instances can run without picking up the same row
SELECT id, aggregatetype, aggregateid, type, payload, attempts
FROM outbox
WHERE published_at IS NULL
AND next_attempt_at <= now()
ORDER BY created_at
LIMIT 100
FOR UPDATE SKIP LOCKED;
A CDC relay, by contrast, reads the database’s transaction log. In the Debezium project’s 2019 outbox article, Gunnar Morling explains that log-based CDC works with very low overhead and in near real time compared with polling. The same article contains a neat detail: even if a row is inserted and deleted within the same transaction, the INSERT event is still in the log and gets published, so the table stays empty and no separate housekeeping job is needed. The price is taking on the operation of Kafka Connect and Debezium.
• • •
Delivery guarantees: at most once, at least once, effectively once
The Confluent documentation defines three semantics. With at-most-once delivery, a message may be lost on failure but is never redelivered. With at-least-once delivery, messages are never lost but may be delivered more than once. Exactly-once means each message is delivered once and only once. Since version 0.11.0.0, Kafka supports exactly-once processing within Kafka via the idempotent producer and transactions; the same documentation stresses, however, that this coordination becomes hard when writing to an external system, and that Kafka provides at-least-once delivery by default.
In practice, across a chain that runs from an ERP to an e-invoicing provider, the realistic target is an effectively-once effect: a message may arrive more than once, but the business outcome happens once. What delivers that is not a broker setting but the idempotent design of the consumer.
| Semantics | What it guarantees | Risk | Cost / complexity | Good fit |
|---|---|---|---|---|
| At most once | No duplicates | Message loss | Low | Metrics, logs, notifications where loss is tolerable |
| At least once | No loss | Duplicate messages | Medium: acks, retries, durable queues | Order, invoice and stock events (with an idempotent consumer) |
| Effectively once | No loss, single business effect | Silent duplicates if the design is wrong | High: outbox, message IDs, dedup table | Payments, e-invoices, accounting entries |
• • •
How do you design idempotency keys and a dedup table?
The Idempotent Consumer pattern on microservices.io recommends that a consumer record the IDs
of the messages it has processed in its database. When processing a message, the consumer starts a
transaction and inserts the message ID into a table whose primary key is
(subscriberId, messageId); if the message was already processed, the insert fails and
the message is ignored. RabbitMQ’s reliability guide likewise notes that messages can be redelivered
after network or node failures, and recommends that consumers handle this through idempotent
design.
CREATE TABLE processed_messages (
consumer_id text NOT NULL, -- e.g. 'einvoice-adapter'
message_id uuid NOT NULL, -- outbox.id (from the message header)
processed_at timestamptz NOT NULL DEFAULT now(),
PRIMARY KEY (consumer_id, message_id)
);
Idempotency on outbound calls
Stripe’s API documentation is a good reference. The client adds an Idempotency-Key
header to each write request; the server saves the status code and body of the first request,
whether it succeeded or failed (500 errors included), and later requests with the same key get the
same result. Stripe suggests values with enough entropy, such as V4 UUIDs, limits keys to 255
characters, notes that keys can be pruned once they are at least 24 hours old, and returns an error
if the same key is reused with different parameters. The same page warns against putting sensitive
data such as email addresses in keys.
Not every e-invoicing provider or bank API offers such a header. If yours does not, you have two options: build a “query first, send only if absent” flow keyed on a unique ID you assign to the invoice, or record each outbound request and its result in a local table and check it before sending again. Which route is possible should be verified against the provider’s documentation during discovery. For how recent e-invoice format changes affect integrations, see our article on the GİB UBL-TR update.
• • •
How should you tune retries, timeouts and jitter?
The Amazon Builders’ Library article “Timeouts, retries, and backoff with jitter” is one of the most solid references on the subject. Its core advice: set both a connection timeout and a request timeout on every remote call; treat retries as “selfish”, because they can add load to a system that is already overloaded; use capped exponential backoff; and limit the number of retries.
The layered-retry trap
The same article gives a striking example: in a five-layer call chain where each layer retries three times, the load on the database can multiply by 243. So retry at a single point in the chain. In an outbox architecture that point is usually the relay and the consumer — and do not leave your HTTP client library’s own hidden retries switched on as well.
Jitter: don’t all come back at once
Marc Brooker’s 2015 post on the AWS Architecture Blog compares “Full Jitter”, “Equal Jitter” and “Decorrelated Jitter” in simulation: exponential backoff without jitter is the clear loser, taking both more work and more time, while Full Jitter makes fewer calls. The rule is simple: pick the wait at random between zero and that attempt’s upper bound.
Which errors should be retried?
The AWS article points out that APIs with side effects are not safe to retry unless they are idempotent, and that client errors (4xx) should not be retried with the same request, while server errors may succeed on a later attempt. Timeouts, connection failures and transient server errors are retried; validation and authorization errors are not — they go straight to the DLQ and to human review.
// Relay worker (pseudo-code): publish, mark; on failure, defer with jitter
const BASE_MS = 1_000
const CAP_MS = 30_000
const MAX_ATTEMPTS = 8
function backoffWithJitter(attempt: number): number {
const ceiling = Math.min(CAP_MS, BASE_MS * 2 ** attempt)
return Math.floor(Math.random() * ceiling) // Full Jitter
}
async function relayBatch(db: Db, broker: Broker): Promise<void> {
await db.transaction(async (tx) => {
const rows = await tx.query(SELECT_PENDING_SQL) // FOR UPDATE SKIP LOCKED
for (const row of rows) {
try {
await broker.publish(row.aggregatetype, {
key: row.aggregateid, // same order → same partition
headers: { 'message-id': row.id }, // consumers dedupe on this
value: row.payload,
}, { timeoutMs: 5_000 })
await tx.query('UPDATE outbox SET published_at = now() WHERE id = $1', [row.id])
} catch (err) {
const attempts = row.attempts + 1
if (attempts >= MAX_ATTEMPTS) {
await moveToDeadLetter(tx, row, err) // alert + human review
continue
}
await tx.query(
`UPDATE outbox
SET attempts = $2,
next_attempt_at = now() + ($3 || ' milliseconds')::interval,
last_error = $4
WHERE id = $1`,
[row.id, attempts, backoffWithJitter(attempts), String(err)],
)
}
}
})
}
• • •
Dead-letter queues, poison messages and circuit breakers
A poison message is one that cannot be processed no matter how many times you try: malformed JSON, a product code with no mapping, a customer account with an invalid tax number. Retrying such a message forever clogs the queue and holds up the healthy messages behind it. RabbitMQ supports routing a message to a dead-letter exchange when it is rejected without requeueing, when its TTL expires, when the queue length limit is exceeded, or, on quorum queues, when the delivery limit is exceeded.
Rules for running a DLQ
- Every message in the DLQ is stored with its original message ID, failure reason, attempt count and correlation ID.
- DLQ depth has an alert, and that alert has an owner (a person or an on-call rotation).
- Once the cause is fixed, reprocessing is safe because consumers are idempotent; the procedure is written in the runbook.
- Messages do not sit in the DLQ forever; retention and any personal data they contain are reviewed against data-protection rules such as KVKK.
Circuit breakers
The circuit breaker Martin Fowler described in 2014 wraps a protected remote call in an object; once failures pass a threshold, the breaker “trips” open and further calls return an error without reaching the remote system at all. After a while it moves to a “half-open” state, and if a trial call succeeds the breaker closes again. Fowler also recommends logging and monitoring state changes. The AWS article, for its part, notes that circuit breakers can introduce modal behaviour that is hard to test, and that limiting retries locally with a token bucket is an alternative. The most tangible benefit: when the e-invoicing provider is down for maintenance, the relay waits instead of dumping every invoice into the DLQ.
• • •
Ordering, sagas and the consumer side
When does ordering matter?
If “order cancelled” is processed before “order created”, things break. Debezium uses the
aggregateid of the outbox row as the message key, and Kafka writes events with the same
key to the same partition, preserving their write order within that partition. If you run a polling
relay with several instances and SKIP LOCKED, two events for the same order can be
published by different instances in a different order. When strict ordering is required, choose a
single relay instance, relays partitioned by key, or CDC — and let the consumer ignore stale events
by checking a version number on the event.
Outbox together with sagas
For business processes that span several services, microservices.io recommends the saga pattern: a sequence of local transactions, each updating its own database and triggering the next step. If a step fails because of a business rule, the previous steps are undone with compensating transactions. For a saga to work reliably, every step must update its database and publish its message atomically — which is why there is usually an outbox underneath a saga.
Consumer pseudo-code
// e-invoice adapter: does nothing if the same message arrives a second time
async function handleOrderConfirmed(msg: Message): Promise<void> {
await db.transaction(async (tx) => {
const res = await tx.query(
`INSERT INTO processed_messages (consumer_id, message_id)
VALUES ('einvoice-adapter', $1)
ON CONFLICT DO NOTHING`,
[msg.headers['message-id']],
)
if (res.rowCount === 0) return // duplicate: already processed
const invoice = mapOrderToInvoice(msg.value) // canonical model → invoice
await tx.query(
'INSERT INTO invoice_requests (order_id, payload, status) VALUES ($1, $2, $3)',
[invoice.orderId, invoice, 'PENDING'],
)
})
// Sending to the provider happens in a separate worker, driven by
// invoice_requests and the idempotency mechanism the provider supports.
}
• • •
Checklist: outbox and retry design
- Are the business data and the outbox row written in the same database, in the same transaction?
- Does every message have a unique ID that travels in a broker header?
- Does every consumer filter duplicates by that ID (dedup table or natural key)?
- Does every remote call have a connection timeout and a request timeout?
- Are retries done in a single layer only, with capped exponential backoff and jitter?
- Do validation and authorization errors skip retries and go straight to the DLQ?
- Are the DLQ alert, owner and reprocessing procedure written down in the runbook?
- Is there a key and relay strategy for events that need ordering?
- Are you tracking the number of pending outbox rows and the age of the oldest one?
You will find how to collect these metrics in our article on observability in enterprise applications, and what to watch on the PostgreSQL side as the outbox table grows in our guide to PostgreSQL read-heavy workloads.
How can Aksiyon Soft help?
In our API and integration projects, the outbox, idempotent consumers, retry policy and DLQ processes are part of the design from the first sprint. If your existing integrations suffer from duplicate invoices or lost orders, we first measure the flow and start from the riskiest point. For monitoring live systems, watching the DLQ and incident response, our maintenance and support service takes over. You can explore the overall architecture on our API and data integration platform solution page.
How we work: discovery maps the flows and failure scenarios, an MVP is built around the first critical flow, a working version is demoed every sprint, and go-live is followed by hypercare and SLA-backed maintenance. We are headquartered in Samsun and work remotely with teams across Türkiye, with planned on-site visits when needed.
Frequently asked questions
Do I need Kafka for the outbox pattern?
No. The outbox is broker-agnostic; the relay can forward messages to RabbitMQ, to Kafka or to an HTTP endpoint. If you want CDC with Debezium, the Kafka Connect ecosystem is a natural choice, but a polling relay works with any broker.
Won’t the outbox table grow and hurt performance?
With polling, published rows should be deleted or archived regularly, and a partial index on pending rows keeps the query fast. With CDC, rows can be inserted and deleted immediately; as Debezium’s outbox article describes, the table stays practically empty.
Is exactly-once impossible?
Kafka supports exactly-once processing between its own topics using transactions. But when the chain includes external systems such as an e-invoicing provider, a bank or an email service, the goal is an effectively-once effect: a message may arrive more than once, and the idempotent consumer ensures the business outcome happens once.
How many times should we retry?
There is no universal number; it depends on how long the other system takes to recover and how urgent the work is. What matters is a bounded number of attempts, capped exponential backoff, jitter, and a DLQ plus alert once the limit is reached. The values are written down during discovery and tuned in production based on metrics.
How should payment webhook notifications be handled?
Acknowledge the notification quickly and lightly, persist it, and do the real work asynchronously. Use the provider’s notification ID in a dedup table so the same notification is not processed twice, and write the “order paid” update and the shipping/invoicing event to the outbox in the same transaction.
Does adding an outbox to an existing system mean a big rewrite?
Usually not. Starting with the riskiest flow, it is enough to change the code that sends the message so that it writes to the outbox, and to add a relay. Other flows are migrated gradually, and adding a dedup table to consumers is often a small change.
Let’s talk about your project
If you are dealing with duplicate invoices, orders that never reach the ERP, or error queues nobody looks at, let’s review your flow together in a short discovery call. Just use the contact form to tell us which systems are connected and where the problem shows up.
Sources
- microservices.io (Chris Richardson) — Pattern: Transactional outbox
- microservices.io (Chris Richardson) — Pattern: Idempotent Consumer
- microservices.io (Chris Richardson) — Pattern: Saga
- Debezium Documentation — Outbox Event Router
- Debezium Blog — Gunnar Morling: Reliable Microservices Data Exchange With the Outbox Pattern (19 February 2019)
- Amazon Builders’ Library — Timeouts, retries, and backoff with jitter
- AWS Architecture Blog — Marc Brooker: Exponential Backoff And Jitter (4 March 2015)
- Stripe API Reference — Idempotent requests
- Confluent Documentation — Kafka Message Delivery Guarantees
- PostgreSQL Documentation — SELECT (locking clause, SKIP LOCKED)
- martinfowler.com — Martin Fowler: CircuitBreaker (6 March 2014)
- RabbitMQ — Reliability Guide
- RabbitMQ — Dead Letter Exchanges
Related posts
Software Buyer Guides
Gaziantep Software Partner: Export ERP, e-Invoicing and B2B Portals
A guide to Gaziantep software needs for textile, carpet and food exporters: export ERP, e-invoice and customs integration, multi-plant production and B2B dealer portals.
Software Buyer Guides
Malatya Software Partner: Apricot Exports, Traceability and Business Continuity
How Malatya software projects can support apricot processing and exports, OIZ textiles and post-earthquake rebuilding: traceability, export documents, cloud backups and business continuity.
