What Happens When the Server Saves but the App Never Knows?

What Happens When the Server Saves but the App Never Knows?

The problem: one tap, two possible realities

request retries and ambiguous outcomes

A user submits feedback. The backend saves it successfully, but the response never reaches the app. The database contains the feedback. The user sees an error.

From the client's perspective, the operation failed. From the server's perspective, it succeeded. Both are behaving consistently with the information available to them, yet the system has reached a state where the user and the database disagree.

This is not an unusual edge case. Mobile networks switch between Wi-Fi and cellular data, connections disappear mid-request, and responses get lost even when the server has completed its work. Retrying the request is the natural response. It is also where a simple feedback form starts becoming a distributed-systems problem.

If the backend treats the retry as a new submission, the same feedback gets stored twice. If it rejects the retry because the record already exists, the user receives an error for an operation that actually succeeded. Either way, the system has failed to preserve the meaning of a single action.

There is a second problem behind the same button. Once feedback has been saved, the team needs to hear about it, perhaps through an email and a Slack notification. Saving the record and sending those notifications are separate operations. The database can commit successfully while an email provider is unavailable, or a notification can be sent before a subsequent database operation fails.

The question, then, is not simply how to make a Send button reliable. It is how to preserve one logical operation across a client, an unreliable network, a database and multiple downstream services, each with its own failure modes.

The client creates the identity

The first decision is to distinguish a user action from the HTTP request that carries it.

Every time the user initiates a new feedback submission, the app generates a unique identifier for that submission. The identifier remains unchanged across automatic retries and manual retries. It is cleared only when the submission reaches a successful outcome, and the next submission receives a new identifier.

This identifier is an idempotency key.

POST /feedback
Idempotency-Key: 8f4a7c2e-...
Content-Type: application/json

The important property is not the format of the key. It is its lifecycle. If a network failure causes the same action to produce five HTTP requests, all five requests must carry the same key. Generating a new key for every attempt would make every retry appear to be a different submission.

The backend associates the key with the authenticated user and uses that identity to locate the original operation. When a request arrives, the server either processes a new submission or returns the result of an existing one.

A repeated request should not be treated as an error simply because the record already exists. If the original submission was successful, the retry should receive a successful response representing that same submission. The client can then show the thank-you screen without knowing whether it was the first request or the fifth.

This is the basic idea behind idempotent request handling: repeated attempts produce the same logical outcome without repeating the underlying operation.

The existence check that is not enough

There is a concurrency problem hidden inside even a correctly designed idempotency check.

Consider two identical requests arriving almost simultaneously. Both check the database for the idempotency key. Neither finds a record, so both proceed to create one.

The problem is that checking whether a record exists and creating it are two separate operations. Under concurrency, both requests can pass the check before either one commits.

An application-level condition does not establish uniqueness. The database must enforce it.

The feedback record therefore needs a uniqueness constraint based on the submission's identity, scoped appropriately to the authenticated user. The server attempts an atomic insert. If another request has already created the record, the database rejects the duplicate through its uniqueness constraint, and the application retrieves the existing result instead.

Conceptually, the operation looks like this:

async def submit_feedback(user_id, key, content):
    existing = await repository.find_by_key(user_id, key)

    if existing:
        return existing

    try:
        return await repository.create_if_absent(
            user_id=user_id,
            idempotency_key=key,
            content=content,
        )
    except DuplicateSubmission:
        return await repository.find_by_key(user_id, key)

This is illustrative pseudocode. The repository's insert must be atomic, and the real implementation must handle concurrent transactions, in-progress requests and database errors explicitly. A uniqueness conflict must not be confused with an unrelated persistence failure.

The distinction matters because the goal is not merely to avoid duplicates during normal operation. It is to make duplicate prevention hold when requests race, workers restart or the network causes the client to retry at exactly the wrong moment.

But even after the feedback record is safe, there is still another operation to account for.

Saving the record is not the same as notifying the team

atomic writes and asynchronous event processing

A naive implementation saves the feedback and then sends an email.

feedback = await repository.create(content)
await email_service.notify_team(feedback)

The sequence looks reasonable until the email service fails. The feedback has already been committed, but the notification has not been delivered. Retrying the entire operation can create another feedback record unless idempotency is enforced independently.

Reversing the order does not solve the problem. Sending the email first creates the possibility of notifying the team about feedback that never makes it into the database.

The underlying issue is that the application is attempting to coordinate two independent operations without a shared atomic boundary. This is the dual-write problem.

The transactional outbox pattern addresses it by changing what the application commits.

Instead of trying to save the feedback and immediately notify the team, the backend saves the feedback record and an event describing its submission in the same database transaction. The event represents the durable intent to perform the downstream work.

The transaction has two responsibilities: persist the feedback and record that its submission must be processed.

If the transaction commits, both the feedback record and the event exist. If it rolls back, neither does.

There is no committed state in which the feedback exists but the application has failed to record that a notification is required.

For a feedback submission, the transaction can be represented as:

BEGIN;

INSERT INTO feedback (...);

INSERT INTO outbox_events (
    event_id,
    event_type,
    aggregate_id
) VALUES (
    ...,
    'feedback.submitted',
    ...
);

COMMIT;

This illustrates the transactional relationship rather than a specific database implementation. The actual schema, transaction API and uniqueness constraints depend on the storage layer.

The event is not the notification itself. It is a durable record of work that still needs to happen. That distinction allows notification delivery to fail independently without losing the original submission.

The outbox moves failure out of the request

Once the transaction commits, a background worker can pick up the event and dispatch it to the appropriate handlers.

One handler sends the email. Another posts to Slack. Each handler tracks its own processing state and can retry independently when a transient failure occurs.

The user's request no longer needs to wait for either service. A slow email provider does not delay the thank-you screen, and a Slack outage does not turn a successfully saved feedback submission into an apparent failure.

The important boundary is the database commit. Once the feedback and its outbox event have been committed, the application has durably recorded both the user's submission and the work required to notify the team.

The background processing system is responsible for eventually carrying out that work.

This does not mean that the notification is delivered instantaneously. A worker may be delayed, a provider may be unavailable, or an event may need several attempts before processing succeeds. The advantage is that these failures no longer require the user to submit the feedback again.

The event remains available for processing, and the worker can resume after transient failures without repeating the original database operation.

There is, however, another retry problem waiting on the other side of the outbox.

The worker can fail after the notification succeeds

Suppose the email handler sends a notification successfully. Immediately afterwards, the worker crashes before recording that the notification was sent.

When the event is processed again, the worker sees no completed status. It sends the email a second time.

This is a different failure from the original request. Idempotency at the API boundary prevents duplicate feedback records, but it does not automatically prevent duplicate external side effects.

Each handler needs its own idempotency strategy.

Where the email provider supports idempotency keys, the handler can derive a stable key from the event and notification identity. Repeated attempts using the same key can then be deduplicated within the provider's documented guarantees.

The handler should also persist delivery state, distinguishing successful processing from intentional skips and failures. Before acting, it should check whether the notification has already been completed. Where concurrent workers are possible, that check must be supported by atomic claiming or another concurrency-control mechanism.

A database status check alone cannot eliminate every duplicate. There is still a window between an external service accepting a notification and the application recording the result. If the worker crashes in that window and the provider offers no idempotent operation, the application may have to choose between risking a duplicate and risking a lost notification.

This is why the phrase exactly once needs to be used carefully. A database can guarantee that a particular record is created once under the appropriate uniqueness constraints. A transaction can guarantee that a record and its outbox event commit together. External delivery requires additional guarantees from the receiving service or a clearly defined recovery strategy.

The system becomes reliable not by assuming failures cannot happen, but by deciding how each failure can be recovered from.

The guarantees are different at each boundary

The complete flow has three separate boundaries.

At the client boundary, the idempotency key ensures that multiple HTTP requests can represent the same user action. At the database boundary, atomic writes ensure that the feedback record and its event cannot be committed independently. At the notification boundary, background handlers and their retry policies make downstream processing recoverable without coupling it to the original request.

Each mechanism solves a different problem. None substitutes for the others.

An idempotency key without a database uniqueness constraint can still fail under concurrency. A transactional outbox without a reliable worker can accumulate events that never get processed. A worker with retries but no idempotency strategy can repeatedly perform the same external action.

The architecture works because the guarantees compose across these boundaries, rather than because any single component promises that everything happens exactly once.

Observability is part of that design too. Pending events, failed handlers, retry counts and delivery delays need to be visible. Otherwise, a system can preserve every record correctly while leaving notifications stuck indefinitely. Durable work is only useful if failures can be detected and recovered from.

Where this fits at Hoomanely

Feedback is one example of a broader pattern across the Hoomanely application.

Logging a health event, saving a vault document, answering a quiz question or publishing a community post all begin with a user action that should produce one logical result. A network failure can cause the client to retry, while a downstream operation such as a notification or reminder can fail independently of the original write.

The same principles apply, although not every feature needs the complete outbox architecture. An operation that only writes a database record may require little more than idempotent handling and appropriate constraints. An operation that must reliably trigger work in another service has a stronger consistency requirement.

The design decision should follow from the consequences of failure: what must be saved, what must happen afterwards, which operations can be repeated safely, and how the system recovers if a process stops between two steps.

These questions become increasingly important as a product adds more integrations and event-driven behavior. Without explicit guarantees, every new downstream action introduces another place where the application can silently diverge from the state users believe it has reached.

What a reliable Send button actually guarantees

The visible interface has barely changed. There is still a text field, a Send button and a confirmation screen.

Underneath it, the application now distinguishes one user action from multiple network attempts, uses database constraints to prevent duplicate submissions, commits the feedback and its event atomically, and processes notifications independently with explicit retry and recovery behavior.

The important result is not that every operation executes exactly once under every possible failure. Distributed systems cannot provide that guarantee across arbitrary independent services without additional assumptions and cooperation.

The result is that each boundary has a defined contract. A retry does not create another submission. A committed record does not lose its notification intent. A temporary downstream failure does not invalidate the user's successful action. And when a worker fails, the system has a defined way to recover rather than relying on someone to discover the missing message later.

The Send button is only the entry point. Reliability comes from preserving the meaning of that single action all the way from the device to the database and beyond.

Read more