Skip to content
Back to journal

A Background Job Is a Customer Promise

Long-running work deserves a customer-facing contract: an honest receipt, visible progress, safe recovery, and a durable outcome.

Exports, imports, video processing, large uploads, report generation, account migrations, and many AI-assisted tasks cannot finish while a customer waits on one screen. Moving that work into the background is often the right engineering decision.

But “we queued it” is not a customer outcome. It is an internal implementation detail.

The person who started the work still needs to know whether the request was accepted, what will happen next, whether they can leave, how to find the result, and what to do if it does not finish. Support needs enough context to investigate without asking them to reconstruct a vanished moment. Engineering needs a way to tell a retry from a second request.

That is why a background job is a customer promise, not a queue entry.

A customer who starts long-running work can confirm what was accepted, understand the current state, return to a durable result, and recover safely when the work cannot complete.

This is a small contract, but it connects product decisions, interface design, APIs, operations, and quality around the same outcome.

Acceptance is not completion

An asynchronous request needs honest language from its first response. HTTP's 202 Accepted status means the request was accepted for processing; it does not say that processing is complete or even guaranteed to succeed later. RFC 9110 makes that distinction explicit.

The product should make it explicit too. “Your export is ready” and “We started preparing your export” are materially different statements. The first gives someone permission to leave and expect a file. The second opens a state that still needs a visible outcome.

At acceptance, show the person a concise receipt:

  • what they asked the product to do;
  • the scope that matters, such as the selected workspace, date range, or destination;
  • a durable reference or direct link to the work; and
  • the next truthful state, such as “preparing,” “waiting for approval,” or “scheduled.”

Do not quietly convert a slow action into a spinner that eventually disappears. If the work survives the current request, its record should survive the current screen.

For an API, return a stable operation resource or identifier with the acceptance response. For a product interface, link that operation to a meaningful history, activity entry, or result page. A random correlation ID pasted into a toast helps neither customers nor support unless the product gives it a place to live.

Name states by what the customer can do

Internal systems may need many transitions: queued, leased, running, retried, delayed, partially written, dead-lettered. Customers do not need a simplified copy of the scheduler. They need states that tell them what is true and what they can do.

Start with a small state model:

Requested → Preparing → Running → Complete
                     ↘ Needs attention
                     ↘ Cancelled

Each state should have a plain-language meaning. “Preparing” means the request exists and has not begun its substantive work. “Running” means work is in progress. “Complete” names the result and links to it. “Needs attention” says what blocked progress and who can take the next safe action. “Cancelled” confirms whether any partial effect remains.

Avoid a status that only reports time passing. “In progress” is useful only when it has a route to a later answer. If there is no reliable progress signal, say less—not more. A fabricated percentage turns uncertainty into an inaccurate promise. A neutral message such as “This may take a few minutes; you can leave this page and we’ll keep the result here” is often more dependable.

The same rule applies to estimated completion times. Show one only when it is based on a signal the product can sustain, and make its uncertainty clear. A precise countdown attached to a variable queue is worse than an honest expectation window.

Give retries an identity

Long-running work will fail sometimes: an upstream service throttles, a worker restarts, a file is malformed, a permission changes, or a system reaches a capacity boundary. The design question is not whether failure is possible; it is how the product distinguishes a safe retry from a second costly action.

First, decide what one customer request means. An operation reference should tie together the accepted request, any automatic attempts, customer-visible status, and final outcome. When a person presses “Try again,” the product should be able to explain whether it is continuing the same work, starting a replacement, or creating an additional result.

That distinction is especially important for consequential actions: charging a card, sending a notification, changing access, publishing content, or importing records. A retry must not quietly duplicate an external effect.

Retry behavior also needs a boundary. Google's SRE guidance notes that retries during broad overload can amplify the original problem, so clients need both a limited retry policy and a signal for when not to retry. Its handling-overload guidance describes retry budgets and the risk of retries multiplying across layers. That is an infrastructure lesson with a direct product consequence: do not show a button that repeatedly submits work when the product has no safe way to honor it.

A useful customer-facing recovery path answers three questions:

  1. What happened to this attempt?
  2. Is retrying safe now?
  3. What outcome will the next attempt replace, preserve, or add?

If the answer is “an administrator must approve this,” “the source data must be fixed,” or “try again after a stated limit resets,” say that. A generic error makes the customer guess whether the request disappeared, is still running, or should be repeated.

Keep the outcome where the work began

Background work breaks a familiar product assumption: the person may not be looking at the original screen when it ends. They may close the tab, switch devices, lose connectivity, or return after an approval changes. Completion therefore needs a durable home.

Choose the place that matches the job:

  • an export belongs in the report or export history that requested it;
  • an import belongs with the target record and its validation summary;
  • a generated artifact belongs with the project or workspace that owns it;
  • an account change belongs in the account's security or activity history.

Notifications can help, but they should not be the only evidence that the job completed. Email can be delayed, a push notification can be disabled, and a temporary banner disappears. The product needs a return path that says what was completed, when it happened, and where the resulting artifact or changed state can be examined.

This also narrows the data boundary. A notification can say “Your export is ready” and link to an authenticated product page; it does not need to include sensitive rows, a signed download URL, or internal failure detail. The operation record can retain the context needed for authorized support and audit without turning a message preview into a data leak.

Design cancellation and partial results on purpose

“Cancel” has at least three meanings: stop work that has not started, ask running work to stop, or hide the result after work has already taken effect. Treating all three as one button produces confusing and sometimes unsafe outcomes.

Define what cancellation guarantees for each operation. Can a queued report be removed before it starts? Can a running import be stopped between records? If an account migration has changed some records, can it be rolled back, resumed, or only reviewed? A customer should not be told “cancelled” if the system may still finish a side effect a moment later.

Partial success needs the same care. “Imported 84 of 100 rows” is not a failure message with a number attached. It is a decision point. Explain which items need attention, whether successful records are live, and whether a retry processes only the unresolved items or starts over. Give the person a safe way to download or correct the failed set rather than forcing manual comparison.

The wording should follow the underlying guarantee. Product copy cannot compensate for an undefined operation boundary.

Test the promise, not just the worker

The happy path proves that a worker can finish. It does not prove that a customer can safely use the feature.

Test the complete contract:

  • start work, leave the page, and return on another device or in a new session;
  • refresh immediately after acceptance and confirm exactly one visible operation remains;
  • interrupt a worker, trigger a retry, and confirm that the final result is not duplicated;
  • reach a capacity limit or dependency failure and confirm that the status gives a truthful next step;
  • cancel before, during, and after work, checking the actual effect in each case;
  • complete a partial import or generation and confirm that recovery preserves the successful work deliberately; and
  • verify that notifications, history, support context, and API status all describe the same outcome.

Measure the experience as well as the queue. Useful signals include time from acceptance to a durable result, operations abandoned because their state was unclear, repeated manual submissions for the same intent, recovery success rate, and support contacts that require a customer to find a missing result. Pair those with operational measures such as backlog age, retry rate, and terminal failures. A fast queue with an ambiguous customer outcome is still a product problem.

Make the invisible work accountable

Background processing is valuable because it lets a product do serious work without holding a person hostage to a loading screen. The trade-off is responsibility: once work leaves the immediate request, the product must carry its state, outcome, and recovery path forward.

Give the action an honest receipt. Give its states useful meaning. Keep retries bounded and identifiable. Preserve the result where the customer can return to it. Define cancellation and partial success before they occur.

Then the queue remains an implementation detail—and the customer keeps a promise they can rely on.

If your next product needs product strategy, design, engineering, and quality to agree on the consequential states behind a simple action, bring BugSquad the challenge.