A scheduled sync can fail quietly. The job may time out, lose its API token, hit a rate limit, or never start after an infrastructure change. If the next run only asks for “records changed since the last successful timestamp,” the missed window can remain missing forever.
Recovering from that gap is not as simple as running the same job again. A safe backfill needs a known source range, a stable record identity, duplicate protection, and a reconciliation step that proves the gap was closed. Without those controls, recovery can create a second problem: duplicate records, repeated emails, conflicting updates, or inflated reporting.
Start with the missed window, not with “everything since last time”
Before replaying data, identify exactly what the failed run was supposed to cover. A useful sync record should capture at least:
- The source system and destination system.
- The intended start and end of the source range.
- The cursor, timestamp, sequence number, or page marker used by the run.
- The number of records discovered, accepted, rejected, and written.
- The run status and the reason it stopped.
That information turns a vague incident into a bounded recovery job. For example, if a nightly customer sync completed at 2:00 a.m. on Monday and the next successful run was Wednesday morning, the recovery range might cover changes from Monday at 2:00 a.m. through Wednesday at the time the normal process resumes. The exact boundary depends on the source system and its timestamp semantics.
Do not assume the job’s local clock, the source system’s clock, and the business reporting date mean the same thing. Store the boundary in an unambiguous format, record the time zone, and document whether the boundary is inclusive or exclusive. A range such as [start, end), meaning “include start, exclude end,” is easier to reason about than an informal phrase like “since Monday.”
Use a cursor that can survive a retry
Incremental syncs commonly use a modified timestamp, an auto-incrementing ID, a source sequence, or a vendor-specific cursor. Each option has tradeoffs.
A timestamp is easy to understand but can miss records when several changes share the same precision, or when a record is updated with a clock that is slightly behind the sync process. An ID works well when IDs increase with creation order, but it may not capture updates to older records. A vendor cursor can be precise, but the integration must store it safely and understand how long it remains valid.
For higher-risk workflows, use a small overlap between runs and deduplicate by a stable source identity. A run might read a few minutes or pages before the last stored cursor, then use an upsert or idempotency key so previously processed records are updated rather than inserted again. The overlap protects against boundary errors; the identity rule keeps the overlap from becoming duplicates.
Document the cursor as part of the integration contract. The team operating the sync should know what it represents, when it advances, and whether it is safe to move it forward after partial success.
Separate discovery from writing
A recovery process is easier to control when it has distinct stages:
- Discover. Read the source range and record the candidate count, identifiers, and source-side status.
- Stage. Store the extracted records, run ID, source range, and payload checksum in a temporary or durable staging area.
- Validate. Check required fields, identifier presence, data types, relationships, and records that already exist in the destination.
- Write. Upsert approved records using stable identity and bounded batches.
- Reconcile. Compare source discovery counts and identities with destination results, including rejects and exceptions.
This design creates a pause point between reading and changing the destination. If the source returns an unexpected number of records, a field suddenly changes shape, or the target API starts rejecting requests, the recovery can stop before it compounds the problem.
For a small system, staging may be a database table or a protected file with a run identifier. For a larger integration, it may be a queue or object-storage batch. The implementation can vary, but the purpose stays the same: make the recovery set visible, repeatable, and reviewable.
Make the write operation idempotent
Backfills are often retried because the first attempt times out after writing some records. The next attempt must be able to recognize work that already succeeded. A common pattern is to define a destination key from the source system name, object type, and stable source ID. If the destination supports an external ID, use it. Otherwise, maintain a mapping table between source and destination identifiers.
For operations that create side effects, store a deterministic operation key for the specific source record and action. The same key should be reused when the recovery retries that operation. AWS describes idempotency as a way to make repeated requests have the same effect as one request, and Stripe documents a similar pattern for safely retrying create or update requests. These are platform examples, but the design lesson is general: retries need an identity they can reuse.
Do not use a new random key every time a backfill retries. A new key can make the same source event look like a new operation. Also do not treat an HTTP success response as proof that the destination is correct. Read back a sample, capture the destination ID, and reconcile the final state.
Keep the recovery bounded
A missed run should not turn into an uncontrolled historical import. Define limits before starting:
- The source time or cursor range.
- The maximum number of records per batch.
- The maximum request rate and concurrency.
- The number of retries for transient failures.
- The conditions that pause the job for human review.
Boundaries protect the source and destination from sudden load. They also make it easier to estimate how long a recovery will take and to identify which batch introduced a problem. If a vendor API returns rate-limit responses, use the platform’s guidance for retry timing and keep the retry policy separate from the backfill’s record-level error handling.
It is also useful to give a recovery run its own identifier and status. A status such as planned, staged, writing, paused, reconciled, or needs-review gives operators a shared language for what happened.
Reconcile the gap from more than one direction
A count comparison is a good start, but it is not enough. Two systems can report the same number of records while containing different records. Reconciliation should include:
- Source candidates versus staged records.
- Staged records versus accepted writes.
- Accepted writes versus destination reads.
- Rejected records and the reason for each rejection.
- Duplicate candidates and how they were resolved.
- Records updated more than once during the recovery.
For important workflows, compare sets of stable IDs and sample the resulting field values. If the sync moves donations, orders, memberships, or customer updates, include the status and amount fields that affect reporting. If it updates a CRM, confirm the destination record and any related activity or campaign attribution that should have been created.
Keep the audit evidence long enough to investigate the incident and answer operational questions. Do not copy unnecessary sensitive data into logs. Record identifiers, counts, hashes, timestamps, and error categories where those are enough to prove what happened.
Test the recovery path before the next failure
A recovery workflow that exists only in a runbook will be slow and fragile during an incident. Test it with a controlled data set and a non-production destination when possible. At minimum, simulate:
- A missed run followed by a normal incremental run.
- A timeout after the destination accepts some records.
- Duplicate source records in the recovery range.
- A record that changes twice before the backfill runs.
- An invalid record that must be rejected without stopping the batch.
- A cursor that is missing, expired, or no longer valid.
Confirm that the normal job and the recovery job cannot advance the same cursor independently. Decide who approves a large backfill, who receives alerts, and how the team stops or resumes the run. The objective is not to guarantee that nothing fails. It is to make failure visible and recovery predictable.
Plan the backfill before you need it
Missed sync runs are operational events, not just engineering inconveniences. The safest response is a small, explicit workflow: identify the gap, stage the source range, validate the records, write with stable identity, and reconcile the result. That process avoids the two tempting shortcuts, rerunning everything and trusting a successful request.
DigitalWerks can help review a scheduled integration, define its cursor and recovery boundaries, add duplicate protection, and build a reconciliation process that operations teams can actually use. Talk with DigitalWerks about a sync workflow that can recover cleanly when a run is missed.