Link copied.
DigitalWerks Insights

API Pagination: How to Sync Every Page Without Missing or Duplicating Records

Reliable API pagination needs more than a page loop. Learn how to use continuation tokens, stable identifiers, durable checkpoints, and reconciliation to prevent missing or duplicate records.
Editorial still life showing data packets moving through pagination checkpoints toward a reconciled archive
DigitalWerks field note

Many APIs return records in pages. That sounds simple until a sync stops after page 3, follows a continuation token incorrectly, or starts a second run from the wrong boundary and writes duplicates. The request can succeed every time while the destination still contains a partial or inconsistent copy of the source.

Reliable pagination is not just a loop around an HTTP request. It is a small data-movement workflow with a cursor, a stopping rule, a stable identity key, durable progress, and a way to reconcile what was read with what was written.

Pagination is a state machine, not a page counter

A paginated API usually returns some combination of records, a continuation token or next-page URL, and metadata about the current response. Some APIs use numbered pages. Others use cursor-based pagination, where the server returns an opaque value that represents the next position. A few combine a limit with a timestamp or record ID.

Your integration should treat the response as state:

  • Input state: the request parameters, cursor, filters, and source scope.
  • Read state: the records returned, the response status, and the next-page value.
  • Write state: which records were accepted, rejected, retried, or left uncertain.
  • Progress state: the last safe cursor or boundary that can be resumed.

That separation matters because “the request completed” and “the page was safely applied” are different events. A worker can receive page 4, lose its database connection while writing, and restart from page 4. If the write was partly successful and the operation is not idempotent, the retry can create duplicates. If it restarts from page 3, it may repeat records. If it advances to page 5 before page 4 is committed, it can create a gap.

Prefer the API’s continuation signal over guessing

When an API provides a next-page URL or opaque cursor, persist and use that value. Do not reconstruct it by incrementing a page number unless the API documentation explicitly defines that behavior.

Cursor values can encode server-side ordering, filters, or a snapshot boundary. Treat them as opaque data. Do not parse them for meaning, store only part of them, or reuse one cursor with a different filter set.

Numbered pages still need a defined ordering. If the source can insert or delete records while the sync runs, page 2 may not represent the same slice of data after page 1 completes. A stable sort by an immutable field, a source-provided snapshot, or a time-bounded extraction window can reduce that risk. Without a stable boundary, pagination can miss a record that moves between pages or read one record twice.

For an incremental sync, record the scope alongside the cursor. A useful checkpoint might include the source object type, filter values, ordering rule, start and end timestamps, page size, cursor, and run ID. A bare cursor is difficult to audit and dangerous to resume under different conditions.

Make the write safe before you make the loop fast

The destination should have a stable source identifier and a defined upsert rule. If the source record ID is 8472, the destination should be able to recognize that the next page contains the same record, even if the record has changed since the previous run.

For each record, decide what happens when the destination already contains the source ID:

  • Insert a new record when the source ID is unknown.
  • Update the existing record when the source version or updated timestamp is newer.
  • Ignore the record when the destination already has an equal or newer version.
  • Quarantine the record when identity, required fields, or mapping rules are invalid.

This is where pagination meets the broader problem of duplicate prevention. A retry-safe write needs a unique constraint or equivalent lookup on the source system and object type. If the workflow creates related records or side effects, give those operations their own idempotency key or reconciliation rule. The page loop should not be responsible for guessing whether a prior attempt finished.

DigitalWerks has also covered why a successful retry needs idempotency keys and why a successful API request is not the same as a successful data sync. Pagination adds another boundary to validate, but the same principle applies: request status is only one signal in the workflow.

Checkpoint after a safe page boundary

A checkpoint should advance only after the page’s records and its outcome have been durably recorded. A practical sequence is:

  1. Read the current cursor or page request from the run state.
  2. Request one page using the exact stored scope.
  3. Validate the response shape, source identifiers, and expected page metadata.
  4. Stage the records with the run ID and page token.
  5. Apply inserts, updates, rejects, and side effects using retry-safe rules.
  6. Record counts and exceptions for the page.
  7. Commit the page outcome and next cursor together.
  8. Continue only when the checkpoint is durable.

The important boundary is between steps 6 and 7. If the process fails before the checkpoint is committed, the page can be retried. If it fails after the checkpoint is committed, the next run can continue. Either way, the destination must tolerate a repeated page because the failure may occur after a write but before the checkpoint, or after the source response but before the destination write.

Keep page-level state separate from run-level state. A run may complete 18 pages and fail on page 19. That is different from a run that received page 19 but had 12 rejected records. The audit record should make both facts visible.

Know when page counts are lying

Page size and total-count metadata are useful for planning, but they are not always a safe completion signal. A source can change while the sync runs. Some APIs omit total counts. Others return an approximate count or apply permissions that change the visible result set.

Use the documented next-page signal as the primary stopping rule. Stop when the API says there is no next page, when the response contains no records and no continuation value, or when the source-specific contract defines another terminal condition.

Set operational guardrails as well:

  • Maximum pages or records per run.
  • Maximum elapsed time and request count.
  • Minimum and maximum accepted page sizes.
  • Detection for a repeated cursor or repeated page fingerprint.
  • Alerts when the page size unexpectedly drops, grows, or becomes zero.
  • A run status that distinguishes complete, partial, failed, and needs review.

A repeated cursor is especially important. Without a guard, one bad response can create an infinite loop that repeatedly reads and potentially rewrites the same page.

Validate the whole result, not just the last page

After a paginated run, compare the source-side read results with the destination-side write results. At minimum, retain the number of pages requested, records received, records inserted, records updated, records skipped, records rejected, and records left uncertain.

Counts are a starting point, not proof. A page can contain duplicates, records can be rejected after being read, and two source objects can collide if the destination key is wrong. Validate a sample of source IDs across the first, middle, and last pages. Check that the extraction scope and filters match the intended run. Compare the newest source timestamp with the newest destination timestamp. Review rejected records and confirm that the next run will not silently move past them.

For high-value workflows, keep a staging table or equivalent audit store. It gives operators a way to answer questions such as: Which cursor produced this record? Was it written on the first attempt? Did it fail validation? Was it replayed? Did the source change between pages?

A practical test matrix for pagination

Before production, test more than a happy path with two full pages. Include:

  • One page smaller than the requested limit.
  • Exactly one full page followed by a terminal response.
  • A missing or malformed continuation value.
  • A repeated continuation value.
  • A timeout after the source returns data.
  • A destination failure after some records are written.
  • A retry of the same page.
  • A record updated between page requests.
  • A deleted or permission-filtered record.
  • One invalid record mixed into an otherwise valid page.
  • An empty result set.
  • A run that stops at the operational page limit and resumes later.

Verify both data and operational behavior. The integration should leave a clear status, a durable checkpoint, a traceable run ID, and enough detail to explain what happened without replaying the entire job manually.

Pagination becomes reliable when progress is explainable

Every page is a boundary where records can be missed, repeated, rejected, or written without a durable checkpoint. The solution is not a more aggressive loop. It is a workflow that preserves the request scope, trusts the documented continuation signal, writes against stable source identity, checkpoints only after safe application, and reconciles the complete run.

If a system is already losing records or creating duplicates, DigitalWerks can review the pagination contract, cursor storage, destination keys, retry behavior, page-level audit trail, and reconciliation checks as one integration workflow. That review is usually more useful than inspecting the HTTP client in isolation.

Useful? Pass it on.Share this field note with someone who can use it.
From insight to implementation

Make the rest of your digital system work this clearly.

DigitalWerks connects strategy, websites, software, analytics, integrations, and AI-ready operations into one dependable system.

Start a conversation