An import can look healthy while quietly dropping records. Every request returns a successful status. The destination contains data. The job finishes on schedule. Yet the source has more records than the destination because the integration fetched only the first page.
This is one of the most common ways a data workflow becomes incomplete without producing an obvious error. The problem is not always authentication, mapping, or an unavailable endpoint. Sometimes the integration simply stops asking for more data.
API pagination is the mechanism an endpoint uses to divide a large collection into smaller responses. When an integration does not understand that mechanism, it can create a partial copy that looks trustworthy until someone compares totals or notices missing records.
Why APIs paginate results
Returning thousands of records in one response is expensive for both sides of an integration. Large responses take longer to generate, consume more memory, and are more likely to time out. Pagination lets the API return a bounded set and gives the client a way to request the next set.
There are two broad patterns:
- Page-based pagination: The client requests a page number and a page size, such as
page=2&per_page=100. The response may include total counts or total pages. - Cursor-based pagination: The response includes a cursor or continuation token. The client sends that value to retrieve the next page, often with a parameter such as
after,next, orstarting_after.
WordPress REST collections, for example, use page and per_page and return total counts in response headers. Stripe list endpoints use cursor-based parameters such as starting_after. The names differ, but the operational requirement is the same: keep retrieving pages until the collection is complete.
The failure pattern that hides in plain sight
Imagine a scheduled job that imports contacts from a CRM into a marketing platform. The developer tests the endpoint with a small sample and receives 100 records. The job then runs in production against 8,400 contacts.
If the code makes one request and stores the response, the job may successfully import the first 100 contacts and report success. It did not fail at the HTTP layer. It failed at the collection layer.
The same problem can affect:
- Orders imported into an ecommerce reporting database.
- Survey responses pulled into a CRM after an event.
- Donation transactions copied into a reconciliation workflow.
- WordPress posts or media retrieved for a migration.
- Support tickets or subscribers synchronized into an operations system.
The risk grows when the source sorts records by a changing field. A job that uses page numbers while new records are being added may see duplicates or skip records as the underlying collection shifts between requests.
Read the response as a contract
Before writing the loop, inspect the endpoint documentation and a real response. Look for four things.
- How does the API indicate more data? It may return a total page count, a boolean such as
has_more, a next-link object, or a continuation cursor. - What value requests the next page? The next request may increment a page number or pass back a token exactly as returned.
- How is ordering defined? A stable sort order matters when records can be created or updated during the import.
- What does an empty page mean? It may indicate completion, but it can also reflect a filter, permission boundary, or temporary inconsistency.
Do not infer the contract from the first response alone. Some APIs place pagination information in headers, some put it beside the data array, and some return a link that should be followed instead of constructing a URL yourself.
Build page retrieval as a separate step
Pagination logic should be easy to test independently from field mapping and destination writes. A useful design separates the workflow into four stages:
- Fetch: Request one page and record the request parameters, response status, and page or cursor position.
- Accumulate: Collect the records or stream them into a staging area without changing their identity.
- Validate: Check counts, duplicates, required identifiers, and page progression.
- Write: Send validated records to the destination with an idempotent create-or-update rule.
This separation makes it possible to answer a basic question: did the source return everything, or did the destination reject something after it was fetched? Without that distinction, an incomplete import and a failed write can look identical.
Choose page size with care
The largest permitted page is not automatically the best page size. A large response may reduce request count but increase timeout risk, memory use, and retry cost. A smaller page can make failures easier to resume and isolate.
Use the endpoint’s documented maximum, then test realistic payload sizes. Include records with long text, nested objects, attachments, or custom fields. Measure response time and memory rather than assuming that a page of 100 small records behaves like a page of 100 complex records.
For long-running imports, persist progress after each successful page. A checkpoint can include the source name, filter set, sort order, last page or cursor, number of records fetched, and timestamp. That creates a recovery point without pretending that a partially completed run is complete.
Protect against duplicates and skips
Page-based retrieval is sensitive to changing collections. If a new record appears between page one and page two, a record may shift positions. The client can fetch one record twice or miss one entirely.
Reduce that risk by using a stable sort field and a bounded extraction window when the API supports it. For example, capture a start timestamp, retrieve records created or updated before that boundary, and process later changes in a separate run. Cursor-based APIs often reduce movement-related problems because the cursor represents a position in the ordered collection, but they still need a documented sort rule and a replay strategy.
Every write should also be idempotent. Use a stable source identifier to update an existing destination record rather than creating a duplicate when a page is retried. If the source has no durable identifier, stop and resolve that design problem before trusting the import.
Validate the collection, not just the requests
A reliable import has checks that would fail loudly when only the first page was retrieved.
- Compare the number of records fetched with the source total when the API provides one.
- Record the number of pages or cursors processed and confirm that progression ended normally.
- Track the first and last source identifiers or timestamps for every page.
- Check for duplicate source identifiers across pages.
- Compare fetched, accepted, rejected, and written counts.
- Sample records from the beginning, middle, and end of the collection.
- Run a reconciliation query after the write and investigate differences.
Do not treat a 2xx response as proof that the dataset is complete. It proves that a request was accepted or processed. Completeness is a separate business and data-quality assertion.
What to log without creating a privacy problem
Pagination logs should contain enough evidence to troubleshoot the run without copying sensitive records into an operations log.
Useful fields include the source endpoint name, filter and sort configuration, page or cursor position, request timestamp, response status, response item count, elapsed time, retry count, and checkpoint status. Log stable internal identifiers or hashed values when needed for correlation, but avoid full names, email addresses, payment details, survey answers, or tokens.
When a page fails, the log should make the recovery path clear. An operator should be able to tell whether the job can safely retry the same cursor, resume from a saved checkpoint, or restart the extraction window and rely on idempotent writes.
A practical pre-launch test
Before an import goes live, create a source dataset large enough to require multiple pages. Include records that are valid, invalid, duplicated, updated during the run, and near the page boundary.
Then verify that the workflow:
- Requests more than one page when the collection requires it.
- Stops only when the documented completion condition is reached.
- Preserves stable source identifiers.
- Resumes safely after a simulated timeout.
- Does not duplicate records after a retry.
- Reports fetched, written, rejected, and reconciled counts separately.
- Alerts when source and destination totals diverge.
Run the test again after changing filters, sort order, page size, authentication, or field mappings. Pagination is part of the integration contract, so a small change elsewhere can affect how much data is retrieved.
Complete imports are designed, not assumed
Pagination is easy to overlook because the first request often works. The integration becomes dependable when it treats the source collection as a sequence that must be fully traversed, measured, checkpointed, and reconciled.
DigitalWerks helps organizations design and validate the data moving between CRMs, websites, forms, analytics tools, ecommerce systems, and operational platforms. If a sync reports success but no one can explain how completeness is proven, a pagination and reconciliation review is a practical place to start.
Sources: WordPress REST API pagination documentation and Stripe API pagination documentation.