Pagination for large result sets
A search for all active profiles in Bihar matches about 140,000 of them in the lab and would match tens of millions at 300 million profiles. The Search API returns at most 50 per request, so it needs a way to ask for the next 50, and the next, without getting slower on every page and without showing a profile twice or skipping one when data changes between requests.
This chapter reproduces the limits of page numbers, then replaces them with cursors. You page with search_after, watch a profile disappear between pages when the index changes, and prevent that with a point in time (PIT). Then the Search API gets an opaque, tamper-evident cursor with an optional stable mode, verified by walking a whole result set page by page and comparing it with PostgreSQL.
You need the lab over TLS and the search-api module from chapter 12. The chapter takes about 45 minutes.
Page numbers with from and size
What it is. from skips a number of hits and size returns the next ones, like OFFSET and LIMIT in SQL.
Why it matters at 300 million profiles. To return hits 10,000 to 10,010, every shard must find and sort its own top 10,010, and the coordinating node must merge all those lists and throw away all but ten. The pagination guide describes the cost: each shard loads the requested hits and those of every previous page into memory. Cost grows with depth and with the number of shards, and a page deep in a result set is the most expensive request the API can receive.
Example. Elasticsearch caps the window with index.max_result_window, which is 10,000 by default:
GET user-profile-read/_search?size=10&from=9995{ "query": { "term": { "state": "bihar" } } }illegal_argument_exception: Result window is too large, from + size must be less than or equal to: [10000] but was [10005].See the scroll api for a more efficient way to request large data sets. This limit can be set by changing the[index.max_result_window] index level setting.Common mistake. Raising index.max_result_window so that page 1,000 works. The limit is a protection, not an arbitrary number: raising it makes the most expensive requests possible. The error message’s own suggestion, scroll, is not the answer for user-facing search either, as a later section explains.
Cursors with search_after
What it is. Instead of saying how many hits to skip, the request says where the previous page ended. Every hit in a sorted search carries its sort values; search_after takes the last hit’s values and returns the hits that sort after them. Each shard then needs only size hits past that point, however deep the page.
Why it matters at 300 million profiles. Page 1,000 costs the same as page 2. The requirement is a total sort order, so that “after this hit” has one meaning. Chapter 11’s sorts all end with the unique userId for exactly this reason.
Example. Active profiles in Patna, newest first, three per page:
GET user-profile-read/_search?filter_path=hits.hits._id,hits.hits.sort&size=3{ "_source": false, "query": { "bool": { "filter": [ { "term": { "city": "patna" } }, { "term": { "accountStatus": "ACTIVE" } } ] } }, "sort": [ { "updatedAt": "desc" }, { "userId": "asc" } ]}{"hits":{"hits":[{"_id":"560997","sort":[1787875200000,"560997"]}, {"_id":"196496","sort":[1787702400000,"196496"]}, {"_id":"385496","sort":[1787702400000,"385496"]}]}}For the next page, send the same request with the last hit’s sort array:
GET user-profile-read/_search?filter_path=hits.hits._id,hits.hits.sort&size=3{ ...the same query and sort..., "search_after": [1787702400000, "385496"]}{"hits":{"hits":[{"_id":"106495","sort":[1787443200000,"106495"]}, {"_id":"311992","sort":[1787270400000,"311992"]}, {"_id":"911997","sort":[1787270400000,"911997"]}]}}PostgreSQL’s ORDER BY updated_at DESC, user_id::text ASC returns the same six profiles in the same order. The ::text is deliberate: userId is a keyword, so it sorts as a string, and "196496" comes before "385496" whatever their numeric values. For a tiebreaker, only uniqueness matters.
Common mistake. Ending a sort with a field that is not unique, such as updatedAt alone. Profiles updated at the same moment then have identical sort values, and search_after skips the ones that tie with the last hit of a page.
A page can shift while the user reads it
search_after asks the current index for what comes next. If profiles change between two page requests, the second page reflects the change, and a profile can move from a page the user has not seen yet to one they already have.
Reproduce it on a small scratch index of nine profiles, updated on 1 to 9 August:
PUT ch13-pages{ "settings": { "number_of_shards": 1, "number_of_replicas": 0 }, "mappings": { "properties": { "userId": { "type": "keyword" }, "updatedAt": { "type": "date" } } }}
POST ch13-pages/_bulk?refresh=true{"index":{"_id":"1"}}{"userId":"1","updatedAt":"2026-08-01T00:00:00Z"}{"index":{"_id":"2"}}{"userId":"2","updatedAt":"2026-08-02T00:00:00Z"}...{"index":{"_id":"9"}}{"userId":"9","updatedAt":"2026-08-09T00:00:00Z"}Page 1, three per page, newest first, returns profiles 9, 8, and 7, and the last sort values are [1786060800000, "7"]. Now, before the user asks for page 2, profile 4 is updated and becomes the newest:
PUT ch13-pages/_doc/4?refresh=true{ "userId": "4", "updatedAt": "2026-08-10T00:00:00Z" }Page 2 with search_after:
{"hits":{"hits":[{"_id":"6","sort":[1785974400000,"6"]},{"_id":"5","sort":[1785888000000,"5"]},{"_id":"3","sort":[1785715200000,"3"]}]}}Profile 4 is gone. It now sorts first, onto page 1, which the user has already read, so it never appears. The same mechanism can also show a profile twice when it moves the other way.
A point in time holds the view still
What it is. A point in time is a lightweight view of an index’s state at the moment it is opened. Searches against the PIT see that state, however the index changes afterwards. Each search extends the PIT’s keep_alive.
Why it matters at 300 million profiles. Most interactive users never page far enough to notice a shifted page. Exports, reconciliation jobs, and admin tools that walk an entire result set do notice, because they promise “every matching profile, once”.
Example. Reset the scratch index to its original state by putting profile 4 back at 4 August. Run page 1 again, and this time open a PIT before profile 4 changes:
POST ch13-pages/_pit?keep_alive=2m{"_shards":{...},"id":"2KLBBAEKY2gxMy1wYWdlcx..."}Move profile 4 to 10 August again, then request page 2 against the PIT. A PIT search has no index in its path; the PIT defines the target:
GET _search?filter_path=hits.hits._id,hits.hits.sort&size=3{ "_source": false, "pit": { "id": "<id from the previous response>", "keep_alive": "2m" }, "sort": [ { "updatedAt": "desc" }, { "userId": "asc" } ], "search_after": [1786060800000, "7"]}{"hits":{"hits":[{"_id":"6","sort":[1785974400000,"6"]},{"_id":"5","sort":[1785888000000,"5"]},{"_id":"4","sort":[1785801600000,"4"]}]}}Profile 4 appears in its old position, because the PIT still sees the index as it was when paging started. A new search without the PIT returns 4, 9, 8 as the first page. Close the PIT when you are done:
DELETE _pit{ "id": "<id>" }{"succeeded":true,"num_freed":1}Trade-off. A PIT is not free. The pagination guide notes that it keeps old segments alive, which costs disk space and file handles until the PIT is closed or expires. Every open PIT is a small resource held on every shard it covers. The reference also says PIT searches add an implicit _shard_doc tiebreaker; with this series’ sort, which already ends in the unique userId, the sort values returned in the lab contained only the two declared keys.
Delete the scratch index:
DELETE ch13-pagesScroll: for batch extraction, not for users
The scroll API also pages through a frozen view, and older code uses it for large exports. The pagination guide is explicit: Elastic no longer recommends scroll for deep pagination, and recommends search_after with a PIT when index state must be preserved beyond 10,000 hits. Scroll contexts are tied to one search and cannot be used for interactive, user-driven paging with a changing page size or a sort chosen per request. For new code, use PIT and search_after for exports as well; chapter 14’s reindexing uses _reindex, which handles its own scrolling internally.
The Search API’s cursor contract
The API exposes pagination as a contract with three rules:
- The cursor is opaque. Clients receive
nextCursorand send it back ascursor. They must not parse or build it, so its format can change. - The cursor belongs to one search. It records a fingerprint of the search parameters. Reusing it with different filters is an error, not a silent jump into another result set.
- Stability is a choice.
stable=trueopens a PIT on the first page and carries its ID in the cursor. The default,false, keeps no server-side state, and accepts that pages can shift, which suits most interactive use.
| Request | Response |
|---|---|
First page: search parameters, no cursor | Hits, total, nextCursor (null if this is the last page) |
Next page: the same parameters plus cursor | The next hits and the next nextCursor |
cursor from a different search, or malformed | 400 Bad Request |
stable=true cursor whose PIT has expired | 410 Gone: start the search again |
size may change between pages; it is not part of the fingerprint.
Stage 1 — Cursors in Kotlin
The request gains cursor and stable, and the response gains nextCursor. In SearchDtos.kt, add the two fields at the end of UserSearchRequest and nextCursor to UserSearchResponse:
@field:Min(1) @field:Max(50) val size: Int = 20, /** Opaque token from the previous page's `nextCursor`; absent for the first page. */ @field:Size(max = 4096) val cursor: String? = null, /** Page over a point-in-time snapshot, so concurrent changes cannot shift pages (chapter 13). */ val stable: Boolean = false,)
data class UserSearchResponse( val total: TotalHits, val hits: List<UserSearchHit>, /** Pass as `cursor` to get the next page; null on the last page. */ val nextCursor: String?,)The cursor itself:
package `in`.o612.eng.usersearch.api.search
import co.elastic.clients.elasticsearch._types.FieldValueimport `in`.o612.eng.usersearch.api.web.UserSearchRequestimport tools.jackson.databind.json.JsonMapperimport tools.jackson.module.kotlin.KotlinModuleimport tools.jackson.module.kotlin.readValueimport java.security.MessageDigestimport java.util.Base64
/** * Where the next page starts: the last hit's sort values, plus a point-in-time id for stable paging. * Serialised as opaque base64url JSON, so clients cannot depend on its contents. */data class SearchCursor( val after: List<Any>, val pitId: String?, /** Ties the cursor to the search it came from. */ val fingerprint: String,) { fun searchAfter(): List<FieldValue> = after.map { value -> when (value) { is String -> FieldValue.of(value) is Int -> FieldValue.of(value.toLong()) is Long -> FieldValue.of(value) is Double -> FieldValue.of(value) else -> throw InvalidCursorException() } }
fun encode(): String = Base64.getUrlEncoder().withoutPadding().encodeToString(mapper.writeValueAsBytes(this))
companion object { private val mapper = JsonMapper.builder().addModule(KotlinModule.Builder().build()).build()
fun decode(token: String, request: UserSearchRequest): SearchCursor { val cursor = try { mapper.readValue<SearchCursor>(Base64.getUrlDecoder().decode(token)) } catch (e: Exception) { throw InvalidCursorException() } if (cursor.fingerprint != fingerprint(request)) throw InvalidCursorException() return cursor }
fun from(lastSortValues: List<FieldValue>, pitId: String?, request: UserSearchRequest) = SearchCursor( after = lastSortValues.map { it._get() }, pitId = pitId, fingerprint = fingerprint(request), )
/** Everything that defines the result set, but not the page size or position. */ fun fingerprint(request: UserSearchRequest): String { val identity = request.copy(cursor = null, size = 0).toString() val digest = MessageDigest.getInstance("SHA-256").digest(identity.toByteArray()) return Base64.getUrlEncoder().withoutPadding().encodeToString(digest).take(16) } }}
/** The cursor is malformed, or belongs to a different search. */class InvalidCursorException : RuntimeException("Invalid cursor for this search")
/** The point in time behind a stable cursor has expired. */class CursorExpiredException : RuntimeException("The cursor has expired; start the search again")The sort values are kept as plain JSON values and turned back into FieldValues for the request: _score becomes a double, updatedAt a long, and userId a string. The fingerprint hashes the request with cursor and size blanked, so everything that defines the result set is covered. It is a tamper-evidence measure, not a security boundary: a caller who crafts a cursor can only page through a search they are already allowed to run.
UserQueryBuilder.build learns to target a PIT and to continue after a position. Replace its beginning, up to .size(request.size):
/** How long Elasticsearch keeps a point in time open between two page requests. */ const val PIT_KEEP_ALIVE = "2m"
/** * @param pitId search this point in time instead of the read alias (stable paging, chapter 13) * @param searchAfter the previous page's last sort values; empty for the first page */ fun build( request: UserSearchRequest, pitId: String? = null, searchAfter: List<FieldValue> = emptyList(), ): SearchRequest = SearchRequest.of { s -> if (pitId != null) { s.pit { p -> p.id(pitId).keepAlive { k -> k.time(PIT_KEEP_ALIVE) } } } else { s.index(UserProfileIndex.READ_ALIAS) } if (searchAfter.isNotEmpty()) s.searchAfter(searchAfter) s.size(request.size)The rest of build, from .source onwards, is unchanged. In UserSearchService, replace search and add two helpers. Import co.elastic.clients.elasticsearch._types.ElasticsearchException:
fun search(request: UserSearchRequest): UserSearchResponse { val cursor = request.cursor?.let { SearchCursor.decode(it, request) } // A stable search opens a point in time on its first page and carries it in the cursor. val pitId = cursor?.pitId ?: if (request.stable) openPointInTime() else null val searchRequest = UserQueryBuilder.build(request, pitId, cursor?.searchAfter().orEmpty())
val response = try { client.search(searchRequest, ResultSource::class.java) } catch (e: ElasticsearchException) { if (pitId != null && e.status() == 404) throw CursorExpiredException() throw e }
val hits = response.hits().hits() val currentPit = response.pitId() ?: pitId val nextCursor = if (hits.size == request.size) { SearchCursor.from(hits.last().sort(), currentPit, request).encode() } else { currentPit?.let { closePointInTime(it) } // last page: release the snapshot now null } val total = response.hits().total() return UserSearchResponse( total = TotalHits(total?.value() ?: 0, exact = total?.relation() == TotalHitsRelation.Eq), hits = hits.mapNotNull { hit -> hit.source()?.let { src -> UserSearchHit( userId = src.userId, fullName = src.fullName, city = src.city, state = src.state, accountStatus = src.accountStatus, updatedAt = src.updatedAt, score = hit.score(), // "fullName.prefix" and "city.text" are index details; callers see "fullName" and "city". highlights = hit.highlight().mapKeys { (field, _) -> field.substringBefore('.') }, ) } }, nextCursor = nextCursor, ) }
private fun openPointInTime(): String = client.openPointInTime { it.index(UserProfileIndex.READ_ALIAS).keepAlive { k -> k.time(UserQueryBuilder.PIT_KEEP_ALIVE) } }.id()
private fun closePointInTime(id: String) { client.closePointInTime { it.id(id) } }Three details carry the design. A full page produces a cursor; a short page is the last one, and closes the PIT immediately instead of waiting for keep_alive. The PIT ID is taken from the response, because Elasticsearch may return a different ID than the one sent. And a 404 from a PIT search means the PIT is gone, which becomes CursorExpiredException.
In ErrorHandling, map the two cursor exceptions, importing both from the search package:
@ExceptionHandler(InvalidCursorException::class) fun invalidCursor(e: InvalidCursorException): ProblemDetail = ProblemDetail.forStatusAndDetail(HttpStatus.BAD_REQUEST, e.message ?: "Invalid cursor")
@ExceptionHandler(CursorExpiredException::class) fun cursorExpired(e: CursorExpiredException): ProblemDetail = ProblemDetail.forStatusAndDetail(HttpStatus.GONE, e.message ?: "Cursor expired")Stage 2 — Verify the pagination
Restart search-api. Request the first page of the Patna list:
curl -s 'localhost:8080/api/users/search?city=patna&sort=UPDATED_AT&size=3'The hits are 560997, 196496, and 385496, the total is {"value": 10000, "exact": false}, and nextCursor is an opaque string. Decoded, purely for inspection, it holds:
{"after":[1787702400000,"385496"],"pitId":null,"fingerprint":"yMfwY_wWXC17vyWY"}Send it back with the same parameters and the next page is 106495, 311992, and 911997: the Dev Tools result and PostgreSQL’s order. Send it with city=gaya instead, or send cursor=not-a-cursor:
{"detail":"Invalid cursor for this search","instance":"/api/users/search","status":400,"title":"Bad Request"}Relevance-sorted searches page the same way; the cursor then starts with the score, for example {"after":[25.586945,1761696000000,"343200"],...} for name=prashant%20kumar&city=patna.
Now walk an entire result set in stable mode, 50 per page, and compare with PostgreSQL. A short script follows nextCursor until it is null:
python3 - <<'EOF'import json, urllib.requestbase = "http://localhost:8080/api/users/search?name=prashant%20kumar&state=bihar&city=patna&size=50&stable=true"seen, cursor, pages = [], None, 0while True: page = json.load(urllib.request.urlopen(base + (f"&cursor={cursor}" if cursor else ""))) pages += 1 seen += [hit["userId"] for hit in page["hits"]] cursor = page["nextCursor"] if not cursor: breakprint("pages", pages, "hits", len(seen), "unique", len(set(seen)))EOFpages 3 hits 117 unique 117PostgreSQL counts 117 active Prashant Kumars in Patna: every profile once, none twice. While a stable walk is in progress, GET _nodes/stats/indices/search?filter_path=nodes.*.indices.search.open_contexts reports one open context; after the last page it reports zero, because the service closed the PIT.
Finally, expiry. Take a stable cursor, close its PIT yourself with DELETE _pit, and send the cursor:
{"detail":"The cursor has expired; start the search again","instance":"/api/users/search","status":410,"title":"Gone"}Checkpoint
A full stable walk returns each matching profile exactly once, and the open-context count returns to zero afterwards. If a walk returns duplicates, check that the sort still ends with userId.
Production note — A stable search that a user abandons holds its PIT until
keep_aliveexpires, two minutes after the last page request. Under heavy use, that is many open contexts per shard. Offerstable=trueto exports and admin tools, keepkeep_aliveshort, and watch open contexts in monitoring (chapter 16). When a result set divides evenly into pages, the last full page still returns a cursor, and one extra request returns an empty page that closes the PIT.
Common mistakes with pagination
- Deep
from/size, or raisingmax_result_windowto allow it. Use cursors. - A sort without a unique tiebreaker.
search_afterskips ties. - Transparent cursors. Clients start building them, and the format can never change again.
- A PIT for every request. Stability costs server-side resources; make it opt-in.
- Forgetting to close a PIT. Close it on the last page, and keep
keep_aliveshort. - Scroll for interactive paging. It is for batch extraction, and PIT with
search_afterhas replaced it there too.
What you built, and what comes next
The Search API pages with opaque cursors that cost the same at any depth, reject reuse with a different search, and optionally hold a point-in-time view so that a complete walk returns every match exactly once. You verified all of it against PostgreSQL.
What is still missing is the write path. The lab index was filled by chapter 05’s shell loader, and the outbox from chapter 04 has been accumulating events ever since.
Chapter 14 builds the indexer: a Kotlin backfill that reads PostgreSQL with keyset pagination and bulk-indexes with backpressure and per-item error handling, the outbox relay designed in chapter 04, and a zero-downtime reindex from user-profile-v1 to a v2 with display fields for facets, ending in an atomic alias swap.