Aggregations and facet counts
A search box rarely returns only hits. Next to the results sit counts: how many matches per state, per city, per gender. This chapter builds those facet counts. You run the aggregations a profile search needs in Kibana Dev Tools and check every number against PostgreSQL, reproduce the case where a top-N count is simply wrong, and add a /api/users/facets endpoint to the Search API whose counts stay useful after the user has already picked a filter.
You need the lab over TLS and the search-api module from chapter 11. The chapter takes about 45 minutes.
How aggregations work, and what they need from the mapping
What it is. An aggregation computes a summary over every document that matches the query: bucket aggregations group documents (per state, per month), and metric aggregations compute numbers (a count, a distinct count). They run in the same request as the search, over the same matching set. The aggregations reference lists every type.
Why it matters at 300 million profiles. A facet over a broad query touches millions of documents. It must read each document’s value quickly, which is what doc values provide: the per-document, column-oriented storage that keyword, numeric, and date fields have by default. Chapter 05’s capability matrix marked state, city, gender, accountStatus, and the date fields as aggregatable for exactly this reason, and chapter 06 showed what happens with a text field instead.
Example. Active profiles per state, top three, with size: 0 because only the counts are needed:
GET user-profile-read/_search?filter_path=aggregations{ "size": 0, "query": { "bool": { "filter": [ { "term": { "accountStatus": "ACTIVE" } } ] } }, "aggs": { "by_state": { "terms": { "field": "state", "size": 3 } } }}{"aggregations":{"by_state":{"doc_count_error_upper_bound":0,"sum_other_doc_count":350174, "buckets":[{"key":"maharashtra","doc_count":140008},{"key":"bihar","doc_count":139925},{"key":"karnataka","doc_count":70317}]}}}PostgreSQL’s GROUP BY state over active profiles gives the same three counts. Two fields in the response matter. sum_other_doc_count is the number of matching documents in states outside the top three. doc_count_error_upper_bound is the worst-case error of the counts shown, and here it is 0.
Common mistake. Displaying bucket keys directly. The keys are the indexed values, and state has chapter 05’s lowercase_ascii normaliser, so the UI would show maharashtra. Either map keys to display names in the application, or aggregate on a sub-field without a normaliser. Chapter 14 adds such a sub-field in the next index version.
When top-N counts are wrong
What it is. A terms aggregation over several shards asks each shard for its own top terms, shard_size of them, and merges those lists. A term that is just outside the top of every shard never reaches the merge, and a term’s count is under-reported when it is missing from some shards’ lists. The terms aggregation reference documents both the default shard_size, larger than size, and the error bounds.
Why it matters at 300 million profiles. The production index has many shards (chapter 08). For fields with a few dominant values, such as state or status, the merge is reliable. For long-tail fields, such as names or pincodes, where many values have similar counts, a “top 10” can contain wrong values with wrong counts.
Example. The lab index has one shard, so its counts are always exact. Copy the Bihar profiles into a five-shard scratch index and ask for the three most common full names:
PUT ch12-sharded{ "settings": { "number_of_shards": 5, "number_of_replicas": 0 }, "mappings": { "properties": { "fullName": { "type": "keyword" }, "state": { "type": "keyword" } } }}
POST _reindex?refresh=true&filter_path=created,failures{ "source": { "index": "user-profile-read", "_source": ["fullName", "state"], "query": { "term": { "state": "bihar" } } }, "dest": { "index": "ch12-sharded" }}{"created":199755,"failures":[]}GET ch12-sharded/_search?filter_path=aggregations{ "size": 0, "aggs": { "names": { "terms": { "field": "fullName", "size": 3, "show_term_doc_count_error": true } } } }{"aggregations":{"names":{"doc_count_error_upper_bound":415,"sum_other_doc_count":199320, "buckets":[{"key":"Fatima Nair","doc_count":172,"doc_count_error_upper_bound":247}, {"key":"Neha Iyer","doc_count":169,"doc_count_error_upper_bound":250}, {"key":"Mohammed Menon","doc_count":94,"doc_count_error_upper_bound":332}]}}}The true top three in PostgreSQL are Aditi Joshi and Harpreet Reddy with 381 profiles each, and Sneha Das with 375. None of them appears, and the counts shown are less than half the real ones. The response does warn you: an overall error bound of 415 on counts below 200 means the result is unreliable. Raise shard_size so that each shard reports enough terms to cover the long tail:
GET ch12-sharded/_search?filter_path=aggregations{ "size": 0, "aggs": { "names": { "terms": { "field": "fullName", "size": 3, "shard_size": 1000, "show_term_doc_count_error": true } } } }{"aggregations":{"names":{"doc_count_error_upper_bound":0,"sum_other_doc_count":198618, "buckets":[{"key":"Aditi Joshi","doc_count":381,"doc_count_error_upper_bound":0}, {"key":"Harpreet Reddy","doc_count":381,"doc_count_error_upper_bound":0}, {"key":"Sneha Das","doc_count":375,"doc_count_error_upper_bound":0}]}}}Exact, because the field has only 600 distinct values and 1,000 per shard covers them all. On a field with millions of distinct values, no shard_size covers everything, and the cost of each shard’s list grows with it.
Trade-off. Large shard_size values buy accuracy with memory and time on every shard. For facets over low-cardinality fields, the default is fine. For “top N” questions over high-cardinality fields, either accept and display the error bound, or answer the question with an exhaustive composite aggregation offline. Delete the scratch index:
DELETE ch12-shardedCommon mistake. Ignoring doc_count_error_upper_bound. When it is large compared with the counts, the list is not a ranking.
Distinct counts are approximate by design
What it is. The cardinality aggregation estimates the number of distinct values with the HyperLogLog++ algorithm, in bounded memory. precision_threshold trades memory for accuracy: the reference gives the memory as about precision_threshold × 8 bytes and says counts are likely accurate up to the threshold.
Why it matters at 300 million profiles. An exact distinct count over hundreds of millions of values needs memory proportional to the values. An estimate with a known, small error needs a few hundred kilobytes.
Example. Every email address in the lab is unique, so the exact answer is 1,000,000:
GET user-profile-read/_search?filter_path=aggregations{ "size": 0, "aggs": { "distinct_emails": { "cardinality": { "field": "email" } }, "distinct_emails_precise": { "cardinality": { "field": "email", "precision_threshold": 40000 } }, "distinct_cities": { "cardinality": { "field": "city" } } }}{"aggregations":{"distinct_cities":{"value":10},"distinct_emails":{"value":994297},"distinct_emails_precise":{"value":1000182}}}The default estimate is 0.57% low; with precision_threshold: 40000, about 320 KB per shard by the formula, it is 0.018% high. The ten distinct cities are counted exactly, because they are far below either threshold.
Common mistake. Presenting a cardinality result as an exact figure in a report or a reconciliation. Use cardinality for dashboards and trends; use PostgreSQL for numbers that must be exact.
Every bucket, page by page, with composite
What it is. A composite aggregation returns buckets for combinations of sources in a stable order, a page at a time, with an after_key to request the next page. Unlike terms, it is exhaustive and exact.
Why it matters at 300 million profiles. Some questions need every bucket, not the top ones: a nightly count per state and city, or a reconciliation that compares Elasticsearch with PostgreSQL group by group. composite answers them without the top-N error and without loading every bucket into memory at once.
Example. State and city combinations, three per page:
GET user-profile-read/_search?filter_path=aggregations{ "size": 0, "aggs": { "places": { "composite": { "size": 3, "sources": [ { "state": { "terms": { "field": "state" } } }, { "city": { "terms": { "field": "city" } } } ] } } }}{"aggregations":{"places":{"after_key":{"state":"karnataka","city":"bengaluru"}, "buckets":[{"key":{"state":"bihar","city":"gaya"},"doc_count":99646}, {"key":{"state":"bihar","city":"patna"},"doc_count":100109}, {"key":{"state":"karnataka","city":"bengaluru"},"doc_count":100159}]}}}Pass the after_key back as "after": {"state": "karnataka", "city": "bengaluru"} in the same request to get the next page, which starts at kerala / kochi. Stop when a response has no after_key.
Counts over time with date_histogram
A date_histogram buckets documents by calendar interval. Profile updates per month since May 2026:
GET user-profile-read/_search?filter_path=aggregations{ "size": 0, "query": { "range": { "updatedAt": { "gte": "2026-05-01" } } }, "aggs": { "per_month": { "date_histogram": { "field": "updatedAt", "calendar_interval": "month" } } }}2026-05 38662026-06 27162026-07 17232026-08 5792026-09 4The response is condensed to month and count; each bucket also carries key in epoch milliseconds. It matches PostgreSQL’s date_trunc('month', updated_at) counts. The four profiles in September are the ones chapter 04 changed.
Facets that survive a selection
A facet panel has one behaviour users expect and a naive implementation breaks. When a user picks city = patna, the city facet must still show Gaya and its count, so the user can switch; only the other facets should narrow. If the city filter is applied to the city facet, it shows a single bucket, Patna.
Elasticsearch offers two ways to get that behaviour. The first is post_filter, which applies after aggregations are computed:
GET user-profile-read/_search?filter_path=hits.total,aggregations{ "size": 0, "track_total_hits": true, "query": { "bool": { "filter": [ { "term": { "accountStatus": "ACTIVE" } }, { "term": { "state": "bihar" } } ] } }, "post_filter": { "term": { "city": "patna" } }, "aggs": { "city_facet": { "terms": { "field": "city" } } }}{"hits":{"total":{"value":70214,"relation":"eq"}}, "aggregations":{"city_facet":{"doc_count_error_upper_bound":0,"sum_other_doc_count":0, "buckets":[{"key":"patna","doc_count":70214},{"key":"gaya","doc_count":69711}]}}}The hits are Patna only; the facet still shows Gaya. post_filter handles one selected facet. With several facet fields, each facet must ignore its own filter and apply all the others, which is what filter aggregations do, one per facet. The Search API uses that pattern.
Stage 1 — Facets in the Search API
Facet filters and the other filters must now be available separately. In UserQueryBuilder.kt, replace the filters function with these three:
/** Exact constraints: filter context, no scoring, cacheable. */ fun filters(request: UserSearchRequest): List<Query> = baseFilters(request) + facetFilters(request).values
/** Constraints that are never faceted. */ fun baseFilters(request: UserSearchRequest): List<Query> = buildList { add(term("accountStatus", request.accountStatus.name)) request.pincode?.let { add(term("pincode", it)) } request.updatedSince?.let { add(updatedAtFrom(it.toString())) } // Rounded to the day, so the same filter is reused, and cached, for a whole day. request.updatedWithinDays?.let { add(updatedAtFrom("now-${it}d/d")) } }
/** Constraints on faceted fields, keyed by field name (chapter 12). */ fun facetFilters(request: UserSearchRequest): Map<String, Query> = buildMap { request.state?.let { put("state", term("state", it)) } request.city?.let { put("city", term("city", it)) } request.gender?.let { put("gender", term("gender", it.name)) } }filters returns the same set of clauses as before, so /search is unchanged. Add the facets request builder:
package `in`.o612.eng.usersearch.api.search
import co.elastic.clients.elasticsearch._types.aggregations.Aggregationimport co.elastic.clients.elasticsearch._types.query_dsl.Queryimport co.elastic.clients.elasticsearch.core.SearchRequestimport `in`.o612.eng.usersearch.api.web.UserSearchRequestimport `in`.o612.eng.usersearch.index.UserProfileIndex
/** * Facet counts for a search. Each facet ignores its own filter and applies all the others, * so a user who picked city=patna still sees how many matches Gaya would give ("disjunctive" facets). */object UserFacetsBuilder {
val FACET_FIELDS = listOf("state", "city", "gender")
private const val BUCKETS_PER_FACET = 50
fun build(request: UserSearchRequest): SearchRequest = SearchRequest.of { s -> s.index(UserProfileIndex.READ_ALIAS) .size(0) .query(baseQuery(request)) .aggregations(FACET_FIELDS.associateWith { field -> facet(field, request) }) }
/** The name clause and the non-faceted filters: what every facet has in common. */ private fun baseQuery(request: UserSearchRequest): Query = Query.of { q -> q.bool { b -> request.name?.let { b.must(UserQueryBuilder.nameQuery(it, request.nameMode)) } UserQueryBuilder.baseFilters(request).forEach { b.filter(it) } b } }
/** A filter aggregation with every other facet's filter, wrapping a terms aggregation on [field]. */ private fun facet(field: String, request: UserSearchRequest): Aggregation { val otherFilters = UserQueryBuilder.facetFilters(request).filterKeys { it != field }.values.toList() return Aggregation.of { a -> a.filter { f -> f.bool { b -> b.filter(otherFilters) } } .aggregations("values", Aggregation.of { t -> t.terms { tt -> tt.field(field).size(BUCKETS_PER_FACET) } }) } }}Each facet is a filter aggregation that applies every other facet’s filter, wrapping a terms aggregation on the facet’s own field. The top-level query carries only what all facets share: the name clause and the non-faceted filters. The facet fields have a handful of values each, so 50 buckets per facet returns all of them.
Add the response types at the end of SearchDtos.kt:
/** Counts per value for each faceted field. Keys are the indexed, normalised values (chapter 12). */data class FacetsResponse(val facets: Map<String, List<FacetBucket>>)
data class FacetBucket(val value: String, val count: Long)In UserSearchService, add the facets method, and import FacetBucket and FacetsResponse:
fun facets(request: UserSearchRequest): FacetsResponse { val response = client.search(UserFacetsBuilder.build(request), Void::class.java) val facets = UserFacetsBuilder.FACET_FIELDS.associateWith { field -> response.aggregations()[field]!!.filter().aggregations()["values"]!!.sterms().buckets().array() .map { FacetBucket(value = it.key().stringValue(), count = it.docCount()) } } return FacetsResponse(facets) }The typed client exposes each aggregation result through a variant accessor: filter() for the filter aggregation, then sterms() for a terms aggregation on a string field. And in UserSearchController, add the endpoint before the /{userId} mapping:
@GetMapping("/facets") fun facets(@Valid request: UserSearchRequest): FacetsResponse = service.facets(request)The facets endpoint takes exactly the same parameters as /search, so a UI sends the same query string to both.
Stage 2 — Verify the facets
Restart search-api, then request facets with a state and a city selected:
curl -s 'localhost:8080/api/users/facets?state=bihar&city=patna'{ "facets": { "state": [ { "value": "bihar", "count": 70214 } ], "city": [ { "value": "patna", "count": 70214 }, { "value": "gaya", "count": 69711 } ], "gender": [ { "value": "FEMALE", "count": 34752 }, { "value": "MALE", "count": 34707 }, { "value": "OTHER", "count": 755 } ] }}The response is reformatted here. Each facet behaves as designed. city ignores the city selection, so Gaya stays visible with its count. state ignores the state selection but applies city=patna, so only Bihar remains. gender applies both. PostgreSQL agrees on every number, for example 34,752 active female profiles in Patna.
With a name, the facets describe the matches of the search. facets?name=prashant&state=bihar returns 4,715 for Bihar, 4,628 for Maharashtra, and 2,401 for Rajasthan in the state facet, the active Prashants per state, and 2,363 for Patna and 2,352 for Gaya in the city facet. All match PostgreSQL.
Checkpoint
/api/users/facets?state=bihar&city=patna returns two city buckets, not one. If it returns only Patna, the filterKeys { it != field } line in UserFacetsBuilder is missing.
Global ordinals: the hidden cost of terms on keywords
A terms aggregation on a keyword field works on global ordinals, a per-shard mapping from every distinct value to a number. Elasticsearch builds them lazily, on the first aggregation after a refresh that changed the shard, so on a busy index that first facet request after each refresh pays the build cost. For the low-cardinality facet fields here, the cost is small. For high-cardinality fields that are aggregated often, the eager_global_ordinals mapping option moves the build to refresh time, so indexing pays instead of the first search. Chapter 15 comes back to this trade-off with measurements.
Common mistakes with aggregations
- Aggregating on
textfields. Usekeywordsub-fields; chapter 06 showed the alternative. - Treating
termsresults over many shards as exact. Readdoc_count_error_upper_bound. - Treating
cardinalityas exact. It is an estimate with a tunable error. - Using
termswith a hugesizeto “get everything”. Usecompositeand page through. - Applying a facet’s own filter to that facet. Users lose the ability to switch values.
- Showing normalised keys as display names.
maharashtrais an index value, not a label.
What you built, and what comes next
You ran the aggregations profile search needs, terms, cardinality, composite, and date_histogram, and checked every count against PostgreSQL. You reproduced a top-N aggregation that returned the wrong names across five shards, and fixed it with shard_size where the field’s cardinality allowed. The Search API now has a facets endpoint that keeps each facet’s alternatives visible after a selection.
The facet values are still normalised keys. Chapter 14’s next index version adds display sub-fields for them.
Chapter 13 adds cursor pagination: why deep from and size is expensive, search_after with a point in time for stable pages, when scroll still has a place, and a cursor-based contract for the Search API.