Searching 300 million user profiles with Elasticsearch
Build user search over 300 million PostgreSQL profiles with Elasticsearch 9 and Spring Boot 4 in Kotlin: mappings, queries, sync, sizing, and operations.
Chapters
The search problem at 300 million profiles
Where PostgreSQL indexes stop fitting flexible user search at 300 million profiles, and where they are still enough, shown with EXPLAIN on a local lab.
Elasticsearch for relational developers
Clusters, shards, replicas, segments, refresh, and mappings, each shown in a live lab and compared with PostgreSQL, including where the table analogy breaks.
The REST API surface, from documents to the cluster
Document, bulk, by-query, task, and inspection APIs in a live lab, with optimistic locking, external versions, and two version traps a sync pipeline must avoid.
Search architecture and keeping Elasticsearch in sync
Why dual writes to PostgreSQL and Elasticsearch fail, how outbox and CDC compare, and an idempotent sync design with row versions, re-reads, and safe deletes.
Designing the user-profile index
A versioned user-profile index behind read and write aliases, a field capability matrix, the full creation request, and a million profiles loaded and checked.
Mapping and analysis in depth
How text and keyword fields, analysers, filters, and normalisers decide what matches, shown with _analyze, with fielddata and mapping explosion reproduced.
Autocomplete, typos, synonyms, and name relevance
edge_ngram, search_as_you_type, and completion compared on a million names, plus fuzziness, live synonym updates, a request-cache trap, and BM25 via _explain.
Shards, settings, and capacity planning for 300 million profiles
Primaries, replicas, refresh, indexing buffer, and translog explained, then a shard-sizing worksheet built from lab measurements and a Rally track you can run.
Choosing an Elasticsearch client for Kotlin and Spring Boot
The Java API Client, its async and low-level variants, and Spring Data Elasticsearch, run from Kotlin, with the JSON-mapper and 10,000-total traps explained.
Building the Spring Boot search service
A Kotlin Spring Boot 4 Search API over TLS: index bootstrap from reviewed JSON, request and response DTOs, exact lookups, and clean failure handling.
Query design for profile search
Query and filter context, bool clauses, cross_fields and fuzzy name matching, highlighting, and the "active Prashants in Bihar" query in JSON and Kotlin.
Aggregations and facet counts
Terms, cardinality, composite, and date histogram aggregations checked against PostgreSQL, why top-N goes wrong across shards, and disjunctive facets in Kotlin.
Pagination for large result sets
Why deep from and size fails, search_after cursors, a point in time for stable pages, where scroll still fits, and an opaque cursor contract for the Search API.
Backfill, the outbox relay, and zero-downtime reindexing
A Kotlin indexer with keyset backfill, BulkIngester backpressure, per-item errors, an outbox relay, and a v1 to v2 reindex ending in an atomic alias swap.
Performance optimisation, measured
Refresh, caches, index sorting, the profile API, and the slow log, each measured on the lab, including one standard tip that made indexing slower here.
Operating it safely: monitoring, snapshots, security, and personal data
Monitoring signals, snapshots with retention, least-privilege API keys, and personal-data controls, including a facet leak that bypassed authorisation.
Failure modes, and the tests that catch them
Every failure mode found while building this series, from resurrected profiles to leaking facets, and the unit and Testcontainers tests that catch them.
From lab to production: rollout phases and checklists
A five-phase rollout with exit criteria, a relevance gate with _rank_eval, an automated readiness check, the questions to answer first, and what to build next.