Elasticsearch for relational developers
This chapter builds the mental model the rest of the series depends on. You create a small scratch index in the chapter 01 lab and use it to watch each core concept behave: a cluster turning yellow because a replica has nowhere to go, a document you can fetch but not yet find, an update that leaves a deleted document behind in an immutable segment, and a mapping that refuses to change. Each concept follows the same pattern: what it is, why it matters at 300 million profiles, a request you run, and the mistake people make with it.
You need the lab from chapter 01 running (docker compose up -d in user-search-lab/). The chapter assumes you know PostgreSQL well, so comparisons point there. It takes about 30 minutes, and the scratch index is deleted at the end.
How to send requests in this chapter
Every request is written in Kibana Dev Tools syntax: a method, a path, and an optional JSON body. Open http://localhost:5601, go to Dev Tools, paste the request, and run it.
GET _cluster/healthThe same request works from a shell. Load the passwords from .env first, as in chapter 01:
set -a; source .env; set +acurl -s -u "elastic:$ELASTIC_PASSWORD" "localhost:9200/_cluster/health?pretty"For requests with a body, add -X PUT (or the method shown), -H 'Content-Type: application/json', and -d '<body>'. The responses in this chapter were captured from the lab and trimmed with the filter_path parameter, which removes response fields you don’t need. Values such as node names and UUIDs will differ on your machine.
Cluster and node: the unit of operation
What it is. A node is one running Elasticsearch process. A cluster is a set of nodes that share one cluster state: the list of indices, their settings and mappings, and which node holds which shard. One node is the elected master and publishes that state to the others.
Why it matters at 300 million profiles. No single node will hold and serve the whole dataset with headroom for failure, so production means several nodes, and the cluster’s job is to spread data and work across them. In PostgreSQL, one primary owns all writes and replicas are whole copies. In Elasticsearch, every data node owns a slice of the data.
Example. Ask the lab’s cluster about itself:
GET _cluster/health{ "cluster_name" : "docker-cluster", "status" : "green", "number_of_nodes" : 1, "number_of_data_nodes" : 1, "active_primary_shards" : 51, "active_shards" : 51, "unassigned_shards" : 0, ...}Fifty-one shards already exist, even though you have created nothing. Kibana and Elasticsearch store their own state in system indices, and those count too. Now list the nodes:
GET _cat/nodes?v&h=name,node.role,heap.max,mastername node.role heap.max master4483bb3f14e5 cdfhilmrstw 1gb *The one node carries every node role, abbreviated to one letter each, including m (master-eligible) and d (data). The * marks it as the elected master.
Common mistake. Treating the lab’s layout as a template. A production cluster usually separates roles, with three small dedicated master-eligible nodes, so that a data node under memory pressure cannot destabilise cluster coordination. Chapter 08 covers sizing and chapter 16 covers operations.
Index: a logical namespace, not a storage unit
What it is. An index is a named collection of documents that share settings and a mapping. It is what you query. Physically, an index is split into one or more shards, and each shard is a complete Apache Lucene index that lives on one node.
Why it matters at 300 million profiles. The index is the boundary for settings you cannot change later, such as the number of primary shards, and for the mapping, which you cannot change incompatibly. Chapter 05 uses versioned index names and aliases because of exactly those two limits.
Example. Create a scratch index with two primary shards, one replica of each, and automatic refresh turned off. You turn it off only so that the refresh experiment later in this chapter is predictable.
PUT ch02-profiles{ "settings": { "number_of_shards": 2, "number_of_replicas": 1, "refresh_interval": "-1" }}{"acknowledged":true,"shards_acknowledged":true,"index":"ch02-profiles"}Common mistake. Creating one index per tenant, per city, or per day “for organisation”. Each index brings at least one shard, and every shard has a fixed memory and coordination cost. Split data into indices only when the lifecycle or settings genuinely differ.
Shards: primaries, replicas, and why the cluster turned yellow
What it is. A primary shard holds a subset of the index’s documents and accepts writes for them. A replica shard is a copy of a primary on a different node. It provides redundancy and can serve searches. The clusters, nodes, and shards overview describes the full model.
The diagram shows a two-node cluster holding an index with two primaries and one replica each. Node 1 holds primary 0 and the replica of shard 1. Node 2 holds primary 1 and the replica of shard 0. Elasticsearch never places a replica on the same node as its primary, so either node can fail without losing data.
Why it matters at 300 million profiles. Primaries divide the data, so they decide how large each shard gets and how many nodes can share indexing work. Replicas multiply storage, and they add search capacity and failure tolerance. Chapter 08 turns both into a sizing calculation.
Example. Check the scratch index’s health and shard placement:
GET _cluster/health/ch02-profiles?filter_path=status,active_primary_shards,active_shards,unassigned_shards{"status":"yellow","active_primary_shards":2,"active_shards":2,"unassigned_shards":2}GET _cat/shards/ch02-profiles?v&h=index,shard,prirep,state,docs,nodeindex shard prirep state docs nodech02-profiles 0 p STARTED 0 4483bb3f14e5ch02-profiles 0 r UNASSIGNEDch02-profiles 1 p STARTED 0 4483bb3f14e5ch02-profiles 1 r UNASSIGNEDBoth primaries (p) started. Both replicas (r) are unassigned, because the only node already holds their primaries. That is what yellow means: all data is available, but not all copies exist. red would mean at least one primary is unassigned, so some data is unavailable.
The replica count is a dynamic setting. Set it to zero for the lab:
PUT ch02-profiles/_settings{ "index": { "number_of_replicas": 0 } }GET _cluster/health/ch02-profiles?filter_path=status,unassigned_shards{"status":"green","unassigned_shards":0}The primary count is not dynamic. Try to change it:
PUT ch02-profiles/_settings{ "index": { "number_of_shards": 4 } }illegal_argument_exception: Can't update non dynamic setting(s) [[index.number_of_shards]] for open indices[[ch02-profiles/...]]. The setting(s) [[index.number_of_shards]] cannot be modified on an index once it iscreated. You will need to create a new index with the desired setting(s) and reindex your data.Principle. The primary shard count is fixed when an index is created. The replica count can change at any time. The index settings reference lists which settings are static and which are dynamic. (The split and shrink APIs can change the primary count by creating a new index from an existing one, with restrictions; they are not an in-place change.)
Common mistake. Reading yellow as “broken”. On a single-node lab it is the normal state for any index with replicas. In a multi-node production cluster, persistent yellow is a real problem, because you have lost failure tolerance without losing data yet.
Documents and routing: how a row finds its shard
What it is. A document is a JSON object stored in an index, identified by an _id. Elasticsearch stores the original JSON as _source and also indexes each field according to the mapping. To decide which primary shard receives a document, Elasticsearch hashes its routing value, which is the _id unless you supply one.
Why it matters at 300 million profiles. Routing is why the primary count is fixed: change the number of shards and the hash sends existing documents to different shards. It also explains why a lookup by _id goes straight to one shard, while a search goes to every shard. For this series, _id is always the PostgreSQL user_id, so re-indexing a profile overwrites the same document instead of creating a duplicate.
Example. Index one profile with an explicit ID:
PUT ch02-profiles/_doc/1{ "userId": 1, "fullName": "Prashant Kumar", "city": "Patna", "state": "Bihar", "accountStatus": "ACTIVE", "updatedAt": "2026-09-01T10:15:00Z"}{"_index":"ch02-profiles","_id":"1","_version":1,"result":"created", "_shards":{"total":1,"successful":1,"failed":0},"_seq_no":0,"_primary_term":1}_version, _seq_no, and _primary_term are the concurrency metadata that chapter 03 uses for optimistic locking and idempotent writes.
Common mistake. Letting Elasticsearch generate IDs for a projection of a relational table. Auto-generated IDs index faster, because Elasticsearch can skip checking whether the ID already exists, as the indexing speed guide notes. But every retry or replay then creates a new document. A projection needs the source’s primary key as _id.
Refresh and near-real-time search
What it is. An indexed document is not immediately searchable. It first goes into an in-memory buffer on the shard and is recorded in the translog, the shard’s write-ahead log. A refresh turns the buffer into a new searchable segment. Until then, search cannot see the document. That is what Elasticsearch means by near-real-time search. By default, refresh happens every second on indices that have received a search recently.
Why it matters at 300 million profiles. Refresh is a trade-off between freshness and indexing cost. Every refresh creates a small segment that must later be merged. During a bulk backfill of hundreds of millions of documents, refreshing every second wastes work; in chapter 14 you disable it for the backfill and restore it afterwards. It also sets the minimum staleness of the projection, on top of whatever delay the sync pipeline adds.
Example. The scratch index has refresh_interval: -1, so nothing refreshes automatically. Document 1 has been indexed but not refreshed. Getting it by ID works:
GET ch02-profiles/_doc/1?filter_path=_id,_version,found,_source.fullName{"_id":"1","_version":1,"found":true,"_source":{"fullName":"Prashant Kumar"}}A search for it does not:
GET ch02-profiles/_search?filter_path=hits.total{ "query": { "match": { "fullName": "prashant" } } }{"hits":{"total":{"value":0,"relation":"eq"}}}Refresh explicitly, then run the same search again:
POST ch02-profiles/_refresh{"_shards":{"total":2,"successful":2,"failed":0}}GET ch02-profiles/_search?filter_path=hits.total,hits.hits._id{ "query": { "match": { "fullName": "prashant" } } }{"hits":{"total":{"value":1,"relation":"eq"},"hits":[{"_id":"1"}]}}A get by ID is real time because it can read the latest version before a refresh, as the reading and writing documents guide describes. Search only sees refreshed segments.
Add four more profiles in one _bulk request. The body is newline-delimited JSON: an action line followed by a document line, and the request must end with a newline.
POST ch02-profiles/_bulk?refresh=true{"index":{"_id":"2"}}{"userId":2,"fullName":"Priya Sharma","city":"Mumbai","state":"Maharashtra","accountStatus":"ACTIVE","updatedAt":"2026-08-11T09:00:00Z"}{"index":{"_id":"3"}}{"userId":3,"fullName":"José Fernandes","city":"Kochi","state":"Kerala","accountStatus":"INACTIVE","updatedAt":"2026-07-02T12:30:00Z"}{"index":{"_id":"4"}}{"userId":4,"fullName":"Prashant Jha","city":"Gaya","state":"Bihar","accountStatus":"ACTIVE","updatedAt":"2026-09-20T08:45:00Z"}{"index":{"_id":"5"}}{"userId":5,"fullName":"Anjali Nair","city":"Kochi","state":"Kerala","accountStatus":"SUSPENDED","updatedAt":"2026-06-15T17:05:00Z"}{"errors":false,"items":[{"index":{"result":"created"}},{"index":{"result":"created"}}, {"index":{"result":"created"}},{"index":{"result":"created"}}]}The ?refresh=true parameter refreshes the affected shards before the response returns, so these documents are searchable immediately. Now look at where the five documents landed:
GET _cat/shards/ch02-profiles?v&h=index,shard,prirep,state,docsindex shard prirep state docsch02-profiles 0 p STARTED 1ch02-profiles 1 p STARTED 4This is routing at work. The hash is deterministic, so you get the same split. It is uniform over many documents and arbitrary over five: with millions of IDs, each primary gets a nearly equal share.
The diagram shows the write path inside one shard. A write goes to the in-memory buffer and to the translog. A refresh turns the buffer into a searchable segment. A flush makes segments durable as a Lucene commit, after which the translog can be trimmed. Background merges combine small segments into larger ones. The translog is what makes an acknowledged write survive a crash between refresh and flush; the translog settings control how often it is synced to disk.
Production note — The one-second default has a condition. If an index has no explicit
refresh_interval, shards that have not received a search forindex.search.idle.after(30 seconds by default) skip background refreshes until the next search arrives, as the index settings reference explains. The first search after a quiet period can wait for that refresh. Setrefresh_intervalexplicitly when freshness matters more than indexing throughput.
Common mistake. Adding ?refresh=true to every write so that tests or the UI can read their own writes. Each forced refresh creates a tiny segment, and on a busy index this causes heavy merge load. Use refresh=wait_for where a caller truly needs to wait, and design the rest of the system to tolerate a second of lag. The refresh parameter reference compares the options.
Segments: immutable files, and what an update really is
What it is. A segment is an immutable, self-contained Lucene index file set inside a shard. Refresh creates segments, and background merges combine them. Because segments are never modified in place, a deleted document is only marked deleted until a merge rewrites its segment without it.
Why it matters at 300 million profiles. Profiles change constantly. Every update in Elasticsearch writes a complete new copy of the document and marks the old one deleted. With a high update rate, deleted documents and merge work become part of the steady-state cost, which is one reason chapter 04 cares about sending only real changes.
Example. List the segments:
GET _cat/segments/ch02-profiles?v&h=shard,segment,docs.count,docs.deleted,sizeshard segment docs.count docs.deleted size0 _0 1 0 7.8kb1 _0 1 0 7.8kb1 _1 3 0 8.4kbShard 1’s segment _0 came from the first refresh and holds document 1. Segment _1 came from the bulk request. Now update one field of document 1:
POST ch02-profiles/_update/1?refresh=true&filter_path=_version,result{ "doc": { "city": "Gaya", "updatedAt": "2026-09-25T11:00:00Z" } }{"_version":2,"result":"updated"}shard segment docs.count docs.deleted size0 _0 1 0 7.8kb1 _0 0 1 10kb1 _1 3 0 8.4kb1 _2 1 0 7.8kbThe partial update rewrote the whole document into a new segment, _2. Segment _0 still exists, now with zero live documents and one deleted one, and it will stay on disk until a merge removes it. The merge settings control that process.
Common mistake. Assuming a partial update is cheaper than a full reindex of the document. The _update API saves network bandwidth, not indexing work: Elasticsearch reads the current _source, applies the change, and indexes the whole result.
Mappings: a schema with search behaviour
What it is. A mapping declares each field’s type, and for text, how it is analysed into searchable terms. It is the closest thing to a table’s DDL, but it decides more than storage: it decides which queries can match a value at all.
Why it matters at 300 million profiles. A mapping mistake is expensive to fix, because most changes need a full reindex, and at this scale that is hours of work and a second copy of the data (chapter 14). It also decides index size: every field you map is indexed, and every text field produces an inverted index.
Example. You never defined a mapping for the scratch index, so Elasticsearch inferred one from the first documents, using dynamic mapping:
GET ch02-profiles/_mapping{ "ch02-profiles" : { "mappings" : { "properties" : { "accountStatus" : { "type" : "text", "fields" : { "keyword" : { "type" : "keyword", "ignore_above" : 256 } } }, "city" : { "type" : "text", "fields" : { "keyword" : { "type" : "keyword", "ignore_above" : 256 } } }, ... "updatedAt" : { "type" : "date" }, "userId" : { "type" : "long" } } } }}Every string became two fields: a text field, analysed into lowercase words for full-text search, and a keyword sub-field that holds the exact value. That default is a guess, and for most profile fields it is wrong both ways: accountStatus never needs full-text search, and fullName needs more careful analysis than the default. See what the guess does to an exact filter:
GET ch02-profiles/_search?filter_path=hits.total{ "query": { "term": { "state": "Bihar" } } }{"hits":{"total":{"value":0,"relation":"eq"}}}GET ch02-profiles/_search?filter_path=hits.total{ "query": { "term": { "state.keyword": "Bihar" } } }{"hits":{"total":{"value":2,"relation":"eq"}}}The text field stored the term bihar in lowercase, so an exact term query for Bihar finds nothing. Chapter 06 explains analysis in full.
Now try to fix the mapping in place:
PUT ch02-profiles/_mapping{ "properties": { "city": { "type": "keyword" } } }illegal_argument_exception: mapper [city] cannot be changed from type [text] to [keyword]Once a field is mapped, its type is fixed for the life of the index. The mapping also rejects documents that contradict it:
PUT ch02-profiles/_doc/6{ "userId": "not-a-number", "fullName": "Rahul Verma" }document_parsing_exception: [1:11] failed to parse field [userId] of type [long] in document with id '6'.Preview of field's value: 'not-a-number'And it silently accepts fields it has never seen:
PUT ch02-profiles/_doc/7{ "userId": 7, "fullName": "Meera Iyer", "nickname": "Meeru" }GET ch02-profiles/_mapping/field/nickname{"ch02-profiles":{"mappings":{"nickname":{"full_name":"nickname","mapping":{"nickname":{"type":"text", "fields":{"keyword":{"type":"keyword","ignore_above":256}}}}}}}}Two new fields, nickname and nickname.keyword, now exist in the index permanently, created by one document. In a projection fed by an event stream, one upstream change can add fields you never designed. Chapter 05 sets dynamic: strict to reject such documents instead.
Common mistake. Shipping with dynamic mapping because “it worked in dev”. It maps every string twice, guesses types from the first value it sees, and lets unknown fields grow the mapping without review. That growth is called mapping explosion, and chapter 06 shows how to prevent it.
Where “an index is a table” breaks
The analogy is useful for the first hour and misleading after that. This table is the reference for the rest of the series.
| Concept | PostgreSQL table | Elasticsearch index |
|---|---|---|
| Storage layout | One heap, plus separate indexes you choose | Shards of Lucene segments; every mapped field is indexed by default |
| Schema change | ALTER TABLE changes most things in place | New fields can be added; changing an existing field needs a new index and a reindex |
| Write visibility | Visible to others at commit | Searchable after the next refresh, typically within a second |
| Transactions | Multi-row, multi-table ACID | Single-document atomicity only; no multi-document transactions |
| Joins | Core feature | None in the Query DSL; ES|QL LOOKUP JOIN enriches rows from a special lookup index, but it is not a general join. Denormalise instead |
| Updates | In place, with MVCC row versions | Always a new copy of the whole document plus a delete marker |
| Uniqueness | UNIQUE constraints on any column | Only _id is unique; nothing else can be enforced |
| Scaling reads | Replicas each hold a full copy | Shards spread data; replicas copy shards |
| Query answer | Exact set of matching rows | Exact filters, or a relevance-ranked list; totals can be approximate |
Two rows matter most for this series. Without transactions or uniqueness constraints, Elasticsearch cannot be a source of truth for profiles, which is chapter 01’s rule. And because changing a mapping means building a new index, index names must be versioned from day one, which is where chapter 05 starts.
Clean up the scratch index
Warning —
DELETEremoves an index and all its documents permanently. Check the name before you run it. Everything inch02-profileswas created in this chapter.
DELETE ch02-profiles{"acknowledged":true}What you now know, and what comes next
You have seen each core concept behave: a cluster of nodes, an index split into primary and replica shards, documents routed by _id, near-real-time search driven by refresh, immutable segments where an update is a new copy plus a delete marker, and a mapping that behaves like a schema but decides what can match.
This chapter used a single node, so it could not show replica allocation, failover, or shard rebalancing across nodes. Chapter 08 discusses them for capacity planning, and chapter 16 for operations.
Chapter 03 tours the REST API you will use for the rest of the series: document APIs, bulk requests, optimistic concurrency with if_seq_no and if_primary_term, external versioning, and the _cat and cluster APIs for inspecting what the cluster is doing.