Series overview
Part 15 of 1883% complete
2026-05-27•12 min read

Performance optimisation, measured

Elasticsearch performance advice is easy to find and easy to apply without checking. This chapter takes the optimisations that matter for profile search and measures each one on the lab, with the method stated, instead of repeating them. Three results are worth knowing in advance. The node query cache ignores the filters you would most expect it to cache. Index sorting made the “newest first” query faster and the index 59% larger. And disabling refresh during a bulk load, the most common indexing tip, made loading slower in both lab runs.

None of these numbers transfers to production. What transfers is the method: change one thing, measure it with the workload that matters, and keep the change only when the gain is larger than its cost. You need the lab over TLS with user-profile-v2 live (chapter 14) and the Rally track from chapter 08. The chapter takes about 60 minutes.

The measurement rules for this chapter

Every number in this chapter comes from the same environment: one Elasticsearch 9.5.4 node in Docker with a 1 GB heap, on a 6-core desktop CPU, with the load generator on the same machine and one million synthetic profiles. Search timings are Elasticsearch’s own took, in whole milliseconds, over 100 requests after 20 warm-up requests. Indexing throughput comes from Rally’s summary report. Surprising results were repeated before being reported.

Needs validation: every setting in this chapter must be re-measured on production-like hardware, data, and shard counts, as chapter 08’s Stage 3 describes. A setting that helps one node with a 1 GB heap can be neutral or harmful on twelve shards across six nodes.

Indexing: what refresh actually cost in the lab

What the guidance says. Elastic’s indexing speed guide recommends, for large bulk loads, using bulk requests with several workers, setting refresh_interval to -1, and setting replicas to 0. Chapter 05’s loader and chapter 14’s backfill both do the last two.

The measurement. Make the refresh interval a track parameter. In rally/user-profile-track/index.json, change the refresh_interval line to:

rally/user-profile-track/index.json
"refresh_interval": "{{ refresh_interval | default('1s') }}"

The lab now speaks TLS (chapter 10), so the Rally container needs the CA certificate. Run the track twice with each value:

Terminal window
set -a; source .env; set +a
docker run --rm --network host \
-v "$PWD/rally/user-profile-track:/rally/track" \
-v "$PWD/certs/ca/ca.crt:/rally/ca.crt:ro" \
elastic/rally:latest race \
--track-path=/rally/track --target-hosts=localhost:9200 --pipeline=benchmark-only \
--client-options="use_ssl:true,verify_certs:true,ca_certs:'/rally/ca.crt',basic_auth_user:'elastic',basic_auth_password:'${ELASTIC_PASSWORD}'" \
--track-params="refresh_interval:'-1'"
Runrefresh_intervalBulk mean throughputCumulative refresh timeCumulative merge time
11s88,747 docs/s2.7 s77.8 s
21s88,025 docs/s2.9 s80.3 s
3-176,020 docs/s0.6 s68.1 s
4-157,003 docs/s0.5 s41.1 s

Refresh itself cost under three seconds of shard time in a load that took about twelve seconds of wall time, and disabling it lowered throughput in both runs, with more run-to-run variation. The cause was not isolated. With a 1 GB heap, the indexing buffer is small, so segments are written frequently whether or not refresh runs; that makes the saving small, but it does not explain the loss. The honest conclusion is narrow: on this setup, the standard tip did not help, and the backfill’s settings are a hypothesis to measure, not a given.

Principle. A performance recommendation is a starting point for a measurement. The indexing guide’s advice is sound for the loads it describes, typically long loads on larger heaps, where frequent refreshes create many small segments and merges fall behind. Whether your backfill is such a load is something Rally answers in an afternoon.

Common mistake. Copying a tuning setting because it appears in every checklist, and never checking that it moved the metric it was meant to move.

Bulk sizing and concurrency

BulkIngester in chapter 14 sends 1,000 operations or 5 MiB per request, whichever comes first, with two requests in flight. Those are reasonable starting points, not measured optima. Larger bulk requests amortise per-request overhead until they start to cost heap and cause 429 rejections. More concurrent requests help until the write thread pool, six threads on the lab’s node, is saturated. Needs validation: vary bulk-size and clients in the Rally track’s bulk-index task, one at a time, and keep the smallest values that reach the throughput plateau.

Two costs in this design are accepted on purpose. Explicit IDs, the PostgreSQL user_id, are slower to index than auto-generated ones, because Elasticsearch must check whether each ID already exists; a projection needs them for idempotency. External versions add a version check per document for the same reason.

Search: where a query spends its time

What it is. Setting "profile": true on a search returns, for each shard, the time spent in every part of the query tree and in collection. See the profile API reference.

Why it matters at 300 million profiles. A slow search is usually slow because of one clause. The profile shows which one, instead of guessing.

Example. Profile chapter 11’s composed request, the active Prashants in Bihar updated in the last 90 days, by adding "profile": true to its body. Condensed from the response:

Query nodeTimeDescription
BooleanQuery3.13 msThe whole query
TermQuery0.09 msaccountStatus:ACTIVE
TermQuery0.08 msstate:bihar
IndexOrDocValuesQuery1.13 msThe updatedAt range
DisjunctionMaxQuery0.38 msThe cross_fields clause over name, city, and state
TermQuery0.16 msfullName:prashant, from the fuzzy clause

Two things stand out. The date range is the most expensive clause, which is why the next section’s cache matters for it. And the fuzzy clause became a plain TermQuery: Lucene rewrote prashant~AUTO into the terms within edit distance that actually exist in the index, and there were none other than prashant itself. Profiling adds overhead of its own, so compare profiles with each other, not with production latency.

Common mistake. Profiling a query once, on a cold cache, and optimising for that result. Profile a warmed query, and the slow one from the slow log, described later in this chapter.

Search: which filters the query cache stores

What it is. The node query cache stores, per segment, which documents matched a filter, and reuses that result for any later query with the same filter. It holds at most 10,000 queries in up to 10% of the heap, evicting the least recently used.

Why it matters at 300 million profiles. A filter shared by many searches, such as “updated in the last 90 days”, can be computed once per segment and reused by every search that includes it, whatever else those searches ask.

Example. Two rules from the reference decide what is cached, and the lab shows both. Only segments with at least 10,000 documents and at least 3% of the shard’s documents are cached. And term queries are not eligible at all, because they are already cheap. Repeat a search with two term filters six times, and the cache statistics do not move:

GET user-profile-v2/_stats/query_cache?filter_path=_all.total.query_cache
{"hit_count":0,"miss_count":113,"cache_count":0}

Replace one filter with a range, {"range": {"updatedAt": {"gte": "now-365d/d"}}}, and repeat it ten times with the name priya:

{"memory_size_in_bytes":121645,"hit_count":15,"miss_count":158,"cache_count":3}

The range filter is now cached for the index’s three eligible segments. Then search for rahul with the same range filter, three times:

{"memory_size_in_bytes":121645,"hit_count":24,"miss_count":167,"cache_count":3}

Nine more hits: three searches times three segments. A different search reused the cached filter. That reuse is only possible because the filter’s value is identical across requests, which is why chapter 11 rounds date math to the day with now-90d/d. An unrounded now-90d produces a new value every millisecond and never repeats.

Common mistake. Expecting filter to mean “cached”. It means “not scored, and eligible for caching if it is expensive, reused, and applied to large enough segments”.

Search: the request cache for facet counts

The shard request cache stores whole shard-level results of size: 0 requests, such as chapter 12’s facet counts, keyed by the request body. Run the same facet request three times after clearing the caches, and the statistics show two hits:

{"request_cache":{"hit_count":2,"miss_count":3}}

It is invalidated when a refresh brings in changed documents, so on an index that the relay updates every second, a cached facet count lives until the next change reaches that shard. It still pays for bursts of identical requests, such as a facet panel loaded by many users at once. Chapter 07 showed its one trap: a synonym change is not a data change, so counts need POST <index>/_cache/clear?request=true after one.

Index sorting: faster in one query, larger on disk

What it is. Index sorting stores documents in each segment in a fixed order, chosen at index creation. A search sorted the same way can stop early, once it has collected enough hits.

Why it matters at 300 million profiles. “Newest profiles first”, with filters, is a common admin query. Without index sorting, Elasticsearch must consider every matching document’s updatedAt value to find the newest, although Lucene already skips efficiently on numeric sorts.

Example. Build two copies of user-profile-v2: one from its definition unchanged, and one with these two settings added under settings.index:

"sort.field": ["updatedAt", "userId"],
"sort.order": ["desc", "asc"]

Fill both with _reindex from user-profile-v2, force-merge each to one segment, and run “active profiles, newest first”, size: 20, 100 times against each:

Indexp50p95MeanStore size
Unsorted2 ms3 ms2.06 ms155.7 MB
Sorted by updatedAt, userId0 ms1 ms0.11 ms248.1 MB

Both return the same hits in the same order. The sorted index answered in a fraction of a millisecond against about two milliseconds for the unsorted one, and it is 59% larger. took has millisecond resolution, so read the difference as “about two milliseconds”, not as a ratio. The size increase is plausibly compression: documents ordered by userId sit next to similar neighbours in the lab data, and ordering them by update time breaks that locality. Index sorting also adds work at indexing time, and it cannot be changed without a reindex. Delete both copies when you are done.

Trade-off. Keep index sorting off by default for this index. Its benefit is limited to queries sorted exactly like the index; relevance-sorted searches, which are most of the Search API’s traffic, gain nothing. Consider it only when one field-sorted query dominates, and when a benchmark at production shard sizes shows the latency gain is worth the disk and indexing cost. In the lab, the gain was about two milliseconds on one million documents.

Other settings, and when they apply

Setting or practiceEffectCostUse when
_source filtering (chapter 10)Less data read and sent per hitNoneAlways, for results
Date math rounded to the dayFilters become cacheable and reusableCoarser windowsAny relative date filter
eager_global_ordinals on facet fieldsThe first facet request after a refresh does not pay to build ordinalsRefresh does the work insteadHigh-cardinality fields aggregated constantly; measure first
Force merge to one segmentFewer, larger segments; faster searchesHeavy I/O; harmful on indices still writtenRead-only indices only
Avoid leading wildcards and scripts in queriesPredictable latencyLess flexible queriesAlways, for user-facing search
More replicasMore concurrent search capacityDisk and indexing work per replicaWhen search, not indexing, is the bottleneck (chapter 08)

The force merge API warns against its use on an index that is still being written, because it can produce segments larger than 5 GB that ordinary merges then ignore. user-profile-v2 receives changes from the relay every second, so it is never a candidate. The search speed guide covers further techniques.

Diagnosis tools: slow log and hot threads

The search slow log records every search slower than a threshold, per index. See the slow log settings. On a scratch index, set a threshold of zero so everything is logged, and run a leading-wildcard query:

PUT ch15-unsorted/_settings
{ "index.search.slowlog.threshold.query.warn": "0ms", "index.search.slowlog.include.user": true }
GET ch15-unsorted/_search?size=1
{ "query": { "wildcard": { "fullName.keyword": { "value": "*kumar" } } } }

In docker compose logs elasticsearch, condensed:

{"log.level": "WARN", "log.logger":"index.search.slowlog.query",
"elasticsearch.slowlog.source":"{\"size\":1,\"query\":{\"wildcard\":{\"fullName.keyword\":{\"wildcard\":\"*kumar\",...}}}}",
"elasticsearch.slowlog.took":"6.6ms", "elasticsearch.slowlog.total_hits":"10090+ hits",
"user.name":"elastic", ...}

The entry names the index and shard, the full query, its time, its hit count, and, with include.user, who ran it. In production, set thresholds at the latency you care about, for example warn at your p99 target, and review the log regularly.

Security note — The slow log records the full query source. For a profile search, that means names, email addresses, and mobile numbers typed by agents, now in a log file. Restrict access to Elasticsearch logs as you would to the data, set their retention, and consider the log in your erasure process. Chapter 16 returns to logs as a store of personal data.

Hot threads. GET _nodes/hot_threads samples each node’s busiest threads over a short interval and prints their stack traces. On an idle lab it shows nothing of interest; during a slow period in production, it answers “what is this node busy doing?” in one request. See the hot threads API.

Remove the scratch index when you finish:

DELETE ch15-unsorted,ch15-sorted

What you measured, and what comes next

You measured refresh during bulk loads, the node query cache, the request cache, and index sorting on the lab, and profiled the Search API’s composed query clause by clause. Two widely repeated recommendations, “disable refresh during loads” and “filters are cached”, turned out to be conditional in ways that only measurement shows. Index sorting was faster for one query shape at a clear cost in disk.

These are single-node, one-million-document results. They show how to test each setting and how large an effect to look for; production values come from repeating the tests at production scale.

Chapter 16 covers operating the cluster: health states, the monitoring signals that predict trouble, snapshots, security with least-privilege API keys and TLS, and the personal-data obligations of a profile index, including erasure across index versions, snapshots, and logs.

ElasticsearchPerformance

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind