Failure modes, and the tests that catch them
This series has met a lot of failure modes, and most of them were found by running the system, not by reading about it. This chapter collects them into one reference: symptom, cause, fix, and the chapter where each appeared. Then it builds the test suite that keeps them fixed. There are unit tests for the query and facet builders, which run in milliseconds and need no cluster, and Testcontainers integration tests that run the real Search API against a real, TLS-enabled Elasticsearch 9.5.4. Building the suite found four more problems, including one in the application code that the tests exist to protect.
You need the user-search project from chapter 16 and Docker. The first test run downloads nothing new if the lab’s Elasticsearch image is present. The chapter takes about 60 minutes.
The failure modes, in one table
The first group comes from the brief this series was written against, and the second from building it.
| Failure mode | Symptom | Cause | Fix | Chapter |
|---|---|---|---|---|
| Elasticsearch treated as the source of truth | Data that exists only in the index; drift nobody can repair | Writes or fixes applied to Elasticsearch directly | Write only to PostgreSQL; rebuild the projection from it | 01, 04 |
| Careless dynamic mapping | Unexpected fields; mapping growth; the field limit error | dynamic left at its default | dynamic: strict, false, or flattened | 02, 06 |
Everything mapped as text | Filters behave like word search; aggregations fail | Default string mapping | Keyword for identifiers and filters; multi-fields for names | 05, 06 |
match on keyword fields, term on text fields | “The profile exists but search can’t find it” | Query and index disagree about terms | Decide query types per field; check with _analyze | 06 |
fielddata enabled on text | Word-level buckets; heap exhaustion; breaker trips | Following the error message’s suggestion | A keyword sub-field | 06 |
| Shard count by folklore | Oversized or overly numerous shards | No measurement | The sizing worksheet and a Rally benchmark | 08 |
Deep from/size | Result window is too large; slow deep pages | Offset paging | search_after cursors, with a PIT when stability matters | 13 |
| Refresh ignored during bulk loads | Slow loads, or writes that never become visible | Refresh left at -1, or never considered | Measure; always restore settings in finally | 08, 14, 15 |
| Dual writes with no recovery path | Silent drift after crashes | Writing PostgreSQL and Elasticsearch in one handler | Transactional outbox or CDC | 04 |
| Delete propagation and reindexing forgotten | Deleted profiles in search results; no way to change mappings | No delete events; no versioned index | Delete events through the outbox; versioned indices behind aliases | 04, 05, 14 |
Sensitive fields returned from _source | Contact details in result lists or highlights | No source filtering | RESULT_FIELDS; highlight only safe fields | 10, 11 |
| Resurrected profiles after a delete | A deleted profile reappears in search | A stale event after the delete tombstone expired (gc_deletes) | Re-read current rows; raise gc_deletes during backfills | 03, 04, 14 |
Lost updates after _update_by_query | A later change is rejected with 409, silently | By-query updates bump the external version | Never modify the projection in place | 03 |
| Outbox events skipped | Changes committed late never reach the index | Polling by id watermark | Claim-and-delete with SKIP LOCKED | 04 |
| Kotlin classes will not deserialise | Cannot construct instance ... (no Creators ...) | Spring Boot 4’s default Jackson3JsonpMapper has no Kotlin module | A JsonpMapper bean with the Kotlin module | 09 |
| Totals stop at 10,000 | Counts shown as exact that are not | track_total_hits default; Spring Data Page hides the relation | Return the relation; track totals where needed | 09 |
cross_fields finds nothing | prashant patna returns 0 | Fields with different search analysers form separate groups | Force one analyser; check with _validate/query?explain | 11 |
| Stale counts after a synonym change | Counts ignore the new synonym | The request cache is not invalidated by synonym reloads | Clear the request cache after synonym changes | 07 |
| Wrong top-N across shards | Top terms with wrong names and counts | Per-shard top lists merged | Read doc_count_error_upper_bound; shard_size; composite | 12 |
| Settings restore fails after a backfill | Index left with refresh disabled | Typed settings API rejects null | Restore explicit defaults; test the failure path | 14 |
| PIT paging refused for a scoped key | 403 on the second page | PIT searches are authorised on concrete indices, not aliases | Grant read on the index versions, or disable stable paging | 16 |
| Facets leak outside the caller’s scope | An agent sees counts for other states | Scope applied as a facet filter, which facets ignore | Scope as a base filter on every query and aggregation | 16 |
| Identifiers in URLs | Emails and phone numbers in access logs | GET lookups with query parameters | POST lookups with a request body | 16 |
The rest of this chapter is the test suite that keeps the application-side rows of this table fixed.
What to test, and at which level
| Level | Tests | Needs | Runs in |
|---|---|---|---|
| Unit | Query and facet builders, cursors | Nothing | Milliseconds |
| Integration | Mapping creation, analysis, filters, autocomplete, pagination, alias swap, authorisation | A real Elasticsearch in a container | Tens of seconds |
| Benchmark | Throughput and latency (chapter 08) | Production-like hardware | Minutes to hours |
Query builders are pure functions from a request to JSON, so most query logic can be tested without a cluster. What cannot be unit-tested is whether Elasticsearch interprets that JSON as intended against the real mapping: that cross_fields blends the groups you expect, or that a synonym reaches a search. Those need a real node. The same Elasticsearch version as production matters, because behaviours in this series, such as the request-cache trap, are version-specific.
Stage 1 — Test dependencies
Add the test dependencies to search-api/build.gradle.kts, after the existing implementation lines:
testImplementation("org.springframework.boot:spring-boot-starter-test") testImplementation("org.springframework.boot:spring-boot-testcontainers") testImplementation("org.springframework.security:spring-security-test") testImplementation("org.testcontainers:testcontainers-elasticsearch") testImplementation("org.testcontainers:testcontainers-junit-jupiter") testImplementation("org.jetbrains.kotlin:kotlin-test-junit5") testRuntimeOnly("org.junit.platform:junit-platform-launcher")}
tasks.test { useJUnitPlatform()}Spring Boot 4.1.1 manages Testcontainers 2.0.5, whose Elasticsearch module is testcontainers-elasticsearch. See the Testcontainers Elasticsearch module and Spring Boot’s Testcontainers support.
Stage 2 — Unit tests for the builders
A small helper turns a request into a navigable JSON tree, using the same JsonpMapper as the application:
package `in`.o612.eng.usersearch.api.search
import co.elastic.clients.elasticsearch.core.SearchRequestimport co.elastic.clients.json.JsonpUtilsimport `in`.o612.eng.usersearch.index.ElasticsearchJsonimport tools.jackson.databind.JsonNodeimport tools.jackson.databind.json.JsonMapper
/** The JSON body a request would send, as a tree that tests can navigate. */fun SearchRequest.toJsonTree(): JsonNode = JsonMapper.builder().build().readTree(JsonpUtils.toJsonString(this, ElasticsearchJson.jsonpMapper()))
/** The elements of a JSON array node. */fun JsonNode.items(): List<JsonNode> = (0 until size()).map { get(it) }The builder tests assert the design decisions of chapters 11 to 16, one per test:
package `in`.o612.eng.usersearch.api.search
import `in`.o612.eng.usersearch.api.web.NameModeimport `in`.o612.eng.usersearch.api.web.SortOrderimport `in`.o612.eng.usersearch.api.web.UserSearchRequestimport org.junit.jupiter.api.Testimport kotlin.test.assertEqualsimport kotlin.test.assertTrue
class UserQueryBuilderTest {
@Test fun `exact constraints are filters and only the name clause is scored`() { val json = UserQueryBuilder.build(UserSearchRequest(name = "prashant", state = "bihar", city = "patna")).toJsonTree() val bool = json.at("/query/bool")
assertEquals(1, bool.at("/must").size(), "one scored clause") val filterFields = bool.at("/filter").items().map { it.properties().first().value.properties().first().key } assertEquals(listOf("accountStatus", "state", "city"), filterFields) }
@Test fun `cross_fields forces one analyser so words may match in different fields`() { val json = UserQueryBuilder.build(UserSearchRequest(name = "prashant patna")).toJsonTree() val crossFields = json.at("/query/bool/must/0/bool/should/0/multi_match")
assertEquals("cross_fields", crossFields.at("/type").asString()) assertEquals("name_search", crossFields.at("/analyzer").asString()) assertEquals("1", json.at("/query/bool/must/0/bool/minimum_should_match").asString()) }
@Test fun `every sort ends with the unique userId`() { for (order in SortOrder.entries) { val sort = UserQueryBuilder.build(UserSearchRequest(sort = order)).toJsonTree().at("/sort") assertEquals("userId", sort.items().last().properties().first().key, "sort $order must be total") } }
@Test fun `prefix mode searches the edge n-gram field without fuzziness`() { val json = UserQueryBuilder.build(UserSearchRequest(name = "pras", nameMode = NameMode.PREFIX)).toJsonTree() val match = json.at("/query/bool/must/0/match/fullName.prefix")
assertEquals("and", match.at("/operator").asString()) assertTrue(match.at("/fuzziness").isMissingNode) }
@Test fun `a point-in-time page names no index and continues after the cursor`() { val request = UserQueryBuilder.build( UserSearchRequest(city = "patna", sort = SortOrder.UPDATED_AT), pitId = "pit-1", searchAfter = SearchCursor(listOf(1787702400000L, "385496"), null, "x").searchAfter(), )
assertTrue(request.index().isEmpty(), "a PIT search must not name an index") assertEquals("pit-1", request.pit()!!.id()) assertEquals(2, request.searchAfter().size) }
@Test fun `the caller's scope restricts searches and every facet`() { val scope = SearchScope(setOf("bihar")) val request = UserSearchRequest(name = "prashant")
val searchFilters = UserQueryBuilder.build(request, scope).toJsonTree().at("/query/bool/filter") assertTrue(searchFilters.items().any { it.at("/terms/state").isArray }, "search must carry the scope")
// Regression test for the facet leak in chapter 16: the scope must be in the shared base query, // which every facet applies, not in the state facet's own filter, which that facet ignores. val facetBase = UserFacetsBuilder.build(request, scope).toJsonTree().at("/query/bool/filter") assertTrue(facetBase.items().any { it.at("/terms/state").isArray }, "facets must carry the scope") }
@Test fun `a facet ignores its own filter and applies the others`() { val json = UserFacetsBuilder.build(UserSearchRequest(state = "bihar", city = "patna")).toJsonTree()
val cityFacetFilters = json.at("/aggregations/city/filter/bool/filter").items().map { it.at("/term").properties().first().key } val stateFacetFilters = json.at("/aggregations/state/filter/bool/filter").items().map { it.at("/term").properties().first().key } assertEquals(listOf("state"), cityFacetFilters) assertEquals(listOf("city"), stateFacetFilters) }}Each test names a rule a future change could break: only the name clause is scored; cross_fields forces one analyser; every sort ends with userId; autocomplete has no fuzziness; a PIT search names no index; the caller’s scope reaches both searches and facets; and each facet ignores only its own filter.
Cursors get their own tests: the sort values’ types survive the round trip, and a cursor is refused for a different search but not for a different page size.
package `in`.o612.eng.usersearch.api.search
import co.elastic.clients.elasticsearch._types.FieldValueimport `in`.o612.eng.usersearch.api.web.UserSearchRequestimport org.junit.jupiter.api.Testimport org.junit.jupiter.api.assertThrowsimport kotlin.test.assertEquals
class SearchCursorTest {
private val request = UserSearchRequest(city = "patna")
@Test fun `a cursor round-trips its sort values with their types`() { val sortValues = listOf(FieldValue.of(25.586945), FieldValue.of(1761696000000L), FieldValue.of("343200")) val decoded = SearchCursor.decode(SearchCursor.from(sortValues, null, request).encode(), request)
assertEquals(listOf(25.586945, 1761696000000L, "343200"), decoded.searchAfter().map { it._get() }) }
@Test fun `a cursor is rejected for a different search, but not for a different page size`() { val token = SearchCursor.from(listOf(FieldValue.of("1")), null, request).encode()
SearchCursor.decode(token, request.copy(size = 50)) assertThrows<InvalidCursorException> { SearchCursor.decode(token, request.copy(city = "gaya")) } assertThrows<InvalidCursorException> { SearchCursor.decode("not-a-cursor", request) } }}Run them:
./gradlew :search-api:test --tests '*UserQueryBuilderTest' --tests '*SearchCursorTest'The test reports record nine tests and no failures:
testsuite name="in.o612.eng.usersearch.api.search.UserQueryBuilderTest" tests="7" skipped="0" failures="0" errors="0"testsuite name="in.o612.eng.usersearch.api.search.SearchCursorTest" tests="2" skipped="0" failures="0" errors="0"Check that the regression test can fail
A test that cannot fail protects nothing. Remove the scope line from UserQueryBuilder.baseFilters, if (scope.states.isNotEmpty()) add(terms("state", scope.states)), and run UserQueryBuilderTest again: the build fails on the caller's scope restricts searches and every facet. Put the line back, and it passes. That is a manual mutation test, and it is worth doing once for every test that guards a security rule.
Stage 3 — A shared, TLS-enabled Elasticsearch for integration tests
All integration test classes share one container, declared once:
package `in`.o612.eng.usersearch.api.it
import org.springframework.boot.testcontainers.service.connection.ServiceConnectionimport org.springframework.boot.testcontainers.service.connection.Sslimport org.testcontainers.elasticsearch.ElasticsearchContainer
/** * One Elasticsearch for every integration test class, with security and TLS on, as in production. * Spring Boot takes the URL, the elastic user's password, and the container's CA from it. */object SharedElasticsearch { @JvmField @ServiceConnection @Ssl // without it, Spring Boot connects over plain HTTP and the TLS-only container closes the connection val elasticsearch: ElasticsearchContainer = ElasticsearchContainer("docker.elastic.co/elasticsearch/elasticsearch:9.5.4") .withEnv("ES_JAVA_OPTS", "-Xms512m -Xmx512m") // Started here, not by Spring: the CA certificate is read while test properties are resolved. .apply { start() }}Two lines in it are there because the first attempts failed:
@Ssl. The Testcontainers image runs with security and TLS enabled, as production does. Spring Boot’s service connection takes the URL and theelasticuser’s password from the container, but it only builds an SSL bundle from the container’s CA certificate when the field is annotated with Spring Boot’s@Ssl. Without it, the connection details reportedHTTPand no SSL bundle, and every request failed withConnection closed by peer..apply { start() }. The application reads the CA certificate path from a property, and properties are resolved before Spring starts containers imported with@ImportTestcontainers. Reading the certificate from a container that had not started failed withcopyFileFromContainer can only be used when the Container is created. Starting it when the object initialises fixes the order; Spring’s later start is a no-op.
The support object supplies the properties the application normally reads from the environment, and seeds the data once per JVM:
package `in`.o612.eng.usersearch.api.it
import co.elastic.clients.elasticsearch.ElasticsearchClientimport co.elastic.clients.elasticsearch._types.Refreshimport `in`.o612.eng.usersearch.index.UserProfileIndeximport org.springframework.test.context.DynamicPropertyRegistryimport java.nio.file.Files
object IntegrationTestSupport { /** Properties the application normally reads from the environment, set for the test container. */ fun register(registry: DynamicPropertyRegistry) { val ca = Files.createTempFile("es-test-ca", ".crt") Files.write(ca, SharedElasticsearch.elasticsearch.caCertAsBytes().orElseThrow()) registry.add("ELASTICSEARCH_CA_CERT") { "file:$ca" } registry.add("ELASTICSEARCH_URIS") { "https://localhost" } // replaced by the service connection registry.add("ELASTICSEARCH_API_KEY") { "" } // the service connection uses the elastic user registry.add("USER_SEARCH_INDEX_BOOTSTRAP") { "true" } registry.add("LAB_AGENT_PASSWORD") { AGENT_PASSWORD } registry.add("LAB_SUPERVISOR_PASSWORD") { SUPERVISOR_PASSWORD } }
const val AGENT_PASSWORD = "agent-test" const val SUPERVISOR_PASSWORD = "supervisor-test" const val PROFILES = 400
@Volatile private var seeded = false
/** Indexes the synthetic profiles once per JVM, and waits until they are searchable. */ @Synchronized fun seed(client: ElasticsearchClient) { if (seeded) return val response = client.bulk { b -> b.refresh(Refresh.WaitFor) TestProfiles.generate(PROFILES).forEach { doc -> b.operations { op -> op.index { i -> i.index(UserProfileIndex.WRITE_ALIAS).id(doc.userId).document(doc) } } } b } check(!response.errors()) { "Seeding failed: ${response.items().firstOrNull { it.error() != null }?.error()}" } seeded = true }}Realistic, anonymised test data
Test data for a personal-data system must be synthetic, never copied or “masked” from production. Masked production data still carries real distributions of rare names, real combinations of city and age, and real people who can be re-identified from them. Generate data instead, and make it realistic where realism affects behaviour:
- Name and place frequencies. Relevance and top-N behaviour depend on how skewed values are. If you need production-like skew, derive aggregate frequency tables from production under your privacy process, and generate from those.
- Script and diacritics. Include accented and non-Latin names if your users have them, or analysis bugs stay invisible.
- Status and time distributions. Filters and date windows need values on both sides of their boundaries.
- Determinism. A generator driven by the row number produces the same data on every run, so a failing test fails the same way twice.
The test generator follows those rules at a small scale; chapter 01’s SQL seed did the same for the lab:
package `in`.o612.eng.usersearch.api.it
import `in`.o612.eng.usersearch.index.UserProfileDocumentimport java.time.Instantimport java.time.LocalDateimport java.time.temporal.ChronoUnit
/** * Deterministic synthetic profiles. Nothing is copied from real data: names, places, and identifiers * are generated, with enough variety to exercise analysis, filters, and paging. */object TestProfiles { private val firstNames = listOf("Prashant", "Priya", "Mohammed", "José", "Lakshmi", "Rahul", "Fatima", "Arjun") private val lastNames = listOf("Kumar", "Sharma", "Khan", "Fernandes", "Iyer", "Jha") private val places = listOf("Patna" to "Bihar", "Gaya" to "Bihar", "Kochi" to "Kerala", "Mumbai" to "Maharashtra") private val statuses = listOf("ACTIVE", "ACTIVE", "ACTIVE", "INACTIVE") private val start: Instant = Instant.parse("2026-01-01T00:00:00Z")
fun generate(count: Int): List<UserProfileDocument> = (1..count).map { id -> val first = firstNames[id % firstNames.size] val last = lastNames[(id / firstNames.size) % lastNames.size] val (city, state) = places[(id * 7) % places.size] UserProfileDocument( userId = id.toString(), fullName = "$first $last", firstName = first, lastName = last, email = "${first.lowercase()}.${last.lowercase()}.$id@example.com", mobileNumber = "9" + id.toString().padStart(9, '0'), city = city, state = state, country = "IN", pincode = "800" + (id % 1000).toString().padStart(3, '0'), dateOfBirth = LocalDate.of(1970, 1, 1).plusDays(id * 37L % 15000), gender = if (id % 2 == 0) "FEMALE" else "MALE", accountStatus = statuses[id % statuses.size], createdAt = start, updatedAt = start.plus(id.toLong(), ChronoUnit.HOURS), ) }}Stage 4 — The first real bugs the suite found
The first integration run failed before any test started. The index bootstrapper from chapter 10 checked for aliases the moment the application started, and the fresh node answered 503, even to a health request that asked to wait for yellow, while it was still recovering its cluster state. A deployment that starts the application alongside a new cluster has the same race. The bootstrapper now retries until the cluster is ready, with a deadline:
package `in`.o612.eng.usersearch.api.index
import co.elastic.clients.elasticsearch.ElasticsearchClientimport co.elastic.clients.elasticsearch._types.ElasticsearchExceptionimport co.elastic.clients.elasticsearch._types.HealthStatusimport co.elastic.clients.transport.TransportExceptionimport `in`.o612.eng.usersearch.index.IndexAdminimport `in`.o612.eng.usersearch.index.UserProfileIndeximport org.slf4j.LoggerFactoryimport org.springframework.boot.ApplicationArgumentsimport org.springframework.boot.ApplicationRunnerimport org.springframework.boot.autoconfigure.condition.ConditionalOnBooleanPropertyimport org.springframework.stereotype.Componentimport java.time.Durationimport java.time.Instant
/** * Creates the current index version and its aliases when none exist. * Enabled only with user-search.index.bootstrap=true, for local development and tests. */@Component@ConditionalOnBooleanProperty("user-search.index.bootstrap")class IndexBootstrapper( private val indexAdmin: IndexAdmin, private val client: ElasticsearchClient,) : ApplicationRunner {
private val log = LoggerFactory.getLogger(javaClass)
override fun run(args: ApplicationArguments) { awaitCluster(Duration.ofSeconds(60)) val targets = indexAdmin.aliasTargets(UserProfileIndex.READ_ALIAS) if (targets.isNotEmpty()) { log.info("{} already points to {}; nothing to create", UserProfileIndex.READ_ALIAS, targets) return } // The current version, not v1: the code depends on its fields, such as the facets' display sub-fields. val index = UserProfileIndex.indexName(UserProfileIndex.CURRENT_VERSION) val created = indexAdmin.createIndex(UserProfileIndex.CURRENT_VERSION) if (indexAdmin.aliasTargets(UserProfileIndex.READ_ALIAS).isEmpty()) indexAdmin.addAliases(index) log.info("Index {} created: {}", index, created) }
/** * A node that has just started answers 503 while it recovers its cluster state, even to a health * request that asks to wait (chapter 17). Retry until the cluster is at least yellow, or give up. */ private fun awaitCluster(limit: Duration) { val deadline = Instant.now().plus(limit) while (true) { try { val health = client.cluster().health { it.waitForStatus(HealthStatus.Yellow).timeout { t -> t.time("10s") } } if (!health.timedOut()) return } catch (e: ElasticsearchException) { log.info("Cluster not ready yet: {}", e.status()) } catch (e: TransportException) { log.info("Cluster not ready yet: {}", e.message) } check(Instant.now().isBefore(deadline)) { "Elasticsearch was not ready within $limit" } Thread.sleep(1_000) } }}The second run failed on the facets test: the agent’s state facet was empty. The bootstrapper created user-profile-v1, but since chapter 14 the facets count .display sub-fields that only v2 has. Chapter 14 warned that code depending on a new mapping must not run against the old one, and the bootstrapper did exactly that. It now creates the version the code is written for, and attaches the aliases itself, because the v2 definition deliberately has none. Add the version constant to UserProfileIndex:
/** The index version the applications' code is written for. */ const val CURRENT_VERSION = 2and give IndexAdmin an optional index name for createIndex, plus a method that attaches both aliases:
/** * Creates an index named [name], `user-profile-v<version>` by default, from the reviewed definition of * [version], including any aliases the definition declares. Returns false, and changes nothing, if it exists. */ fun createIndex(version: Int, name: String = UserProfileIndex.indexName(version)): Boolean { if (client.indices().exists { it.index(name) }.value()) return false /** Points both aliases at [index], for a cluster that has none yet. */ fun addAliases(index: String) { client.indices().updateAliases { u -> u.actions { a -> a.add { ad -> ad.index(index).alias(UserProfileIndex.READ_ALIAS) } } .actions { a -> a.add { ad -> ad.index(index).alias(UserProfileIndex.WRITE_ALIAS).isWriteIndex(true) } } } }The rest of createIndex is unchanged. Both bugs were in chapter 10’s code and would have reached anyone bootstrapping a new environment. Neither can be found by a unit test.
Stage 5 — Integration tests
The first class exercises the Search API against the real index, through its service and, for authorisation, over HTTP:
package `in`.o612.eng.usersearch.api.it
import co.elastic.clients.elasticsearch.ElasticsearchClientimport `in`.o612.eng.usersearch.api.search.UserSearchServiceimport `in`.o612.eng.usersearch.api.web.FacetsResponseimport `in`.o612.eng.usersearch.api.web.NameModeimport `in`.o612.eng.usersearch.api.web.SortOrderimport `in`.o612.eng.usersearch.api.web.UserSearchRequestimport `in`.o612.eng.usersearch.index.IndexAdminimport `in`.o612.eng.usersearch.index.UserProfileIndeximport org.junit.jupiter.api.BeforeEachimport org.junit.jupiter.api.Testimport org.springframework.beans.factory.annotation.Autowiredimport org.springframework.beans.factory.annotation.Valueimport org.springframework.boot.test.context.SpringBootTestimport org.springframework.boot.testcontainers.context.ImportTestcontainersimport org.springframework.test.context.DynamicPropertyRegistryimport org.springframework.test.context.DynamicPropertySourceimport org.springframework.web.client.RestClientimport kotlin.test.assertEqualsimport kotlin.test.assertTrue
@SpringBootTest(webEnvironment = SpringBootTest.WebEnvironment.RANDOM_PORT)@ImportTestcontainers(SharedElasticsearch::class)class UserSearchIntegrationTest {
@Autowired lateinit var client: ElasticsearchClient @Autowired lateinit var indexAdmin: IndexAdmin @Autowired lateinit var service: UserSearchService @Value("\${local.server.port}") var port: Int = 0
@BeforeEach fun seed() = IntegrationTestSupport.seed(client)
@Test fun `the bootstrapper creates the current version, strict, behind both aliases`() { val current = UserProfileIndex.indexName(UserProfileIndex.CURRENT_VERSION) assertEquals(setOf(current), indexAdmin.aliasTargets(UserProfileIndex.READ_ALIAS)) assertEquals(setOf(current), indexAdmin.aliasTargets(UserProfileIndex.WRITE_ALIAS)) val mapping = client.indices().getMapping { it.index(current) }[current]!!.mappings() assertEquals("strict", mapping.dynamic()!!.jsonValue()) }
@Test fun `autocomplete, typos, synonyms, and filters find the expected profiles`() { val all = TestProfiles.generate(IntegrationTestSupport.PROFILES).filter { it.accountStatus == "ACTIVE" }
fun count(request: UserSearchRequest) = service.search(request).total.value
assertEquals(all.count { it.firstName == "Prashant" }.toLong(), count(UserSearchRequest(name = "pras", nameMode = NameMode.PREFIX))) assertEquals(all.count { it.fullName == "Prashant Kumar" }.toLong(), count(UserSearchRequest(name = "prashnat kumr"))) assertEquals(all.count { it.firstName == "Mohammed" }.toLong(), count(UserSearchRequest(name = "mohd"))) assertEquals(all.count { it.firstName == "José" }.toLong(), count(UserSearchRequest(name = "JOSE"))) assertEquals(all.count { it.city == "Patna" }.toLong(), count(UserSearchRequest(city = "PATNA"))) }
@Test fun `a stable walk returns every match exactly once`() { val expected = TestProfiles.generate(IntegrationTestSupport.PROFILES).count { it.accountStatus == "ACTIVE" && it.state == "Bihar" } val seen = mutableListOf<String>() var request = UserSearchRequest(state = "bihar", sort = SortOrder.UPDATED_AT, size = 7, stable = true) while (true) { val page = service.search(request) seen += page.hits.map { it.userId } request = request.copy(cursor = page.nextCursor ?: break) } assertEquals(expected, seen.size) assertEquals(seen.size, seen.toSet().size, "no profile twice") }
@Test fun `an agent's facets never include states outside their scope`() { val rest = RestClient.create("http://localhost:$port") val facets = rest.get().uri("/api/users/facets") .headers { it.setBasicAuth("agent", IntegrationTestSupport.AGENT_PASSWORD) } .retrieve().body(FacetsResponse::class.java)!! assertEquals(listOf("Bihar"), facets.facets.getValue("state").map { it.value })
val supervisorFacets = rest.get().uri("/api/users/facets") .headers { it.setBasicAuth("supervisor", IntegrationTestSupport.SUPERVISOR_PASSWORD) } .retrieve().body(FacetsResponse::class.java)!! assertTrue(supervisorFacets.facets.getValue("state").size > 1) }
companion object { @JvmStatic @DynamicPropertySource fun properties(registry: DynamicPropertyRegistry) = IntegrationTestSupport.register(registry) }}The expected counts are computed from the same generator that seeded the index, so each assertion states the rule in terms of the data: every active Prashant for pras, every active Prashant Kumar despite the typos in prashnat kumr, every Mohammed for mohd, every José for JOSE, and every profile in Patna for PATNA. The stable walk pages seven at a time through every active profile in Bihar and checks that each appears exactly once. The authorisation test is chapter 16’s facet leak, end to end.
The second class tests the alias swap, and the rollback path in its finally block:
package `in`.o612.eng.usersearch.api.it
import co.elastic.clients.elasticsearch.ElasticsearchClientimport co.elastic.clients.elasticsearch._types.Refreshimport `in`.o612.eng.usersearch.index.IndexAdminimport `in`.o612.eng.usersearch.index.UserProfileIndeximport org.junit.jupiter.api.BeforeEachimport org.junit.jupiter.api.Testimport org.springframework.beans.factory.annotation.Autowiredimport org.springframework.boot.test.context.SpringBootTestimport org.springframework.boot.testcontainers.context.ImportTestcontainersimport org.springframework.test.context.DynamicPropertyRegistryimport org.springframework.test.context.DynamicPropertySourceimport kotlin.test.assertEquals
@SpringBootTest@ImportTestcontainers(SharedElasticsearch::class)class AliasSwapIntegrationTest {
@Autowired lateinit var client: ElasticsearchClient @Autowired lateinit var indexAdmin: IndexAdmin
@BeforeEach fun seed() = IntegrationTestSupport.seed(client)
@Test fun `swapping aliases moves reads and writes to the new index, and back`() { val current = UserProfileIndex.indexName(UserProfileIndex.CURRENT_VERSION) val candidate = "$current-candidate" indexAdmin.createIndex(UserProfileIndex.CURRENT_VERSION, name = candidate) // The candidate holds one profile that the current index does not. val newcomer = TestProfiles.generate(IntegrationTestSupport.PROFILES + 1).last() client.index { it.index(candidate).id(newcomer.userId).document(newcomer).refresh(Refresh.WaitFor) }
try { indexAdmin.swapAliases(fromIndex = current, toIndex = candidate) assertEquals(setOf(candidate), indexAdmin.aliasTargets(UserProfileIndex.READ_ALIAS)) assertEquals(setOf(candidate), indexAdmin.aliasTargets(UserProfileIndex.WRITE_ALIAS)) assertEquals(1L, client.count { it.index(UserProfileIndex.READ_ALIAS) }.count()) } finally { indexAdmin.swapAliases(fromIndex = candidate, toIndex = current) // the rollback path client.indices().delete { it.index(candidate) } } assertEquals(setOf(current), indexAdmin.aliasTargets(UserProfileIndex.READ_ALIAS)) assertEquals(IntegrationTestSupport.PROFILES.toLong(), client.count { it.index(UserProfileIndex.READ_ALIAS) }.count()) }
companion object { @JvmStatic @DynamicPropertySource fun properties(registry: DynamicPropertyRegistry) = IntegrationTestSupport.register(registry) }}It creates a candidate index from the current definition, adds one profile the live index does not have, swaps both aliases, checks that reads now see only the candidate, and swaps back. It leaves the cluster as it found it, so the test classes can run in any order.
Run the whole build, all modules and all tests:
./gradlew buildtestsuite name="in.o612.eng.usersearch.api.search.SearchCursorTest" tests="2" skipped="0" failures="0" errors="0"testsuite name="in.o612.eng.usersearch.api.search.UserQueryBuilderTest" tests="7" skipped="0" failures="0" errors="0"testsuite name="in.o612.eng.usersearch.api.it.AliasSwapIntegrationTest" tests="1" skipped="0" failures="0" errors="0"testsuite name="in.o612.eng.usersearch.api.it.UserSearchIntegrationTest" tests="4" skipped="0" failures="0" errors="0"In the lab the build took about 48 seconds, almost all of it the container starting and the first context loading. Both test classes use the same Spring context and the same container, because their configuration is identical and Spring caches it.
Checkpoint
./gradlew build succeeds with 14 tests. If the integration tests fail with Connection closed by peer, the @Ssl annotation is missing; with copyFileFromContainer can only be used when the Container is created, the container is not started before the properties are read.
What the suite does not cover
- The indexer. The relay and backfill need PostgreSQL too. A
PostgreSQLContaineralongsideSharedElasticsearch, loading chapter 04’s SQL, makes that suite possible: insert, update, and delete rows, runrelayOneBatch, and assert on the index. It is left as an exercise in chapter 18. - Performance. Tests assert correctness, not speed. Latency regressions are for Rally (chapter 08), run on a schedule against a production-like environment.
- Upgrades. A test suite pinned to 9.5.4 says nothing about 9.6. Run it against the new version as the first step of every upgrade.
What you built, and what comes next
You have a reference of every failure mode this series met, and a test suite that guards the application-side ones: nine unit tests for the builders and cursors, including a mutation-checked regression test for the facet leak, and five integration tests against a TLS-enabled Elasticsearch 9.5.4. Building the suite fixed two real defects in the index bootstrapper.
Chapter 18 closes the series: a phased rollout from proof of concept to production with exit criteria for each phase, a production-readiness checklist, the questions to answer before going live, and the exercises to build next.