API composition and aggregator services
Chapter 4’s mobile-bff fanned out to order-service and payment-service and combined their results, with a deliberately simple “fail the whole request if either call fails” policy. This chapter formalizes that composition technique properly and replaces the fail-closed shortcut with a real partial-failure strategy.
1. Problem the Pattern Solves
The order-confirmation screen’s popularity has grown: it now also needs delivery estimate data from a new logistics-service, in addition to order-service and payment-service. With three dependencies instead of two, mobile-bff’s current all-or-nothing policy means the confirmation screen fails completely whenever any one of three independent services is briefly slow — a customer can’t see their order status just because the delivery-estimate service, entirely unrelated to whether their order was placed and paid for, is having a bad moment.
Product feedback is specific: “if you can’t tell me the delivery estimate right now, that’s fine, just show me the order status you do have.” This requires the aggregator to make a case-by-case decision about which failures are acceptable to degrade around and which aren’t — order status and payment confirmation are essential to this screen’s purpose; delivery estimate is a nice-to-have.
Forces in tension:
- Completeness vs. availability. Waiting for every dependency to respond guarantees a complete answer but ties the aggregate response’s availability to the least reliable of all its dependencies — worse than any single dependency’s own availability.
- Simplicity of fail-all vs. correctness of partial response. A single boolean “did everything succeed” is trivial to implement but treats an optional field’s failure identically to a critical field’s failure — usually the wrong trade-off, as Northwind’s product feedback demonstrates directly.
- Latency budget vs. number of aggregated calls. Every dependency added to an aggregator adds another call whose latency contributes to the aggregate’s total time (if sequential) or whose failure mode needs handling (if parallel) — aggregation doesn’t eliminate per-call cost, it just reshapes how failures and latency compose.
- Consistency of the aggregate view vs. independent per-call freshness. Each fanned-out call reflects its own service’s state at a slightly different instant — the aggregate response is never perfectly point-in-time consistent across all fields, a fact usually acceptable for a display screen but worth naming explicitly.
2. Core Idea
API composition means an aggregator (which can be a dedicated service, a BFF as in Chapter 4, or a component of one) calls multiple downstream services in parallel, then merges their results into one composed response — with an explicit policy for what to do when one or more of those calls fails, times out, or returns incomplete data.
Participants:
- Aggregator —
mobile-bff, unchanged in identity from Chapter 4, now with an explicit per-dependency criticality classification driving its merge policy. - Fan-out calls — the parallel calls to
order-service,payment-service, and the newlogistics-service, each independently subject to the timeout/circuit-breaker protection from Chapter 16. - Merge policy — the explicit rule set determining what response to produce given which subset of calls succeeded — the actual design work this pattern requires, and the part Chapter 4 deliberately deferred.
Commonly confused with:
- Backend for Frontend (Chapter 4). A BFF’s defining trait is owning a client-specific contract; API composition is the technique for fulfilling that contract when it requires data from multiple services. A BFF often uses composition internally (as
mobile-bffdoes), but composition can also stand alone as a dedicated aggregator service with no particular client ownership — for instance, an internal reporting aggregator with no single “client” driving its shape. - The Saga pattern (Chapter 11). A saga coordinates a sequence of state-changing operations across services, with compensation for partial failure. An aggregator composes read-only calls into one response — there’s no state change to compensate, only a response-shaping decision to make when some reads fail. Conflating the two risks treating a harmless missing read as if it needed a saga’s compensating-action machinery.
- GraphQL. GraphQL provides client-driven field selection and can implement a form of composition (resolvers fanning out to different data sources) as an incidental capability of its query model — but it’s a different technology choice with its own trade-offs (Chapter 4’s “commonly confused with” section touches on this), not simply “composition, but with a query language.”
3. When to Use It
Strong indicators:
- A client-facing view genuinely needs data from multiple independent services, and the client shouldn’t have to make multiple round trips itself (Chapter 4’s original motivation).
- Some of the aggregated data is more important to the view’s core purpose than other parts — a real distinction exists between “this field is essential” and “this field is nice to have,” as Northwind’s order-confirmation screen demonstrates.
- The downstream services are independently reachable in parallel — composition’s main benefit (avoiding the sum of sequential latencies) requires the calls not to have data dependencies on each other’s results.
Concrete use cases:
- E-commerce, as here: an order-confirmation or dashboard screen drawing from order, payment, and logistics data is a canonical composition case.
- Travel booking dashboards: a trip summary composing flight status, hotel confirmation, and car rental details from three separate provider integrations, where a slow car-rental API shouldn’t prevent showing confirmed flight and hotel details.
- Healthcare patient dashboards: composing lab results, appointment schedules, and billing status from separate systems, where each has a different criticality to the specific screen being rendered.
- Internal operational dashboards: aggregating health and metrics data from many services into one view, where a single service’s monitoring endpoint being briefly unreachable shouldn’t blank out the entire dashboard.
Prerequisites:
- Explicit, product-reviewed criticality classification per aggregated field or dependency (Section 5) — this is the actual design decision the pattern requires, not an engineering-only call.
- Per-call resilience (Chapter 16’s timeout, circuit breaker, bulkhead) on every fanned-out call — an aggregator without these on its individual calls just moves the cascading-failure risk one layer up instead of solving it.
- A clear answer for what “partial success” looks like on the wire — the response schema itself needs to represent “this field is present” versus “this field is currently unavailable,” not just silently omit fields with no signal.
4. When Not to Use It
- All aggregated data is equally critical, and a partial response would be meaningless or misleading. If a payment-confirmation receipt genuinely needs both the charge amount and the line items to make sense to the customer, a response missing one isn’t useful even though it’s “partially successful” — sometimes fail-all-or-nothing, chosen deliberately rather than by default, is the correct policy for a specific view.
- The downstream calls have data dependencies on each other. If
logistics-service’s call actually needs a value returned byorder-service’s call first, the calls can’t run in parallel — that’s a sequential orchestration concern (closer to a saga’s step-by-step shape, even for reads), not composition’s parallel fan-out-and-merge model. - A single downstream service already has everything the view needs. Aggregation for its own sake, when one call would suffice, adds complexity and multiple points of failure for no benefit.
- Overengineering signal: building a generic, configuration-driven “aggregation framework” applicable to any future composed view, before a second real composition need (beyond order confirmation) has actually appeared — keep
mobile-bff’s aggregation logic specific and simple, per Chapter 4’s own guidance against speculative BFF infrastructure.
5. Implementation Example
Explicit criticality classification, the actual design decision this pattern requires, made visible in code rather than left implicit:
package `in`.o612.eng.northwind.mobilebff
import java.util.UUIDimport java.util.concurrent.CompletableFuture
class OrderConfirmationAggregator( private val orderClient: OrderServiceClient, // CRITICAL private val paymentClient: PaymentServiceClient, // CRITICAL private val logisticsClient: LogisticsServiceClient, // OPTIONAL) { fun aggregate(orderId: UUID): CompletableFuture<OrderConfirmationView> { val orderFuture = orderClient.getOrderAsync(orderId) val paymentFuture = paymentClient.getReceiptAsync(orderId) val logisticsFuture = logisticsClient.getDeliveryEstimateAsync(orderId) .exceptionally { null } // optional: a failure here becomes "no data," never a propagated exception
return CompletableFuture.allOf(orderFuture, paymentFuture, logisticsFuture).thenApply { // If either CRITICAL future failed, .join() below throws and // the whole aggregation fails — the correct behavior for // fields this view cannot meaningfully render without. val order = orderFuture.join() val payment = paymentFuture.join() val logistics = logisticsFuture.join() // already defaulted to null on failure above — never throws
OrderConfirmationView( orderId = order.id, status = order.status, totalCharged = payment.amount, deliveryEstimate = logistics?.estimatedDate, // null, with an explicit "unavailable" signal below deliveryEstimateAvailable = logistics != null, ) } }}
data class OrderConfirmationView( val orderId: UUID, val status: String, val totalCharged: java.math.BigDecimal, val deliveryEstimate: java.time.LocalDate?, val deliveryEstimateAvailable: Boolean, // explicit signal, not just a null the client has to interpret)The .exceptionally { null } on logisticsFuture is the entire mechanism implementing “optional” — a failure there is caught and converted into an absent value before the allOf/join() combination, so it can never fail the aggregate response. orderFuture and paymentFuture have no such handling — their exceptions propagate through join() and fail the whole aggregation, which is the correct, deliberate behavior for critical fields.
The controller, translating a failed aggregation into an honest error response rather than a partial one for critical-field failures:
package `in`.o612.eng.northwind.mobilebff
import org.springframework.http.HttpStatusimport org.springframework.http.ResponseEntityimport org.springframework.web.bind.annotation.*import java.util.UUIDimport java.util.concurrent.CompletionException
@RestController@RequestMapping("/orders")class OrderConfirmationController(private val aggregator: OrderConfirmationAggregator) {
@GetMapping("/{orderId}/confirmation") fun confirmation(@PathVariable orderId: UUID): ResponseEntity<*> = try { ResponseEntity.ok(aggregator.aggregate(orderId).get()) } catch (e: CompletionException) { // A CRITICAL dependency failed — the response, even partial, // wouldn't be meaningful. Fail honestly, not silently degrade. ResponseEntity.status(HttpStatus.SERVICE_UNAVAILABLE) .body(mapOf("error" to "Unable to load order confirmation right now")) }}Note explicitly what changed from Chapter 4: that chapter’s version failed the whole request on any failure, including logistics. This version distinguishes — a logistics-service failure produces a complete, successful response with deliveryEstimateAvailable = false; only an order-service or payment-service failure produces the 503.
Test asserting the partial-failure policy explicitly — the actual behavior this pattern needs verified, not just “the happy path works”:
package `in`.o612.eng.northwind.mobilebff
import org.junit.jupiter.api.Testimport org.assertj.core.api.Assertions.assertThatimport java.util.concurrent.CompletableFuture
class OrderConfirmationAggregatorTest {
@Test fun `logistics failure produces a complete response with delivery estimate marked unavailable`() { val aggregator = OrderConfirmationAggregator( orderClient = stubOrderClient(succeeds = true), paymentClient = stubPaymentClient(succeeds = true), logisticsClient = stubLogisticsClient(succeeds = false), // simulated failure )
val result = aggregator.aggregate(sampleOrderId).get()
assertThat(result.deliveryEstimateAvailable).isFalse() assertThat(result.status).isNotNull() // critical fields still present }
@Test fun `order-service failure fails the whole aggregation`() { val aggregator = OrderConfirmationAggregator( orderClient = stubOrderClient(succeeds = false), // simulated failure — CRITICAL paymentClient = stubPaymentClient(succeeds = true), logisticsClient = stubLogisticsClient(succeeds = true), )
assertThatThrownBy { aggregator.aggregate(sampleOrderId).get() } .hasCauseInstanceOf(RuntimeException::class.java) }}6. Step-by-Step Flow
- Client action. The mobile app requests the confirmation screen, unaware of how many services back it.
- API request.
mobile-bfffans out three calls in parallel — no sequential waiting, since none of the three calls depends on another’s result. - Service behavior. Each downstream service responds independently, each already protected by its own Chapter 16 resilience configuration (timeout, circuit breaker) at the client level.
- Database interaction. Handled entirely within each downstream service, unaffected by the aggregation happening above them.
- Inter-service communication. Three independent, parallel HTTP calls — the aggregator’s entire job is managing their combination, not their individual mechanics.
- Error or failure handling.
logistics-service’s timeout is caught and converted to an absent, explicitly-flagged field; an equivalent failure fromorder-serviceorpayment-servicewould instead fail the whole response, per the criticality classification in Section 5. - Observability signals. Track per-dependency failure rate within the aggregator specifically — a rising
logistics-servicefailure rate that never surfaces as a user-visible error (because it’s optional) could otherwise go unnoticed without deliberate monitoring at this layer. - Final response. The client receives a
200 OKwithdeliveryEstimateAvailable: false— a complete, honest, immediately renderable response that reflects exactly what was actually knowable at the time, rather than either a stale guess or a needless failure.
7. Production Concerns
- Timeouts, retries, idempotency. Every fanned-out call needs its own timeout (Chapter 16) independent of the others — an aggregator’s total latency is bounded by its slowest critical call (optional calls can time out without holding up the response, if implemented with a tight enough per-call timeout relative to the aggregate’s overall budget).
- Data consistency. The composed response reflects each service’s state at a slightly different instant — acceptable for a display screen, but worth documenting explicitly if a consumer of the aggregate response might mistake it for a single, perfectly consistent snapshot.
- API versioning. The aggregate response schema (
OrderConfirmationView) ismobile-bff’s own contract, versioned independently of any of its three downstream services’ contracts — exactly the benefit Chapter 4 established for BFF-owned contracts. - Authentication and service-to-service trust. Each fanned-out call needs its own service-to-service credential, same as any Chapter 6 synchronous call — aggregation doesn’t change the trust model, just the fan-out shape.
- Logging, metrics, tracing, correlation IDs. Propagate one correlation ID across all three parallel calls, so a trace shows them as siblings under one aggregate request — essential for diagnosing which specific downstream call was slow when the aggregate’s total latency spikes.
- Kubernetes deployment, autoscaling. The aggregator’s own resource needs scale with request volume, same as any service — no special considerations beyond ensuring its own thread/connection pool can actually sustain the fan-out concurrency it generates (three outbound calls per one inbound request, tripling connection-pool pressure relative to a pass-through service).
- Testing strategy. Test every combination of critical-success/critical-failure/optional-success/optional-failure explicitly, as Section 5 does for two of the four relevant combinations — this is where an aggregator’s correctness actually lives, not in the individual downstream clients’ own logic.
- Migration strategy. Add a new field to the aggregate response as an optional dependency first (as
logistics-servicewas), even if product eventually wants it treated as critical — this lets the aggregator ship incrementally without a new integration ever risking the existing critical fields’ availability during its own early rollout and stabilization.
8. Common Mistakes
- Treating every aggregated field as equally critical by default. Failing the whole response because an optional field’s call failed, as Chapter 4’s original version did, produces unnecessary failures for functionality the client doesn’t actually need to render. Fix: classify every aggregated field’s criticality explicitly and deliberately, with product input, as Section 5 does.
- Silently omitting a failed optional field with no signal. Simply leaving
deliveryEstimatenull with nodeliveryEstimateAvailableflag forces the client to guess whether null means “no estimate exists” or “we couldn’t fetch it right now” — a real ambiguity with different correct UI treatments. Fix: always pair an optional field with an explicit availability signal, as shown. - Running fan-out calls sequentially instead of in parallel. Awaiting
order-service’s response before starting thepayment-servicecall, when neither depends on the other, needlessly sums their latencies instead of taking the maximum. Fix: issue independent calls in parallel (asCompletableFuture.allOfdoes here), reserving sequential calls only for genuine data dependencies. - No per-call timeout, only an aggregate one. Relying solely on an overall aggregation timeout without individual per-call limits means one very slow optional call can still consume most of the aggregate’s time budget before its own failure is even detected. Fix: apply Chapter 16’s resilience patterns to each individual fanned-out call, not just the aggregate as a whole.
- Building a generic aggregation framework before a second real use case exists. Over-engineering
OrderConfirmationAggregatorinto configurable, reusable “aggregation infrastructure” for hypothetical future composed views adds complexity with only one actual consumer. Fix: keep the aggregator specific and simple, generalizing only once a second, genuinely similar composition need appears. - Ignoring the point-in-time inconsistency across composed fields. Presenting the composed response as if it were a single atomic snapshot, when in practice
order-service’s andpayment-service’s data were fetched microseconds apart and could theoretically reflect different moments, can mislead a consumer relying on the response for something more rigorous than display. Fix: document this inconsistency explicitly wherever the aggregate response might be used for anything beyond a UI screen.
9. Decision Guide
| Problem signal | Use this pattern? | Why | Alternative |
|---|---|---|---|
| A client view needs data from multiple independent services, avoidable client-side round trips | Yes | Fans out in parallel, composes one response, reduces client complexity | — |
| Some aggregated data is essential, some is a nice-to-have | Yes, with explicit criticality classification | Enables a genuinely useful partial response instead of all-or-nothing failure | — |
| All aggregated fields are equally essential to the view’s purpose | Fail-all-or-nothing, deliberately chosen | A partial response would be meaningless for this specific view | Deliberate all-critical policy, not a default |
| Downstream calls have data dependencies on each other | No (not pure composition) | Calls can’t run in parallel; this is sequential orchestration, not fan-out composition | Sequential call chain, or a saga-shaped read orchestration |
| Only one downstream service is actually needed | No | Aggregation for one dependency adds complexity with no combination benefit | Direct call to the single service |
10. Hands-On Exercise
Extend it: add a fourth dependency, loyalty-service, providing the customer’s points earned from this order — classify it explicitly as critical or optional, justify the choice, and implement it following Section 5’s pattern.
Simulate a failure: make both payment-service and logistics-service fail simultaneously for one request, and confirm the aggregator correctly fails the whole response (since payment-service is critical) rather than attempting a partial response that happens to also be missing the optional field — the aggregation logic shouldn’t need special-casing for multiple simultaneous failures if the criticality classification is implemented correctly.
Decision question, with justification required: product now wants the confirmation screen to show a warning banner (“delivery estimate temporarily unavailable”) specifically when logistics-service failed, rather than just omitting the field silently. Does this change belong in mobile-bff’s aggregation logic, or in the mobile client’s own rendering logic, given the deliveryEstimateAvailable flag already exists? Justify using the separation of concerns established across this series (Chapter 4’s BFF-owns-the-contract principle in particular).
11. Key Takeaways
- API composition fans out to multiple downstream services in parallel and merges their results — its real design work is the merge policy, not the fan-out mechanics.
- Classify every aggregated field’s criticality explicitly and deliberately, ideally with product input — this determines whether that field’s failure should degrade the response or fail it entirely.
- Pair every optional field with an explicit availability signal; a bare null is ambiguous between “doesn’t exist” and “couldn’t be fetched right now,” and those need different client treatment.
- Issue independent calls in parallel, never sequentially, unless a genuine data dependency requires it — sequential calls needlessly sum latencies that parallel calls would only take the maximum of.
- Apply per-call resilience (Chapter 16) to every fanned-out call individually — an aggregator without this just relocates the cascading-failure risk rather than solving it.
- Not every composed view should have a partial-failure policy — some views are only meaningful as a complete whole, and fail-all-or-nothing is sometimes the correct, deliberately chosen policy.
- Keep aggregation logic specific to its actual consumer; don’t generalize into reusable aggregation infrastructure before a second genuine use case justifies it.