Part 5: Validation as a first-class feature
A DSL that generates XML you only hope is valid is a code generator with a faith problem. Validation here is not a linter bolted on afterward — it is the stage that makes the DSL trustworthy, and it runs on the immutable AST before any XML exists. This part builds BpmnValidator, which returns structured diagnostics rather than throwing on the first problem.
Structured diagnostics
A diagnostic is data, not an exception. It carries enough to print a useful message and to feed an IDE-style report or a CI gate:
enum class Severity { ERROR, WARNING }
data class Diagnostic( /** Stable machine-readable code, e.g. "BPMN-FLOW-001". */ val code: String, val severity: Severity, val message: String, /** Element the diagnostic is attached to. */ val elementId: String? = null, /** Other elements involved, e.g. both ends of an unresolved reference. */ val relatedIds: List<String> = emptyList(), /** Practical suggestion for fixing the problem. */ val remediation: String? = null,)Stable codes matter more than polished messages: tests assert on codes, CI can suppress a specific warning, and documentation can reference BPMN-FLOW-001 without depending on prose. The distinction between ERROR (output would be invalid or undeployable) and WARNING (structurally legal but suspicious — unreachable node, boundary event with no handler) is what lets validation gate a build without crying wolf.
The shape of the validator
class BpmnValidator {
fun validate(collaboration: CollaborationDefinition): List<Diagnostic> { val d = mutableListOf<Diagnostic>() validateIdentifiers(collaboration, d) for (process in collaboration.processes) { validateProcess(collaboration, process, d) } validateCollaboration(collaboration, d) return d.sortedWith(compareBy({ it.severity }, { it.code }, { it.elementId })) }}Deterministic ordering of diagnostics — same model, same list, same order — is a small discipline that pays off in tests and review diffs.
Identifier validation
BPMN ids are XML IDs: unique across the entire document, not just within a process. The check is a single pass over every id-bearing element:
private fun validateIdentifiers(c: CollaborationDefinition, d: MutableList<Diagnostic>) { val seen = LinkedHashMap<String, String>()
fun check(id: String, kind: String) { if (id.isBlank()) { d += /* BLANK_ID error */; return } val previous = seen.putIfAbsent(id, kind) if (previous != null) { d += Diagnostic( DiagnosticCodes.DUPLICATE_ID, Severity.ERROR, "Duplicate id '$id' used by both $previous and $kind; " + "BPMN ids are XML IDs and must be unique document-wide", elementId = id, remediation = "Rename one of the elements.", ) } }
check(c.id, "collaboration") c.participants.forEach { check(it.id, "participant") } c.processes.forEach { p -> check(p.id, "process") p.lanes.forEach { check(it.id, "lane") } p.allNodes().forEach { check(it.id, "flow node") } p.allSequenceFlows().forEach { check(it.id, "sequence flow") } } c.messages.forEach { check(it.id, "message") } c.messageFlows.forEach { check(it.id, "message flow") }}allNodes() and allSequenceFlows() — the flattened views that include subprocess contents — are what make document-wide uniqueness cheap to check.
Process validation: one pass per scope
validateScope runs once for the process and once for every embedded subprocess — the subprocess interior is a closed scope with its own start/end requirements:
private fun validateScope( scopeId: String, nodes: List<FlowNode>, flows: List<SequenceFlow>, d: MutableList<Diagnostic>,) { val byId = nodes.associateBy { it.id }
if (nodes.none { it is StartEvent }) d += /* NO_START_EVENT error */ if (nodes.none { it is EndEvent }) d += /* NO_END_EVENT error */
val outgoing = flows.groupBy { it.sourceRef } val incoming = flows.groupBy { it.targetRef } // …flow resolution, direction rules, reachability…}Reference resolution. A sequence flow whose endpoints don’t resolve within its own scope is an error — this also catches cross-scope smuggling, because a subprocess flow referencing a parent node won’t resolve in the subprocess’s byId map, and vice versa.
Direction rules. Start events can’t have incoming flows; end events can’t have outgoing; boundary events can’t have incoming (they’re entered via attachment) and should have outgoing — a boundary event with nowhere to send its token fires and discards.
for (n in nodes) { val out = outgoing[n.id].orEmpty() val inc = incoming[n.id].orEmpty() when (n) { is StartEvent -> if (inc.isNotEmpty()) d += directionError(n.id, "start event", "incoming") is EndEvent -> if (out.isNotEmpty()) d += directionError(n.id, "end event", "outgoing") is BoundaryEventNode -> { if (inc.isNotEmpty()) d += directionError(n.id, "boundary event", "incoming") if (out.isEmpty()) d += Diagnostic( DiagnosticCodes.BOUNDARY_NO_OUTGOING, Severity.WARNING, "Boundary event '${n.id}' has no outgoing sequence flow; it will fire and discard the token", elementId = n.id, ) } else -> Unit }}Reachability. A node that no start event can reach is dead modeling — usually a forgotten edge. Computed by DFS over outgoing edges:
val reachable = mutableSetOf<String>()val stack = ArrayDeque<String>()fun drain() { while (stack.isNotEmpty()) { val cur = stack.removeLast() if (!reachable.add(cur)) continue outgoing[cur].orEmpty().forEach { stack.addLast(it.targetRef) } }}nodes.filterIsInstance<StartEvent>().forEach { stack.addLast(it.id) }drain()// Boundary events are entered via attachment, not sequence flow: they are// reachable whenever their host activity is reachable.nodes.filterIsInstance<BoundaryEventNode>() .filter { it.attachedToRef in reachable } .forEach { stack.addLast(it.id) }drain()Common pitfall: the first version of this pass seeded only start events, and every node downstream of a boundary event was flagged unreachable. Boundary events have no incoming edges by design — treat reachable hosts as reachability roots for their boundary events. The same two-pass trick works in reverse for “can reach an end event” (walk
incomingbackwards).
Gateway validation
Three checks encode gateway semantics:
if (g is ExclusiveGateway || g is InclusiveGateway) { g.defaultFlow?.let { def -> if (out.none { it.id == def }) { d += /* DEFAULT_FLOW_INVALID: default must name an outgoing flow */ } } if (out.size > 1) { val undecided = out.filter { it.condition == null && it.id != g.defaultFlow } if (undecided.isNotEmpty()) d += /* MISSING_CONDITION warning */ }}if (g is ParallelGateway) { val conditioned = out.filter { it.condition != null } if (conditioned.isNotEmpty()) d += /* CONDITION_FORBIDDEN error */}Conditions on outgoing flows are also checked at the edge level: BPMN permits conditionExpression only on flows leaving activities and exclusive/inclusive gateways. A conditioned flow leaving a parallel gateway is an error (the condition is silently ignored by engines — modeling intent that evaporates). Conditions leaving an event or a basic receive/send chain are equally invalid.
The limit worth stating: this validator does not simulate token semantics. It cannot prove that a complex net of inclusive joins never deadlocks — only that the obvious shapes (unreachable joins, starved defaults, conditioned parallel splits) are absent. Full token-flow verification is a model-checking problem; Part 10 lists it as an extension.
Lane and collaboration validation
Lanes check three things: every flowNodeRef resolves to a top-level node of the same process; no node claims membership in two lanes; and — by our layout convention — every top-level node is in a lane (warning, not error, because a laneless node is legal BPMN even if unrenderable in our layout). Boundary events are exempt from the lane rule: they render attached to their host and inherit its lane visually.
for (mf in c.messageFlows) { mf.messageRef?.let { ref -> if (ref !in messageIds) d += /* MSG_REF_UNRESOLVED */ } val srcPool = participantOfNode[mf.sourceRef] ?: mf.sourceRef.takeIf { it in participantIds } val tgtPool = participantOfNode[mf.targetRef] ?: mf.targetRef.takeIf { it in participantIds } if (srcPool == null || tgtPool == null) d += /* MSG_FLOW_ENDPOINT: endpoint resolves to nothing */ if (srcPool == tgtPool) d += Diagnostic( DiagnosticCodes.MSG_FLOW_SAME_POOL, Severity.ERROR, "Message flow '${mf.id}' connects two elements inside participant '$srcPool'; " + "message flows must cross participant boundaries (use a sequence flow inside one process)", elementId = mf.id, relatedIds = listOf(mf.sourceRef, mf.targetRef), )}Collaboration checks also reject gateways as message endpoints (a gateway can’t send or receive a message — it’s control flow, not a participant in one) and warn when no process is executable, which is legal BPMN but pointless for deployment.
Boundary and subprocess checks
when (val host = byId[b.attachedToRef]) { null -> d += /* BOUNDARY_TARGET: host doesn't exist in this scope */ is ActivityNode -> Unit // the only legal target else -> d += Diagnostic( DiagnosticCodes.BOUNDARY_TARGET, Severity.ERROR, "Boundary event '${b.id}' attaches to '${b.attachedToRef}', which is a ${host::class.simpleName}; " + "boundary events may only attach to activities (tasks and subprocesses)", elementId = b.id, remediation = "Attach the boundary event to an ActivityNode.", )}This is the check that makes Part 4’s design decision enforceable: attach a boundary timer to a catch event and you get BPMN-BND-001, not a corrupted diagram.
Diagram validation — a separate pass
validateDiagram(collaboration, layout) runs after layout and checks the DI half of the contract:
fun validateDiagram(c: CollaborationDefinition, layout: DiagramLayout): List<Diagnostic> { val d = mutableListOf<Diagnostic>() for (p in c.processes) { for (n in p.allNodes()) { val shape = layout.shapes[n.id] ?: run { d += /* DI_MISSING_SHAPE error */; continue } if (shape.width <= 0 || shape.height <= 0) d += /* DI_BAD_BOUNDS */ } // lane ⊆ participant containment for (lane in p.lanes) { val laneBounds = layout.shapes[lane.id] ?: continue val poolBounds = participant?.let { layout.shapes[it.id] } if (poolBounds != null && !poolBounds.contains(laneBounds)) d += /* DI_LANE_OUTSIDE_POOL */ for (ref in lane.flowNodeRefs) { val nodeBounds = layout.shapes[ref] ?: continue if (!laneBounds.contains(nodeBounds)) d += /* DI_NODE_OUTSIDE_LANE */ } } // boundary event must touch its host's border for (b in p.allNodes().filterIsInstance<BoundaryEventNode>()) { val bs = layout.shapes[b.id]; val host = layout.shapes[b.attachedToRef] if (bs != null && host != null && !bs.intersects(host)) d += /* DI_BOUNDARY_DETACHED */ } for (f in p.allSequenceFlows()) checkEdge(f.id, layout, d) } for (mf in c.messageFlows) checkEdge(mf.id, layout, d) return d}Every visible semantic node gets a shape; every sequence flow and message flow gets an edge with ≥2 waypoints. These are rendering invariants — a model can be semantically perfect and still open as a blank canvas without them.
Testing the validator — on purpose-broken models
The test suite proves detection, not just acceptance:
@Testfun `conditions after a parallel gateway are rejected`() { val c = bpmnCollaboration("badPar") { participant("pool") { executable = true process("proc") { lane("work") { startEvent("start"); parallelGateway("split") serviceTask("a") {}; serviceTask("b") {} parallelGateway("join"); endEvent("done") } sequenceFlow("f0", "start", "split") conditionalFlow("f1", "split", "a", "\${x}") sequenceFlow("f2", "split", "b") sequenceFlow("f3", "a", "join"); sequenceFlow("f4", "b", "join") sequenceFlow("f5", "join", "done") } } } assertTrue(DiagnosticCodes.CONDITION_FORBIDDEN in codes(c))}
@Testfun `message flows may not connect two nodes in the same pool`() { /* … MSG_FLOW_SAME_POOL … */ }The boundary-attachment test is more interesting because the DSL can’t express the invalid case — that’s the point of the type system — so the test constructs the invalid AST directly:
@Testfun `boundary events may not attach to events or gateways`() { val invalid = /* …a process whose "wait" is an IntermediateCatchMessageEvent… */ val badNode = BoundaryTimerEvent( id = "badBoundary", attachedToRef = "wait", timer = TimerDefinition.duration("PT1H"), ) val patched = invalid.copy( processes = listOf(proc.copy(nodes = proc.nodes + badNode)), ) assertTrue(DiagnosticCodes.BOUNDARY_TARGET in codes(patched))}This is the honest division of labor: the DSL prevents invalid models at compile time; the validator catches the ones that can be expressed — wrong references, wrong endpoints, unresolved shapes — and the tests prove each one is actually caught.
Common pitfalls
- Validating per-process only — subprocesses are their own scope; skipping their interior check lets a subprocess exist with no start event.
- Warning fatigue — if everything is an error, nothing is. Reserve
ERRORfor invalid output;WARNINGfor “legal but probably not what you meant.” - Assertions instead of diagnostics — an
assert/checkthrown on the first problem reports one issue at a time. Collect all diagnostics so a single run prints every problem. - Only validating semantics — a model can be semantically valid and have a broken or missing diagram. Both halves need checks.
Reference implementation
BpmnValidator, Diagnostic, and the full DiagnosticCodes registry live in Validation.kt in the bpmn-validation module of pcnixsys/bpmn-dsl-kotlin. The procurement example validates with zero errors and zero warnings; the test suite carries a negative test per diagnostic family.
Next
Part 6 writes the validated AST out as BPMN 2.0 XML using StAX — namespaces, incoming/outgoing bookkeeping, conditionExpression, event definitions, and the extension-element hook the Flowable adapter plugs into.