Series overview
Part 6 of 1060% complete
2026-09-20•4 min read

Part 6: Generating BPMN 2.0 XML safely from the AST

With a validated AST in hand, serialization is almost mechanical — which is the point. Every design decision was made upstream; this part is about writing XML correctly: no string concatenation, deterministic output, and a clean seam for engine-specific extensions.

Why StAX, not string templates

XMLStreamWriter (JDK built-in, javax.xml.stream) handles escaping, quoting, and namespace bookkeeping. String concatenation looks tempting for XML this regular — right up until a task name contains & or a condition expression contains <, and your “generator” ships malformed XML. StAX also streams: a 500-node process doesn’t materialize anything but the output characters.

The constraint StAX adds in non-repairing mode — where namespaces are written manually — is discipline: declare every prefix on the root once, write every element with an explicit (prefix, localName, namespace) triple, and the output is deterministic because nothing is inferred.

Namespaces and the definitions root

xml/BpmnXmlWriter.kt
object BpmnNs {
const val BPMN = "http://www.omg.org/spec/BPMN/20100524/MODEL"
const val BPMNDI = "http://www.omg.org/spec/BPMN/20100524/DI"
const val DC = "http://www.omg.org/spec/DD/20100524/DC"
const val DI = "http://www.omg.org/spec/DD/20100524/DI"
const val XSI = "http://www.w3.org/2001/XMLSchema-instance"
const val FLOWABLE = "http://flowable.org/bpmn"
const val CAMUNDA = "http://camunda.org/schema/1.0/bpmn"
}
xml/BpmnXmlWriter.kt
w.writeStartDocument("UTF-8", "1.0")
w.writeStartElement("bpmn", "definitions", BpmnNs.BPMN)
w.writeNamespace("bpmn", BpmnNs.BPMN)
w.writeNamespace("bpmndi", BpmnNs.BPMNDI)
w.writeNamespace("dc", BpmnNs.DC)
w.writeNamespace("di", BpmnNs.DI)
w.writeNamespace("xsi", BpmnNs.XSI)
w.writeNamespace("bpmndsl", "https://o612.in/bpmn/dsl")
for ((prefix, uri) in engine.namespaces()) {
w.writeNamespace(prefix, uri)
}
w.writeAttribute("id", "${c.id}_definitions")
w.writeAttribute("name", c.name ?: c.id)
w.writeAttribute("targetNamespace", targetNamespace)

bpmndsl is the library’s own extension namespace — process variables serialize there. Engine namespaces come from the EngineExtensionSerializer (Part 9), so the writer emits xmlns:flowable only when the Flowable adapter is in use.

Document skeleton

BPMN’s definitions element is a flat container. Order of the major groups doesn’t affect validity, but a fixed order makes output diffable:

xml/BpmnXmlWriter.kt
writeMessages(c, w) // <bpmn:message> declarations
writeErrors(c, w) // <bpmn:error> for errorEventDefinition refs
writeCollaboration(c, w) // participants + message flows
for (p in c.processes) {
writeProcess(c, p, w) // executable and non-executable processes
}
if (layout != null) {
writeDiagram(c, layout, w) // Part 7
}
Generated excerpt
<bpmn:message id="rfqRequested"/>
<bpmn:message id="supplierQuotationReceived"/>
<bpmn:collaboration id="enterpriseProcurement" name="Enterprise Procurement and Supplier Onboarding">
<bpmn:participant id="buyerOrganization" name="Buyer Organization" processRef="buyerProcurementProcess"/>
<bpmn:participant id="supplier" name="Supplier" processRef="supplierInteractionProcess"/>
<bpmn:messageFlow id="msgRfqToSupplier" sourceRef="sendRfq" targetRef="receiveRfq" messageRef="rfqRequested"/>
<!-- …six more message flows… -->
</bpmn:collaboration>

Process: laneSet before flow elements

Inside <bpmn:process>, the schema sequence places laneSet ahead of flowElements. The writer emits documentation, the variables extension, the lane set, then nodes and flows:

xml/BpmnXmlWriter.kt
w.writeStartElement("bpmn", "process", BpmnNs.BPMN)
w.writeAttribute("id", p.id)
p.name?.let { w.writeAttribute("name", it) }
w.writeAttribute("isExecutable", p.isExecutable.toString())
writeDoc(p.documentation, w)
writeProcessVariables(p, w)
writeLaneSet(p, w)
writeScopeNodes(p.nodes, p.sequenceFlows, w)
writeScopeFlows(p.sequenceFlows, w)
w.writeEndElement()
Generated excerpt
<bpmn:process id="buyerProcurementProcess" name="Buyer Procurement Process" isExecutable="true">
<bpmn:extensionElements>
<bpmndsl:processVariables>
<bpmndsl:variable name="requisitionId" type="string" required="true"/>
<!-- … -->
</bpmndsl:processVariables>
</bpmn:extensionElements>
<bpmn:laneSet id="buyerProcurementProcess_laneSet">
<bpmn:lane id="requestingDepartment" name="Requesting Department">
<bpmn:flowNodeRef>requisitionCreated</bpmn:flowNodeRef>
<bpmn:flowNodeRef>submitRequisition</bpmn:flowNodeRef>
<bpmn:flowNodeRef>correctRequisition</bpmn:flowNodeRef>
</bpmn:lane>
<!-- …five more lanes… -->
</bpmn:laneSet>

The node writer — the one function that does the mapping

Every flow node serializes through a single when on the sealed hierarchy — adding a node type without updating the writer is a compile error, which is exactly what when exhaustiveness buys you:

xml/BpmnXmlWriter.kt
val tag = when (n) {
is StartEvent -> "startEvent"
is EndEvent -> "endEvent"
is IntermediateCatchMessageEvent -> "intermediateCatchEvent"
is BoundaryEventNode -> "boundaryEvent"
is UserTask -> "userTask"
is ServiceTask -> "serviceTask"
is ReceiveTask -> "receiveTask"
is SendTask -> "sendTask"
is EmbeddedSubprocess -> "subProcess"
is ExclusiveGateway -> "exclusiveGateway"
is ParallelGateway -> "parallelGateway"
is InclusiveGateway -> "inclusiveGateway"
is EventBasedGateway -> "eventBasedGateway"
}
w.writeStartElement("bpmn", tag, BpmnNs.BPMN)
w.writeAttribute("id", n.id)
n.name?.let { w.writeAttribute("name", it) }
if (n is BoundaryEventNode) {
w.writeAttribute("attachedToRef", n.attachedToRef)
if (!n.cancelActivity) w.writeAttribute("cancelActivity", "false")
}
if (n is ReceiveTask) n.messageRef?.let { w.writeAttribute("messageRef", it) }
if (n is SendTask) n.messageRef?.let { w.writeAttribute("messageRef", it) }
if (n is ExclusiveGateway || n is InclusiveGateway) {
n.defaultFlow?.let { w.writeAttribute("default", it) }
}
for ((k, v) in engine.elementAttributes(n)) { /* vendor attrs, Part 9 */ }

incoming/outgoing bookkeeping

BPMN flow elements redundantly declare their edges — <bpmn:incoming>/<bpmn:outgoing> children repeat what sourceRef/targetRef on the flows already say. Engines don’t need them; strict tools warn without them. The writer derives them rather than trusting the author:

xml/BpmnXmlWriter.kt
val incoming = flows.filter { it.targetRef == n.id }
val outgoing = flows.filter { it.sourceRef == n.id }
// …
for (f in incoming) { w.writeStartElement("bpmn","incoming",BpmnNs.BPMN); w.writeCharacters(f.id); w.writeEndElement() }
for (f in outgoing) { w.writeStartElement("bpmn","outgoing",BpmnNs.BPMN); w.writeCharacters(f.id); w.writeEndElement() }

Redundancy that is computed, never authored, cannot drift out of sync.

Event definitions

Event definitions are nested elements that turn a generic element into a typed event — <bpmn:boundaryEvent> plus <bpmn:timerEventDefinition> is a timer boundary event; plus messageEventDefinition is a message boundary event:

xml/BpmnXmlWriter.kt
is BoundaryTimerEvent -> {
w.writeStartElement("bpmn", "timerEventDefinition", BpmnNs.BPMN)
w.writeAttribute("id", "${n.id}_timer")
writeTimerBody(n.timer, w)
w.writeEndElement()
}
xml/BpmnXmlWriter.kt
private fun writeTimerBody(t: TimerDefinition, w: XMLStreamWriter) {
val (tag, value) = when {
t.duration != null -> "timeDuration" to t.duration
t.cycle != null -> "timeCycle" to t.cycle
else -> "timeDate" to t.date
}
w.writeStartElement("bpmn", tag, BpmnNs.BPMN)
w.writeAttribute("xsi", BpmnNs.XSI, "type", "bpmn:tFormalExpression")
w.writeCharacters(value!!)
w.writeEndElement()
}
Generated excerpt
<bpmn:boundaryEvent id="quotationSlaExceeded" name="72h supplier SLA" attachedToRef="waitForQuotation">
<bpmn:outgoing>flowSlaToReminder</bpmn:outgoing>
<bpmn:timerEventDefinition id="quotationSlaExceeded_timer">
<bpmn:timeDuration xsi:type="bpmn:tFormalExpression">PT72H</bpmn:timeDuration>
</bpmn:timerEventDefinition>
</bpmn:boundaryEvent>

Condition expressions

xml/BpmnXmlWriter.kt
f.condition?.let { cond ->
w.writeStartElement("bpmn", "conditionExpression", BpmnNs.BPMN)
w.writeAttribute("xsi", BpmnNs.XSI, "type", "bpmn:tFormalExpression")
if (cond.language != "juel") w.writeAttribute("language", cond.language)
w.writeCharacters(cond.body)
w.writeEndElement()
}
Generated excerpt
<bpmn:sequenceFlow id="flowBudgetBranch" name="budget approval" sourceRef="approvalsSplit" targetRef="approveBudget">
<bpmn:conditionExpression xsi:type="bpmn:tFormalExpression">${budgetApprovalRequired}</bpmn:conditionExpression>
</bpmn:sequenceFlow>

The extension seam

EngineExtensionSerializer is the one seam through which vendor-specific output enters otherwise-standard XML:

xml/BpmnXmlWriter.kt
interface EngineExtensionSerializer {
fun namespaces(): Map<String, String> = emptyMap()
fun elementAttributes(node: FlowNode): Map<String, String> = emptyMap()
fun extensionElements(node: FlowNode): List<ExtensionElement> = emptyList()
object None : EngineExtensionSerializer
}

The writer consults it in three places: extra xmlns: declarations on definitions, prefixed attributes on each node element (flowable:candidateGroup="procurement-officers"), and children inside <bpmn:extensionElements> (for structured extensions like flowable:field). ExtensionElement itself serializes recursively — prefix, local name, attributes, children, text:

xml/BpmnXmlWriter.kt
private fun writeExtensionElement(e: ExtensionElement, w: XMLStreamWriter) {
if (e.prefix != null) {
val ns = engine.namespaces()[e.prefix]
?: error("Extension prefix '${e.prefix}' has no registered namespace")
w.writeStartElement(e.prefix, e.name, ns)
} else {
w.writeStartElement("bpmn", e.name, BpmnNs.BPMN)
}
for ((k, v) in e.attributes) w.writeAttribute(k, v)
e.text?.let { w.writeCharacters(it) }
e.children.forEach { writeExtensionElement(it, w) }
w.writeEndElement()
}

The error() on an unknown prefix is deliberate fail-fast: a malformed namespace is worse than a crash at authoring time.

Determinism

Everywhere the writer iterates, it iterates the AST’s declaration order — LinkedHashMap and List preserve it end to end. Running the writer twice on the same input produces byte-identical output, which Part 8 turns into a test. This is not a nice-to-have: diffs of generated XML are how generated workflows get reviewed.

Verification

bpmn-moddle — the parser that backs bpmn-js — reads the generated file with zero warnings. More meaningfully for correctness: BPMN semantic XML without DI is deployable but renders as a blank canvas; with DI it both deploys and draws. That requires Part 7.

Common pitfalls

  • Writing attachedToRef on the activity instead of the boundary event — the attribute belongs to the boundaryEvent element pointing at its host.
  • Skipping incoming/outgoing — some strict parsers warn or reject; compute them from the flow list.
  • String-concatenated XML — first & in a label and the file is malformed. Use a real writer.
  • Declaring xmlns:flowable unconditionally — emit engine namespaces only when the adapter actually uses them.
  • Emitting nodes and flows interleaved with laneSet — laneSet precedes flowElements in the schema.

Reference implementation

The StAX writer — BpmnNs, the Xml element wrapper, EngineExtensionSerializer, and writeToString — is BpmnXmlWriter.kt in the bpmn-xml module of pcnixsys/bpmn-dsl-kotlin.

Next

The semantic XML alone would open in a viewer as nothing — all structure, no coordinates. Part 7 builds the layout engine that produces BPMN DI: columns from longest-path layering, rows from lanes, boundary events pinned to their hosts, and orthogonal edge routing — deterministic, complete, and honest about its limits.

KotlinBPMN

Type to search the site.

↑↓ navigate⏎ openPowered by Pagefind