Five Agents, Five Documents: How We Build Payment Software with AI at DOKU
The AI-assisted development pipeline I walked through at Kelas Beta: one agent per role, every output a document in git, and a human gate between each step.
On 23 September I gave a session at Sekolah Beta’s Kelas Beta (Hacker track) at Garuda Spark Innovation Hub in Jakarta, titled How to Design AI Systems for Digital Payments. The room was backend engineers who’d already integrated a payment gateway or two, came in at 7 pm after a full day of work, and had heard “AI plus payments” enough times to be suspicious of it. Fair enough.
I split the talk into three places AI can show up in a payment system: building it, operating it, and deciding inside the transaction flow. This post covers the first one, which is also the lowest-risk one. When AI helps you build a payment system it touches documents and code, never money. What’s at stake is quality, and quality has a well-understood control: review.
One agent per role, not one agent for everything
The pipeline we use internally is built on Claude Code, and it’s organised around the roles that already exist on a delivery team rather than around the model:
| Step | Who runs it | What comes out |
|---|---|---|
/new-feature |
Product owner | functional-spec.md |
/tech-design |
Solution architect | design.md and openapi.yml |
/impl-plan |
Tech lead | an ordered implementation plan |
/test-plan |
QA | test cases, run in parallel with the plan |
/implementation |
Developer | code and tests |
Five steps, and the thing I’d point at first is the right-hand column. Every agent’s output is a document a person can read, argue with and diff, and every one of those documents goes into git next to the code. The implementation agent is the only one that writes code, and by the time it runs, it’s working from a spec, a design, an API contract and a test plan that four different people have already signed off.
I think this is the part most teams skip when they “adopt AI”: they give one assistant to every engineer and let each of them prompt their way from a ticket straight to a pull request. You get speed, but the decisions that used to be written down (what the feature is for, why this design over the other one, which failure modes the tests must cover) now live in a chat history nobody else will ever open.
The gate is the point
Each step has three parts: the agent generates, an AI reviewer critiques the output against the previous document, and then a human approves or sends it back. The human approval is not optional. That’s a rule of the process, and it’s also the reason I’m comfortable using this for software that moves real money.
A concrete example of what a gate catches: the tech-design reviewer checks that every acceptance criterion in the functional spec maps to something in the design. If the PO wrote “a merchant can’t be charged twice for the same invoice” and the design has no idempotency key anywhere, that’s a finding before a single line of code exists. Catching it at that stage costs a comment. Catching it in production costs a reconciliation incident and an apology to a merchant.
The honest counter to all this is that it’s a lot of ceremony. Five documents and five approvals for a feature can feel heavier than just writing the code, especially for a small change, and a senior engineer could reasonably say they’d get there faster alone. I don’t think that’s wrong for small changes (I wouldn’t run all five steps for a one-line fix either), but for anything that touches a transaction flow, I’d rather pay the ceremony up front than discover the missing decision later.
What holds it together
If you take one thing from this layer, I’d make it this: the model isn’t what makes the pipeline safe. A smarter model writes a more convincing spec, and a more convincing spec is harder to review, not easier. What makes it safe is structural. Every stage stops at a person, and every stage’s output is a file in version control that the next stage has to read.
Two pieces of that structure deserve a mention:
CLAUDE.mdas a permanent contract. Every repository carries one: the stack, the conventions, the things the agent must never do. It’s written once and read on every run, which is cheaper and more reliable than repeating the same constraints in every prompt. I wrote about this in more detail in Spec Before Code.- Documents over chat. A decision made in a conversation with an agent is gone when the session ends. A decision in
design.mdgets reviewed, versioned and quoted back in the next incident review.
The number I didn’t put on a slide
There’s an internal estimate floating around for how much faster this pipeline is than doing the same work by hand. I left it out of the talk, and I’m leaving it out here, because it isn’t measured. There’s no sample size, no baseline, and no agreed definition of “by hand”. Kelas Beta’s audience would’ve taken it apart in the Q&A, and they’d have been right to.
What I can say without hedging is narrower: the documents exist, they’re reviewable, and decisions that used to live in someone’s head now live in git. Whether that makes us faster in the long run, I’m honestly not sure yet. I’d want a few quarters of real delivery data before claiming it, and I’d be suspicious of anyone who claims a multiplier without one.
The cost is real too. Writing and reviewing a functional spec takes the PO most of an afternoon, and on a busy week that’s an afternoon nobody wants to spend. We spend it anyway for anything near money.
Where this leads
Building is the safe layer because nothing in it can move a rupiah. The next post is about the layer where that stops being true, when an agent calls a real payment system, and what DOKU built an MCP server for. The third walks through the demo the room could try from their own phones that night.
Series from my Kelas Beta session, 23 September 2026: 1 · AI-assisted development · 2 · What DOKU’s MCP server is for · 3 · A chat cashier that can’t count money