At Commonwealth Bank I lead agentic AI in Retail Banking Services. Anything we ship has to do what it is meant to do, excel at it, and be safe, compliant and trustworthy in a regulated environment. That is built by design and driven by evals, not discovered afterwards.
It is the best job I could have designed for myself: agentic systems are complex systems, and complex systems are what I am trained to probe, understand and align. A customer-facing agent has to be fast and cheap enough to be worth using, and safe and trustworthy enough to be allowed near a customer. Where it lands between those two is an architectural decision, and architecting agentic systems is where I am at my best.
The build runs on two things: design and evals, with squads of AI agents doing the building. The design is the architecture. The evals guarantee compliance, safety and trust before anything ships, and they take whatever form the system demands: behavioural, regulatory, and beyond. The methods for building this way are new, and developing them is part of the work.
The method is scientific: assumptions are hypotheses, and hypotheses get tested. The trade-offs hide in specifics. A small language model can be expensive when it holds a whole GPU for two thousand calls a month. A subagent that looks free in isolation adds latency to every request. Retrieval is easy for simple things and a different beast at the edge cases and in compliance.
"How do you know a system this complex is doing what it is meant to do?"
I have spent my career asking careful questions of things that cannot explain themselves. For a decade that was the brain: human, subtle, the most complex system we know. A person cannot tell you how they decide, so you present controlled evidence, record the choices, and infer what is happening from the pattern of errors. An agentic system is a machine, not a mind, but it also acts without being able to account for itself. It answers to the same kind of asking, and I had a decade of training in how to ask.
The study I return to most began at RIKEN in Japan and finished at Stanford, and was published in PNAS in 2016: people carry their last few trials with them, and failure weighs differently from success. Agents, in their own way, do the same. Behaviour drifts with what accumulates, and catching that is exactly what behavioural evals are for.
I am based in Sydney, Australia. If you are putting AI to work on problems that matter to real people, email is the best way to reach me.