About

Twenty years on one question: how decisions get made when the evidence is incomplete. First in brains, now in systems.

Today the systems are at Commonwealth Bank: customer-facing agents for investment education and trading guidance, along with scam detection, entity resolution, and the safety and risk evaluation a regulated environment requires. I design them, contribute code to them, and design the evaluations that tell us how they are behaving; leading a team that builds things has never meant giving up building them.

Most of the craft lives in a tension: a customer-facing agent has to be fast and cheap enough to be worth using, and safe and trustworthy enough to be allowed near a customer at all. Those pull against each other, and where you land is a design decision rather than something you discover afterwards. The rest of the job is keeping the person who will actually use the thing in the room while it is being designed.

I work this way because of where I started. For a decade before industry I studied human decisions with psychophysics: the person under study cannot tell you how they decide, so you present controlled evidence, record the choices, and infer the mechanism from the pattern of errors. The research I return to most came out of that, begun at RIKEN in Japan and finished at Stanford: people do not weigh each new piece of evidence afresh when they decide. They carry their last few trials with them, and failure weighs differently from success. PNAS published it in 2016.

The crossing went through AI Labs, the bank's applied research team: hard problems taken through to production, novel enough that three papers came out of them. Leaving academia for industry changed the subject matter, not the method.

Underneath both halves sits the same question: when does a system have enough evidence to answer, and when should it hand the decision to a person? Psychophysics trains a useful distinction here, between how well a system discriminates and how willing it is to say yes, which is the thing accuracy alone misses. It also trains a habit: when something resists solving, it is rarely the method that is wrong. It is usually a decision made early, about what counts or who this is for, that nobody has gone back to check.

I am based in Sydney, Australia, and if you are building AI systems that help people, email is the best way to start a conversation.