No Winning Square
Who wears the risk when an AI agent gets it wrong.
In October 2025, Deloitte handed money back to the Australian government. The firm was paid $440,000 for an assurance review of the welfare compliance system, and the report shipped with academic references that didn’t exist, including papers attributed to a real Sydney University professor who had never written them, and a quote from a Federal Court judgment that nobody ever wrote. Dr Chris Rudge, a University of Sydney researcher, caught it, noting the first fabricated reference as either “an AI hallucination or the world's best kept secret.” Deloitte returned the final instalment, quietly republished the report with a new appendix disclosing that a generative model had been used, and declared the matter resolved. The long-term consequences remain to be seen, but as Peter Evans-Greenwood recently wrote, the case signals a significant erosion of the legitimacy of the Big Four consulting model.
Eighteen months earlier, Air Canada ran the same play with less grace. Its chatbot told a grieving customer he could book a full-fare flight and claim the bereavement discount afterwards — a policy the airline has never offered and one that contradicted the real policy hyperlinked in the bot’s answer. When he made the claim, the airline refused, and before a tribunal it made two submissions that deserve framing: (1) that the chatbot was a separate legal entity responsible for its own actions, and (2) that it shouldn’t have been trusted in the first place. The tribunal ordered the airline to pay, calling the first argument “remarkable”, and on the second, found that a customer has no obligation to double-check one part of a company’s website against another.
Two brands, two agentic workflows, and a single pattern: a machine erred, and the brand wore the consequences. One layer deeper, it signals an erosion of trust, with the blame landing squarely on the brand.
Both success and failure need a target
Let’s take agent-mediated brand interactions and sort them on two axes: (1) who owns the agent making the call, and (2) whether the interaction was successful.
| Unsuccessful interaction | Successful interaction | |
Customer’s agent | “[brand] misled my assistant” Outcome: Brand wears the blame | “My assistant knows me” Outcome: Agent gets the credit |
Brand-owned agent | “[brand] is terrible” Outcome: Not only did you fail, but your systems are also awful | “Thanks for doing your job” Outcome: The minimum level of expectation |
Caption: Agentic commerce is perhaps the worst version of two-up going for retailers
When a customer’s own AI assistant recommends something that disappoints, the early research suggests that blame lands on the brand, rather than the model that steered them wrong. When the brand’s own agent errs, blame lands harder on the brand — Deloitte and Air Canada live in this square, and consumers hold brands almost universally responsible for what their chatbots say and do. And if the recommendation is successful? Well, obviously that just means my assistant is brilliant! The brand’s efforts more or less disappear. However, that brilliance doesn’t carry over to a brand’s AI. That’s just the new minimum expected level of service.
So, blame, more blame, someone else gets the credit, or no credit. There is no square in which the brand builds trust through this exchange.
Merchants are beginning to concede this in public: at a Fortune conference in June, one executive acknowledged that terms and conditions might shield a company from legal liability for an agent’s errors, but not the “perceptual liability”, which still hits hard. Lawyers are very good at moving indemnity. But blame is fundamentally a trust game.
Why the grid has no winning square
The empty square isn’t bad luck. It’s the direct consequence of what delegation does to trust. Returning to the recent lament on replacing my humble toastie machine, trust is not transactional; it is a relational process:
- Interpretation, the good reasons you can gather in advance;
- Suspension, the leap you take despite not knowing a product suits your needs until you use it; and
- Expectation, the settled confidence that comes from the eventuating experience.
The suspension leap is the one bit that only a customer can take, and it is being eroded.
And traditionally, because the customer takes it, they co-own what follows. When a purchase you took a punt on disappoints, some of the loss files under “my bad”. Brands have always been partially insulated by virtue of that co-ownership — every disappointed customer who thinks “well, I chose it” absorbs some of the blame on the brand’s behalf. Delegation removes the leap, and with it, the sense of co-ownership. In agent-led commerce, disappointment is directed outward, whole and undiluted, onto the brand. The grid with no winning square is what a market looks like once the customer no longer owns the wager.
Regulators have started drawing similar conclusions. The UK’s competition authority now advises that businesses answer for their AI agents the way they answer for employees.
What to do with a board with no winning squares
- Be deliberate about where you want the interaction to live.
Every “should we build our own agent” conversation is a liability decision you need to know how to defend. Owned agents may give you an advantage at the point of synthesis, but they carry maximum exposure. While an owned agent makes sense for a business like Walmart, with its own Sparky, rather than handing the transaction over to ChatGPT, it likely won’t make commercial sense for the majority of brands.
- Cap the charm.
A personable agent raises expectations, and the research consistently shows that warmth is punished hardest when it fails. Microsoft’s Clippy taught us this over two decades ago, but, as the bon mot goes, the lesson will be repeated until it is learned. The repayment terms on a brand’s agent leading with personality are especially brutal when things go wrong.
- Treat opportunities to make good as a premium brand asset.
How you respond when things go wrong is one of the few trust-producing moments left on the board. Air Canada was offered the cheapest make-good in aviation history and chose instead to argue for the personhood of its own chatbot. Thinking about your marketing as divorced from customer experience is no longer just an organisational silo risk.
What this means for brand strategy
Making brand-level decisions about AI deployment is not one-size-fits-all. Countless frameworks are emerging, but none give you the complete answer; they can only tell you where to look. Trust is moving, and the millions of small, private wagers customers place on brands are shifting to spaces like social and LLMs, where brands have less influence over them, yet bear more of the blame when things go wrong. That is a core mechanism agentic commerce now threatens, and one worth more than a passing thought in strategy conversations. 
Shane is a brand experience strategist and Senior Industry Fellow in the School of Economics, Finance & Marketing at RMIT. He advises brands on MarTech, brand experience and consumer behaviour. Despite decades of evidence to the contrary, he still foolishly thinks he is the one making choices about the brands he buys.