Insights · AI Delivery

Where AI agents do not work yet.

Agents genuinely run work inside Australian organisations. They also fail in four recurring places. Here is the line, and how to tell which side your use case sits on.

The useful question about AI agents is no longer whether they work. In Australian organisations they demonstrably do: agents triage inbound requests, reconcile records across systems, draft and route documents, and close multi-step tasks that used to sit in someone's queue for three days. The useful question is narrower and much more practical: where do they stop?

That question gets asked far less than it should, because most of what is published about agents is written by people selling them. So here is the other half, drawn from the engagements where we have had to say no, or had to stop and rebuild something underneath before the agent could work at all.

Four places agents reliably fail.

These are not exotic edge cases. They are the four we meet most often, and three of the four have nothing to do with the model.

  • Where the data is contested. An agent that has to decide which of three systems holds the true customer address will pick one, confidently, and be wrong a predictable fraction of the time. Humans handle this by knowing which system to trust for which field. That knowledge usually exists nowhere but in people's heads, so the agent cannot inherit it.
  • Where the task is accountable rather than repetitive. If a regulator, a board or a customer can ask "who decided this, and on what basis?", the work needs a named human in the decision, not merely in an approval queue afterwards. Agents are excellent at assembling the evidence for that decision. They are a poor place to put the decision itself.
  • Where the process is undocumented and genuinely variable. Agents encode a process. If the real process is forty exceptions wearing a trench coat, encoding it produces a brittle system that fails loudly on day nine. The honest sequence is to fix the process first and automate second, which is slower and much less fun to demo.
  • Where the tool surface will not hold. Many Australian enterprises run systems whose integration surface is a nightly file drop or a screen. An agent can be made to work against that, but it inherits every fragility of the connection, and the running cost of maintaining it usually exceeds the labour saved.

The pattern underneath all four.

Three of those four failures are data and process problems that the agent merely exposes. This is the single most useful thing to understand before commissioning agentic work: an agent is an amplifier. Point it at a clean, well-governed, well-integrated process and it multiplies throughput. Point it at an ambiguous one and it multiplies the ambiguity, faster, and with a confident tone of voice.

It follows that the cost of an agent programme is rarely in the agents. It is in the governed data foundation underneath them — knowing which system is authoritative for which field, what the lineage is, who is allowed to see what. Firms that only build agents tend to discover this in month three, as a variation.

A test you can run this week.

Before scoping any agent, put the candidate task through four questions. They take an afternoon and they are a better predictor of success than any pilot.

  • Can you name, for every field the task touches, the one system that is authoritative? If not, you have a data problem wearing an AI costume.
  • If this goes wrong 1 in 50 times, who finds out, and how fast? An undetectable failure mode disqualifies the task regardless of the error rate.
  • Is the process written down at a level someone outside the team could follow? If not, write it down first — you will frequently find the automation is no longer the interesting part.
  • Does every system involved have an interface that a vendor supports and will keep supporting? Screen-level integration is a liability you are choosing to take on.

A task that passes all four is a strong agent candidate. A task that fails one is usually worth fixing first. A task that fails three is a business-process piece of work that will not be improved by adding a model to it.

Why the firm you pick matters more than the model you pick.

Models are close to a commodity, and they change under you every few months. What does not change is whether the people doing the work can tell a data problem from an AI problem on the first visit, and say so when the honest answer is "not this one, not yet."

RUBIX has been doing data and analytics work in Australia since 2011 — fifteen years, across banks, super funds, insurers and government, in an operating and regulatory environment that does not resemble the one most global AI vendors design for. That is the specific reason we can build both halves: the data platform and the agents that stand on it, with one team accountable for whether the thing actually works in production rather than in a demonstration. It is also why we are comfortable telling clients which of their four candidate use cases to drop.

If you are weighing agentic work now, the most valuable first step is usually not a pilot. It is an AI readiness assessment that tells you which of your candidate tasks sit on the working side of the line, and what would have to be true for the rest to join them. That, and a frank conversation with an Australian AI consulting team that will tell you when the answer is no.

TL;DR: AI agents work well where data is authoritative, the process is documented, failure is detectable and the integrations are supported. They fail where the data is contested, the decision is accountable, the process is genuinely variable, or the tool surface is fragile. Three of those four are data and process problems that the agent only exposes — which is why agentic programmes are won or lost underneath the agents.

General information only, not legal, regulatory or financial advice. Current as at September 2026.