AI Tools Change Nothing Until the Work Does
Most organizations have already bought AI tools and hosted the workshops, but they're still bolting new technology onto old structures. This post …
Read article
Motion is not progress. Plenty of organisations are busy with AI: workshops, strategy decks, vendor briefings, ethics frameworks, a standing agenda item that never quite becomes a decision. It feels like momentum. Almost none of it moves a number the business cares about.
This is the central point of JD Trask’s keynote, “New Zealand’s AI export opportunity”. His argument is blunt: mistaking activity for outcomes is the main way companies waste their first year with AI. The fix is a small, funded pilot, tied to one real business result, run over roughly 30 days and owned by a person who can be held to it.
MIT’s 2025 GenAI Divide report found that 95% of the enterprise GenAI pilots it examined had produced no measurable impact on profit and loss. BCG’s 2024 AI adoption research found that 74% of companies were struggling to achieve and scale value from AI. Gartner’s June 2025 forecast predicts that more than 40% of agentic AI projects will be cancelled by the end of 2027, driven by escalating costs, unclear business value or inadequate risk controls.
None of those reports blame committees. They point to unclear value, weak data, poor controls and rising cost, and each of those describes a project that was never anchored to a specific result someone owned. That is the gap a 30-day pilot closes.
The failures rarely look like failures. They look like careful, responsible work that never ships.
“Use AI to improve customer experience” is a category, not a problem. You cannot fund it, measure it or finish it. A team handed a goal that wide spends months mapping, scoping and consulting. Progress needs a boundary: one workflow, one team, one number.
Ask who is accountable for the result of an AI initiative and you often get a group, not a name. Shared ownership across a committee means nobody feels the weight of shipping. When the outcome is a slide instead of a dollar figure, the work drifts toward whatever is easiest to present.
Controls matter. But governance can quietly become the whole programme. Drafting the policy, defining the principles and debating the guardrails all feel productive, and none of them require putting a working tool in front of a real customer. Two quarters can pass with nothing running.
Some teams stall because they are waiting: for the market to settle, for a company-wide platform decision, for the ideal use case, for someone else to prove it first. The technology keeps moving, so the wait never ends. One narrow test teaches you more than another month of watching.
This next point is a JD Trask and Autohive opinion, not a research finding, so treat it as a view rather than a fact.
Appointing a Chief AI Officer looks decisive but often does the opposite of what leaders intend. It signals that AI is a specialist function that lives with one person, separate from the operators who run the business. Responsibility for outcomes drifts to a title that sits away from the frontline, while the people who own the revenue, the margin and the customer stay a step removed.
The operating leader who owns a function should own the AI change inside it, backed by an executive sponsor who can fund the test and act on the result. AI adoption is a leadership responsibility, not a department you can hire your way out of.
Thirty days is the cadence Autohive recommends for a narrow first test, not a proven industry standard. It is short enough to force focus and long enough to produce a real signal. A good pilot has six things.
Settle these before you touch any tooling. If you cannot name the owner or the metric, you are not ready to build yet.
Start from the result you want, then work back to the mechanism, rather than starting with the technology and hunting for somewhere to apply it.
Practical first workflows tend to be repetitive, high-volume and easy to measure:
Each has an obvious owner and a cost you can already feel, which is what makes it a good candidate. Pick the one where a small improvement would be noticed by someone senior.
Judge the pilot on a business result, not on how much the tool got used.
The metrics worth measuring are the ones a CFO would recognise: revenue, margin, staff hours returned to higher-value work, response time, error rates, conversion and cost to serve. Pick the single one that matters most for your chosen workflow and hold the test to it.
Usage statistics are the weak alternative. “Logins”, “prompts run” and “queries handled” measure activity, not value. A tool can be used constantly and change nothing. Usage numbers are also easy to grow and easy to celebrate, which is why they get reported when the real result is thin. If the only number going up is engagement with the tool itself, the pilot has not proven anything.
Here is a concrete example drawn from JD’s keynote. It is illustrative, not a customer case study, and it makes no claim about time saved or results achieved.
A hotel group receives a steady stream of online reviews. Replying to each one in the right tone takes real staff time, and slow or inconsistent responses hurt the brand. A custom agent could collect the incoming reviews and draft a reply to each in the group’s approved voice. A team member then reviews every draft, edits where needed and approves it before it is published.
The agent handles the volume and the first draft. A person keeps judgement over anything that reaches a customer. The success metric would be something like response time or the share of reviews answered within a target window, not the number of drafts the agent generated.
If you build the test on Autohive, you can configure a custom agent with strict operating instructions, written around one specific task, with chosen models and explicitly assigned tools. You can ground the agent in verified workspace files so its answers come from your material rather than guesswork, and connect approved business integrations so it works where the workflow already lives.
Before putting an agent into team circulation, you can test prompts and inspect responses inside the builder against actual source documents. If the pilot handles a recurring sequence, you can link triggers and actions in a visual workflow or schedule the job with a named owner.
For the human review point, shared team chats let people inspect, edit or reject drafts using @mentions before anything is actioned. There is no automated approval gate. The control is a person staying in the loop and signing off.
That last step matters as much as the rest. Scaling the tests that work and stopping the ones that do not is how you build a track record instead of a backlog of pilots that never conclude.
A short, funded test of a single real workflow, with a named owner, a defined human review point and one success metric agreed up front. The 30 days is a focus device, not a rule: long enough to get a real signal, short enough to stop you drifting into a project with no end.
No. You need one accountable owner and an executive sponsor who can fund the test and act on the result. A committee can coordinate later, once there is evidence worth coordinating around. Standing up governance before you have run anything usually delays the first real result rather than improving it.
By a business result, not by usage. Choose one metric that matters for the workflow, such as response time, staff hours returned, conversion, error rate or cost to serve, and compare it against the baseline you recorded before the pilot began. If the only thing that moved is how much the tool got used, treat that as a warning sign.
You need one workflow, one owner and one number, run for a month. Not a grand strategy. The organisations that get somewhere with AI keep stacking up small, finished tests, ones that either worked or got shut down. A polished framework rarely makes that pile.
Turn one AI discussion currently on your calendar into a 30-day pilot with a named owner and a measurable outcome.
Most organizations have already bought AI tools and hosted the workshops, but they're still bolting new technology onto old structures. This post …
Read articleOne title can't own a general-purpose technology, so CEOs and leadership teams need to set the direction, funding and pace for AI themselves rather …
Read article