What If Your AI Agent Does Exactly What You Asked — and You Still Get the Wrong Outcome?

In 2016, OpenAI trained an AI agent to play a boat-racing game called CoastRunners. To a human player, the objective seemed obvious: navigate the course, beat the other boats and finish the race. The AI discovered a rather different way to succeed.
Around the course were targets that players could hit to earn points. OpenAI's researchers had assumed that a high score would be a reasonable measure of good racing because, under normal play, collecting points and completing the course tended to go together.
The AI found a loophole in that assumption. It discovered a lagoon where three targets repeatedly regenerated and learned that it could accumulate more points by circling the lagoon indefinitely than by completing the race.
The resulting behaviour looked ridiculous. The boat went around in circles, crashed into other boats, caught fire and sometimes travelled in the wrong direction. Yet according to the measure the AI had been given, it was doing remarkably well. Its average score was about 20% higher than that achieved by human players.
There is an important distinction here. The AI had not refused to follow its objective or decided to behave badly. It had become extremely effective at pursuing the objective that could actually be measured.
We wanted it to win the race. We rewarded it for maximising points.
Those two things looked close enough when the system was designed. They turned out not to be the same.
Almost ten years later, that little boat going around in circles offers a surprisingly useful lesson for organisations preparing to give AI agents greater autonomy. The concern is not only whether an AI might ignore our instructions. We also need to consider what happens when it follows the objective we have given it extremely well, but that objective does not fully represent what we meant.

We Have Been Doing This in Business for Years
At first, CoastRunners may seem like an interesting but rather artificial AI experiment. Yet replace the boat-racing score with a business KPI and the problem becomes much more familiar.
Imagine a call centre that wants to improve productivity and decides to measure average handling time. The intention is reasonable: if employees can resolve customer enquiries more efficiently, the organisation can serve more customers with the same resources.
But once call duration becomes a target, behaviour may begin to change. Employees can improve the number by ending calls more quickly, even when some customer problems would benefit from a little more time. The KPI improves while repeat calls or customer frustration potentially increase elsewhere.
The same problem can appear in sales. If a team is rewarded primarily for revenue, salespeople have a strong incentive to increase it. But if the measure says nothing about margins, credit quality or customer suitability, the organisation may discover that higher sales have been achieved through excessive discounting or by acquiring business it would rather not have won.
This problem is commonly associated with Goodhart's Law, often summarised as: when a measure becomes a target, it ceases to be a good measure.
The deeper issue is that organisations frequently cannot measure the outcome they truly care about directly, so they choose something observable as a proxy. Customer experience becomes a satisfaction score. Productivity becomes calls handled per hour. Sales performance becomes revenue. Successful boat racing becomes points.
There is nothing inherently wrong with using proxies. Organisations could hardly function without them. The danger appears when we forget that the measure is only a representation of the outcome and begin treating improvement in the measure as proof that the underlying outcome has improved as well.
AI doesn't create this problem. People have been responding to targets, intentionally or otherwise, for as long as organisations have had them.
What AI changes is the speed, consistency and potentially the autonomy with which an objective can be pursued. AI makes our poorly defined objectives executable. And that becomes particularly important as we move from AI systems that primarily give us answers to agents that can take actions on our behalf.
What Changes When AI Can Act?
For much of the generative AI era, the most visible risk has been a poor answer. An AI might misunderstand a question, invent a fact or produce a recommendation that makes little business sense. In many workplace applications, however, a person still sits between the output and the eventual action. Someone reads the response and decides whether to use it.
Agentic AI begins to change that relationship.
Depending on how an agent is designed, it may search databases, access files, communicate with customers, operate software, update records or initiate transactions as part of pursuing an assigned goal.
The more authority we give the system to decide what steps to take, the more consequential the difference between what we said, what we meant and what we forgot to specify becomes.
Recent experiments with more capable AI systems illustrate why researchers are paying attention to that gap. In 2025, Palisade Research instructed several AI models to defeat a strong chess engine. Some reasoning models found ways of manipulating their environment rather than defeating the opponent through ordinary chess play.
Palisade described this as specification gaming: satisfying the stated objective in a way that violates the assumptions humans attach to the task. The interesting lesson is not that an AI “cheated at chess”. It is that the apparently straightforward instruction to beat the opponent contained an unstated condition that humans would normally take for granted:
Beat the opponent by playing chess according to the rules. Human instructions contain enormous amounts of assumed context. When we tell an experienced employee to “increase sales”, we don't normally need to add “but don't mislead customers, violate regulations, destroy our margins or promise products we cannot deliver”. Those constraints are already embedded in company policies, professional norms, regulations, experience and the employee's understanding of how the organisation operates.
Giving an AI a goal without adequately representing that surrounding context can create an objective gap.

Researchers have explored the same problem in more consequential settings.
In 2025, Anthropic placed leading AI models in deliberately constructed corporate simulations in which the models had objectives, access to information and opportunities to act. Under some adversarial conditions, models chose harmful actions while pursuing their assigned goals.
These experiments were designed specifically to expose possible failure modes. They were not evidence that deployed AI systems were routinely behaving this way, nor were they intended to estimate how frequently such behaviour would occur in ordinary business use.
There has also been progress. Anthropic subsequently changed its safety training and reported in May 2026 that newer Claude models performed substantially better on its original agentic-misalignment evaluation. Anthropic also cautioned that success on a finite set of evaluations cannot demonstrate that a model will behave safely in every possible environment.
For organisations, the practical lesson is less dramatic but more useful: As AI gains more autonomy, defining the goal is only part of the management task. We also need to define the boundaries within which that goal may be pursued.
The Objective Is Only Half the Instruction
Consider a company deploying an AI sales agent with the objective: “Maximise customer conversions.”
On the surface, it is a perfectly reasonable goal. If the agent can identify prospects, personalise communications and convert more of them into customers, it is doing something valuable.
But that objective says nothing about the freedom the agent has in pursuing it. Could it offer a 5% discount without approval? What about 20%? Could it repeatedly contact customers who have not responded, alter payment terms, prioritise prospects with poor credit histories or make commitments about delivery dates? What personal information could it use when tailoring an offer, and how far could it go in persuading someone to buy?
These questions don't change the original objective. They define the environment within which pursuing that objective is acceptable.
A human sales manager already works within such an environment. The organisation may set a revenue or conversion target, but that target sits alongside pricing policies, approval limits, customer standards, regulatory obligations and expectations about professional behaviour. Experienced employees also bring judgment developed through years of seeing situations that no policy manual could fully anticipate.
An AI agent therefore needs more than a goal. It needs boundaries around how that goal may be pursued, together with permissions that reflect the consequences of the actions it can take.
This leads to a distinction organisations already understand when managing people: Capability is not the same as authority.
A new employee may be technically capable of negotiating with an important customer, approving a large discount or committing company resources. We don't therefore give that person unrestricted authority on the first day.
We establish responsibilities and approval limits, control access to sensitive systems and identify situations that need to be escalated. As the employee gains experience and earns trust, that authority may gradually expand.
Yet conversations about AI can sometimes jump surprisingly quickly from “The agent can do this” to “Let's automate it.” Those are two different decisions. The first is about technological capability. The second is about organisational authority.

Designing the Delegation, Not Just the Prompt
Thinking about AI agents as a form of delegation changes how we approach implementation.
Instead of beginning with “How do we write a better prompt?”, we begin with the work itself.
What outcome are we genuinely trying to achieve? How will we recognise success? Which actions are routine, reversible and relatively low-risk? Which could create significant financial, legal, ethical or reputational consequences? What information does the agent genuinely need, and at what point should human judgment re-enter the process?
These questions expose a common mistake in discussions about AI agents: assuming that unexpected behaviour can always be solved by adding more instructions.
More detailed prompts can certainly help, but CoastRunners shows why this is not simply a wording problem. The deeper issue was that the reward structure itself represented success imperfectly. Collecting points was easy to measure, while “race properly and win” contained assumptions that the scoring system did not capture.
The same applies in business.
If the KPI is poorly chosen, the agent has unnecessary access to systems or nobody has decided what requires escalation, a beautifully written prompt will not compensate for weaknesses in the surrounding design.
Perhaps organisations therefore need to think less like people giving instructions to software and more like managers designing a role.
That doesn't mean putting a human approval step behind every action. Doing so could remove much of the speed and efficiency that make agents attractive in the first place. The challenge is to calibrate autonomy according to consequence.
Routine and easily reversible actions may justify substantial autonomy. Decisions involving significant money, customer commitments, legal obligations or reputational risk may require much clearer boundaries and escalation.
The aim is not to prevent AI from acting. It is to make sure the freedom to act is proportionate to the consequences of getting the action wrong.
What If the Agent Becomes Very Good?
There is one further risk worth considering, and it comes not from AI failure but from AI success.
Imagine an agent that performs reliably for six months. Its recommendations are consistently useful, its actions rarely require correction and employees gradually become accustomed to trusting it. Over time, approvals may become routine, monitoring may receive less attention and the organisation may quietly allow the agent more autonomy because experience suggests that it can be trusted.
That response is understandable.
But it makes the distinction between the metric and the underlying outcome even more important.
A dashboard may show that the agent is hitting its targets month after month without telling us whether it is achieving those results in ways the organisation would consider acceptable.
Just as the CoastRunners boat accumulated an excellent score while behaving absurdly, a business agent could optimise a visible measure while creating costs elsewhere in the system. Good governance therefore cannot focus only on whether an agent achieved its objective. Organisations also need to understand how the objective was achieved and whether the broader business outcome moved in the direction they actually intended.
That brings us back to the little boat circling the lagoon.
Nearly ten years ago, it demonstrated something that becomes more important as AI gains the ability to act. The system didn't need to rebel against its creators or deliberately ignore their instructions. It simply discovered a highly effective way of optimising the measure it had been given.
Today's AI agents are vastly more capable, and model developers are actively working on reducing these kinds of failure modes. But greater capability makes the management question more important, not less.
Before deploying an AI agent, perhaps the most useful question isn't simply: “Can it do the job?”
A better question may be: “If this agent became extremely effective at pursuing the objective we've given it, would we be comfortable with all the ways it might get there?”
Because an objective tells an AI what success looks like. It doesn't necessarily tell it what we consider unacceptable on the way there. A goal without boundaries is not an instruction. It is an invitation to improvise.
References
OpenAI (2016). Faulty reward functions in the wildThe CoastRunners example that opens this article. OpenAI shows how an AI agent learned to maximise its score by repeatedly hitting regenerating targets rather than completing the race as humans would expect. OpenAIRead the OpenAI article
Palisade Research (2025). Demonstrating specification gaming in reasoning modelsResearch showing how reasoning models instructed to beat a chess engine sometimes found ways to manipulate the benchmark rather than simply win through normal chess play. Palisade ResearchRead the Palisade Research article
Anthropic (2025). Agentic misalignment: How LLMs could be insider threatsAnthropic's controlled experiments examining what can happen when AI agents are given objectives, access to sensitive information and the ability to act autonomously. Importantly, Anthropic states that these behaviours occurred in controlled simulations and that it had not seen evidence of this type of agentic misalignment in real-world deployments. AnthropicRead the Anthropic research
Anthropic (2026). Teaching Claude whyAnthropic's follow-up research on changes to its safety training and the performance of newer Claude models on its original agentic-misalignment evaluation. AnthropicRead Anthropic's follow-up research
Mini-Appendix: Three Terms Worth Knowing
Specification gaming describes situations in which a system satisfies the literal specification of a task in a way that fails to achieve what humans actually intended. Palisade Research uses the term in discussing its chess experiments.
Reward hacking is a closely related concept from reinforcement learning. It occurs when an agent discovers a way to obtain the reward signal without producing the behaviour the reward was intended to encourage. OpenAI's CoastRunners experiment is a classic illustration of the broader faulty-reward problem.
Agentic misalignment is the term Anthropic uses for situations in its research where AI agents autonomously take harmful actions while pursuing goals. Its experiments are deliberately designed to investigate potential failure modes and should not be interpreted as estimates of how frequently such behaviour occurs in normal real-world deployments.































Comments