Companies rushing to integrate artificial intelligence into their operations are discovering that measuring success is far more complicated than deploying the technology. Many organizations remain in what observers call the administrative phase, using AI for modest automations such as drafting documents, answering emails, or conducting research. But as enterprises push for more strategic applications, a critical question is emerging: how does the AI know whether what it did actually improved the business?
The challenge lies in the difference between outputs and outcomes. Current AI metrics often focus on answer quality, task completion, latency, and cost. Yet businesses ultimately care about customer churn, margin, conversion, claims accuracy, delivery times, and customer lifetime value. An agent could complete its assigned task perfectly and still harm the company. A common example is a discount offered to secure a renewal. The customer renews, appearing as a success, but the margin collapses, the customer learns to wait for discounts, and similar customers may demand the same treatment. What counts as success depends entirely on the scoreboard.
This dynamic is captured by Goodhart's law: when a measure becomes a target, it ceases to be a good measure. Poorly chosen metrics become dangerous when optimized aggressively. In business, optimization rarely means just one thing. Improving one metric can easily damage another, whether in sales, manufacturing, or customer service. As a result, AI agents need to operate within a framework of objectives and feedback, much like traditional management. OpenAI now frames enterprise evaluation as a cycle of specify, measure, and improve, turning vague business goals into actionable and measurable expectations.
The industry is also recognizing that evaluation must be continuous. AI cannot simply be approved once and then forgotten. Evaluation has to become an active part of operations, with full workflows captured, behavior graded, regressions detected, and prompts, routing, and guardrails improved. Pre-deployment benchmarks are useful, but they happen before systems are deployed, and companies operate in changing environments. According to the National Institute of Standards and Technology, controlled pre-deployment testing cannot capture every unexpected behavior or consequence that may emerge in real-world use. The useful question is whether the system worked in the business, not merely in the test.
Recent work by McKinsey suggests that many organizations are scaling agents faster than they can redesign the work beneath them, and boards and CFOs are starting to demand clearer value. Token costs are only one piece of the puzzle; agentic workflows must be judged economically at the workflow level. Task completion is not the same as business success. The managerial translation is straightforward: companies must move from asking whether an agent succeeded to asking whether the company improved.
The precedent of AlphaZero offers an intuitive parallel. The algorithm developed by DeepMind did not become strong because it could describe chess. It improved because it acted, saw the consequences of its actions, and had an unambiguous objective: to win. It learned through repeated trial and error, favoring moves that increased its chances of winning. Companies have much messier goals, but the principle survives: learning requires feedback tied to an objective. The hard part is choosing the scoreboard. Increasing sales is not enough. What about margin, churn, returns, compliance, or brand image? Multiple objectives can and do conflict, and those conflicts are not engineering problems alone. They are management problems.
Boards and executives eventually have to decide what should improve, what must never be sacrificed, and what trade-offs are acceptable. Once systems can continually optimize, the objective itself becomes a governance decision. Steering only makes sense if there is a clear destination and a way to measure progress toward it. The missing control function must learn from repeated enterprise episodes toward an explicit objective. As companies move from intelligence to steering, setting objectives, and learning, the questions CEOs should ask are simple: What exactly is our AI optimizing? And how will we know whether its actions improved the business?