In 2021, I wrote about measuring the impact of executive coaching. The financial formula has not changed:
ROI = (Benefits achieved - coaching costs) / coaching costs x 100
The difficult part was never the formula. It was deciding which benefits to count and how much of the change to attribute to coaching.
Suppose sales increase after a leader completes a coaching engagement. How much credit belongs to better leadership, and how much to pricing, a new product, market conditions, or the sales team? We can assign coaching a percentage of the gain, but the number is still a judgment. When the coach, coaching company, or program sponsor selects both the benefit and the attribution percentage, the result can become self-serving very quickly.
That helps explain why executive coaching has accumulated eye-catching ROI claims that are difficult to audit. Controlled research gives us good reason to believe coaching improves performance and behavior. A 2023 meta-analysis of 37 randomized controlled studies found a meaningful overall effect. It did not produce a universal five-times, seven-times, or ten-times financial return.
AI does not eliminate this attribution problem. But it may help us build a much better evidence trail.
The problem with measuring what is easy
Traditional coaching ROI models often rely on retention, productivity, engagement, revenue, etc. These are useful when they relate directly to the reason coaching was commissioned. They are much less useful when selected simply because they are available and can be converted into dollars.
Retention is a good example. If an organization is losing critical people because a leader creates a punishing team environment, improved retention may be central to the business case. If retention is already strong and the leader was coached to improve strategic judgment, using turnover savings to justify the engagement tells us little.
Several tech-driven coaching platforms now offer ROI calculators and real-time dashboards. Their dashboards can track participation, assessment results, goal progress, and cohort trends. Some connect coaching data with performance ratings, promotions, productivity, or turnover, and compare participants with nonparticipants. For broad programs involving hundreds or thousands of employees, standardized measures make sense. They provide consistency and allow HR leaders to see patterns at scale.
But standardized measurement also has limits. Platform models commonly begin with outcomes that can be calculated consistently across a large population. The organization's specific reason for coaching can become secondary to the measures the system already knows how to process.
Selection bias is another complication. Volunteers may be more motivated or development-oriented. If they outperform nonparticipants, comparison groups help only when we understand who entered each group and why.
Why senior-level coaching requires a different lens
At senior levels, the sample is small and the mandate is highly specific. One executive may need to lead an integration, another to restore confidence after a failed transformation, and another to build a credible successor. A universal scorecard will miss much of what matters.
The potential leverage is also different. A senior executive's decisions, relationships, and operating habits affect teams, investments, customers, and strategic choices far beyond the individual. That does not guarantee that executive coaching produces a higher percentage return than coaching elsewhere. It does mean that a meaningful change in one pivotal role can create substantial organizational value.
This is why coaching investment should be concentrated where the leadership challenge and potential organizational consequence are greatest, and why the measurement should be designed around that leader's mandate. For senior executives, the goal is not more data. It is continuous measurement of the most relevant data.
What AI can change: from snapshots to telemetry
Telemetry is the stream of digital evidence generated by the systems in which work already happens: meeting platforms, customer-service systems, CRM tools, project-management software, calendars, collaboration networks, and short employee surveys.
AI can classify, summarize, and connect this information at a scale that would be impractical to manage manually. Instead of asking stakeholders six months later whether a leader became more effective, an organization may be able to observe agreed signals repeatedly during the engagement.
The key is to establish the measurement chain before coaching begins:
Coaching goal -> Signal AI can track continuously -> Business value
Here are practical examples, each tied to a common executive coaching goal:
| Coaching goal | Signal AI can track continuously | Business value |
|---|---|---|
| Delegate more, micromanage less | Decisions escalated to the leader; the leader's meeting load and after-hours work | Executive time freed for strategy; faster decisions |
| Lead better meetings | Share of meetings ending with a clear decision and owner, from AI meeting summaries; team meeting hours | Hours recovered across the team; shorter cycle times |
| Develop the team | 1:1 time with direct reports; internal promotions; regretted attrition | Lower replacement costs; a stronger succession bench |
| Lead change clearly | Recurring themes in pulse-survey comments and town-hall questions; adoption of the new process | Faster adoption; less rework |
| Sharpen customer focus | Volume and sentiment of escalated customer issues in the leader's area; time to resolve | Retained accounts; protected revenue |
Continuous does not mean causal
More data can produce more confidence without producing more truth. AI can spot a pattern; it cannot tell us what would have happened without coaching.
Executives work in complex systems. Strategy changes, incentives, reorganizations, new hires, market shifts, and other development efforts may all influence the same outcome. A meeting transcript is also a proxy, not an objective account of leadership quality. Automated classifications can be wrong, culturally biased, or stripped of context.
For that reason, AI-assisted coaching measurement should follow several rules:
- Define the business reason and two or three target behaviors first. Do not collect data and hunt afterward for a favorable story.
- Establish a baseline. Use several observations before the engagement when possible, rather than a single pre-coaching score.
- Triangulate. Combine a behavioral signal, stakeholder feedback, and an operating outcome. No one measure should carry the argument.
- Test other explanations. Use a comparison team, phased rollout, or historical trend where feasible, and record material changes in the business environment.
- Report contribution, not invented causation. State the evidence, the competing explanations, and the confidence level. Use a range rather than one heroic percentage.
- Protect trust. Collect only what was agreed, obtain informed consent, separate coaching conversations from employer reporting, limit access and retention, and keep development data from quietly becoming surveillance or performance evidence.
Only after that work should an organization convert a benefit to dollars and apply the familiar ROI formula. In some engagements, a financial calculation will be credible. In others, the most honest result will be strong evidence of behavior change and business contribution without a precise dollar return.
A better promise for coaching ROI
AI will not deliver the holy grail of perfectly isolating coaching from every other influence on an executive’s performance. Nor should we hand an algorithm the authority to decide what good leadership means.
What AI can do is reduce the burden and delay of gathering relevant evidence. It can help organizations move from an annual snapshot to a carefully governed stream of signals, reveal patterns sooner, and identify when expected change is not occurring.
That is a meaningful advance. The future of executive coaching ROI is not a universal calculator or a larger claim. It is the continuous collection of the most relevant evidence, selected because it reflects what the organization actually needed from the coaching in the first place.
Measure fewer things. Measure them earlier and more often. Most important, measure what matters.
Sources
- Erik de Haan and Viktor O. Nilsson, “What Can We Know about the Effectiveness of Coaching? A Meta-Analysis Based Only on Randomized Controlled Trials,” Academy of Management Learning & Education, 2023.
- Paul Lawrence and Alana Whyte, “Return on Investment in Executive Coaching: A Practical Model for Measuring ROI in Organisations,” Coaching: An International Journal of Theory, Research and Practice, 2014.
- National Institute of Standards and Technology, AI Risk Management Framework 1.0, 2023.
- Current public measurement materials from EZRA, CoachHub, BetterUp, and Bravely, reviewed October 2026.
Use of artificial intelligence
The author used ChatGPT for initial brainstorming, developing the article outline and providing editorial support including grammar review and style consistency. All conceptual frameworks, research synthesis, and practical recommendations represent the author’s original work and professional expertise. The author reviewed and edited the content produced by the artificial intelligence tool and is fully responsible for the information contained in the article.