How Do You Tell a Real Leadership Intervention From a Retreat?

You cannot tell from what participants say afterward. The largest available meta-analysis of leadership training — 335 independent evaluations, 26,573 employees — shows that whether a program changes behavior back at work depends on a short list of design features, not on how powerful the experience felt. Two questions settle it before you sign: which of four outcomes the provider measures, and which design features the program actually contains.

Every experiential format eventually faces the same accusation: that it is an expensive weekend with a story attached. Ropes courses faced it. Wilderness retreats faced it. Equine-assisted programs face it now. The accusation is sometimes correct and sometimes lazy, and the difference is testable — but not by the method most buyers use, which is asking people who just came back how it went.

Why do leadership programs feel effective and change nothing?

Because satisfaction and behavior are different outcomes, and most organizations only measure the first. The standard evaluation framework, unchanged since Kirkpatrick set it out in 1959, distinguishes four levels: reactions (did they like it), learning (do they know more), transfer (are they behaving differently at work), and results (did anything organizational move). Roughly nine in ten corporate training evaluations collect reaction data. Far fewer collect anything else.

This is not a small measurement gap. Reactions and transfer are not proxies for one another. A program can produce excellent reactions because it was well-catered, emotionally vivid, and held somewhere beautiful, and produce no detectable change in how anyone runs a meeting eight weeks later. The reverse also happens: uncomfortable programs that participants rate poorly and that shift behavior anyway.

What does the evidence actually say about leadership training?

It works, and considerably better than the business press suggests. Lacerenza and colleagues, publishing in the Journal of Applied Psychology in 2017, aggregated 335 evaluations conducted between 1951 and 2014, restricted to employee samples and to designs with either pre-post measurement or a control group. The overall corrected effect was δ = .76 — a substantial effect by the standards of organizational interventions. Transfer specifically came in at δ = .82.

Two honest qualifications. The authors note their database can only contain programs someone chose to evaluate and report, which probably skews toward programs that worked; their own publication-bias tests came back clean, but that particular selection effect is not one a bias test can detect. And the newest study in the set is from 2014, which predates most of what is currently sold as leadership development.

The headline finding still stands, and it reframes the buyer’s question. The issue is not whether leadership training can work. It is which version of it does.

Why is transfer harder for senior leaders than for first-line managers?

Because at senior level, learning was never the binding constraint — transfer is. The meta-analysis split results by the trainee’s level in the hierarchy, and the pattern is stark. For first-line leaders, transfer came in at δ = 1.99. For middle managers, .54. For senior leaders, .37 — roughly four times weaker than at the bottom of the hierarchy.

What makes this finding useful is what it does not show. Learning effects did not differ significantly by level. Neither did organizational results. Senior leaders absorbed the content as well as anyone and, once they applied it, moved organizational outcomes as much as anyone. The authors read this as evidence against both the ceiling explanation (that senior leaders have nothing left to learn) and the motivation explanation (that they are not interested), pointing instead at the classic transfer problem: established behavior is hard to change on the job, and seniority means more of it is established.

That conclusion sits beside a separate question about whether a capability can be rebuilt once it has gone. The practical consequence here is uncomfortable for anyone buying development for an executive cohort.

Which design features predict that training will transfer?

Six, and all six are visible in a proposal before anyone signs anything. Ask the provider to point at each:

  1. A needs analysis conducted before the program was designed — not a catalog offering with your logo applied. Programs built from one showed stronger learning and transfer.
  2. Sessions spaced over time rather than one massed block. Spacing had no detectable effect on learning, but improved both transfer and organizational results.
  3. Practice, not only information. Programs combining information, demonstration, and practice outperformed single-method programs on transfer by more than a standard deviation.
  4. Feedback built into the design — the single clearest transfer moderator in the analysis.
  5. Delivery face-to-face rather than virtual, where transfer was the criterion.
  6. A stated evaluation level. If the provider reports reactions only, they are reporting the outcome that predicts least.

One caution on how to read that list. Several of these moderator estimates are extremely heterogeneous — the credibility interval around the needs-analysis effect on transfer crosses zero despite a large average. The direction of each feature is trustworthy; the magnitudes are not, and any provider quoting you a precise multiplier from this literature is overselling it.

Two regional notes. In the Gulf, the destination program is close to a market default, which collides directly with feature 2: a five-day off-site block is massed by definition, whatever else it is. In Central and Eastern Europe, cost discipline pushes the opposite way, toward virtual delivery — and virtual delivery showed transfer of δ = .22 against 1.10 face-to-face, the largest single gap in the moderator set.

How does this test apply to experiential formats like equine-assisted programs?

Applied honestly, the format passes on design and fails on evidence quality — which is a more useful verdict than either the industry or its critics offer. On the six features, structured equine-assisted leadership work scores well on three: it is practice-based rather than informational, it is face-to-face by necessity, and feedback is immediate and difficult to argue with, since the animal responds to what the participant is actually doing rather than to what they intended. It typically fails feature 2, because it is usually sold as a compressed multi-day block.

The evidence base is the weaker half. The most recent peer-reviewed study, published in Frontiers in Veterinary Science in November 2025, followed eight leaders for twelve months after a five-day program and reported durable changes in communication and emotional regulation. It is a qualitative study: eight participants purposively selected from fifty, self-reported in retrospective interviews, from one sector in one country, with several of the researchers affiliated with the institution delivering the program. The authors state plainly that they cannot isolate the program’s effect from ordinary leadership maturation. That is a creditable piece of work and it is not evidence of effectiveness in the sense feature 6 requires.

I have written on leadership learning in equine settings before, and my own view has not changed: the mechanism is real and the measurement is thin. The physiological account of what happens during that contact is well-established in the underlying literatures. What remains unestablished is whether it converts into behavior at work more reliably than a well-designed conventional program — and on the current evidence, nobody is entitled to claim it does.

Frequently asked questions

Are longer leadership programs more effective?

Longer duration was associated with better organizational results, and the relationship was linear rather than curved once outliers were removed. It showed no relationship with learning or transfer. Duration buys you results-level outcomes, not behavior change.

Is 360-degree feedback worth the cost?

On this evidence, it is not clearly better than feedback from a single source. Programs using 360-degree feedback showed no statistically significant advantage over single-source feedback on learning, transfer, or results — despite roughly 90% of large organizations using it. The number of feedback sources appears to matter less than whether feedback exists at all.

Is off-site training always worse than on-site?

No. On-site programs outperformed off-site ones on organizational results, but the two were statistically indistinguishable on learning and on transfer. The argument for on-site is about downstream organizational outcomes and cost, not about whether participants change their behavior.

Sources

  • Lacerenza, C. N., Reyes, D. L., Marlow, S. L., Joseph, D. L., & Salas, E. (2017). Leadership training design, delivery, and implementation: A meta-analysis. Journal of Applied Psychology, 102(12), 1686–1718. Full text read and verified.
  • Sivagurunathan, R., Senathirajah, A. R. bin S., Sivagurunathan, L., Arokiasamy, L., Qazi, S., Haque, R., & Su, Y. (2025). Equine-assisted learning and leadership transformation: an exploratory qualitative study of workplace behavior. Frontiers in Veterinary Science, 12, 1700029. Open access, full text read and verified.
  • Rajfura, T., & Karaszewski, R. (2018). Horse sense leadership: what can leaders learn from horses? Journal of Corporate Responsibility and Leadership, 5, 61–83.
  • Kirkpatrick, D. L. (1959) — four-level framework, cited via Lacerenza et al.
  • Baldwin, T. T., & Ford, J. K. (1988) — the transfer problem, cited via Lacerenza et al.