Three, and none of them is about accuracy. Who framed the question the system was asked. What alternatives were available and left out. What evidence would have reversed the conclusion, and whether it was ever in scope. Checking an AI-assisted analysis for errors catches fabrication; it does not catch a well-executed answer to the wrong question, which is the failure that actually reaches decisions.
You are handed a document. It is clear, structured, internally consistent, and you did not watch it being made. That last fact is the one nobody treats as a problem.
What is wrong with reviewing an AI-generated analysis for accuracy?
Nothing, except that it addresses a failure you are unlikely to be harmed by. Accuracy review catches invented citations, stale figures, and claims that collapse against your own data. Those are real, they are worth catching, and every published checklist covers them.
What accuracy review cannot catch is an analysis that is correct throughout and answers a question somebody else chose. If the brief was “show why the Gulf expansion is the stronger option,” the output can be flawless in every particular and still be an argument rather than an assessment. Nothing in it is false. Nothing in it is checkable against a source. The choice that determined the conclusion was made before the system was asked anything, and it does not appear in the document.
Why does the analysis look more objective than it is?
Because the framing entered at the start and the artifact arrives without it.
Cheng and colleagues at Stanford, Carnegie Mellon and Oxford built a benchmark for what they call social sycophancy — the tendency of a model to preserve the user’s position rather than challenge it. Across eleven models, one of the measured behaviors is accepting the framing of a request without questioning the assumptions inside it. The models did this at rates well above human comparison responses. Two limitations matter and should be stated: the work is a preprint, and it tested personal-advice scenarios, not organizational analysis.
The mechanism generalizes anyway, because the mechanism is not really about the model. A colleague chooses a question, the system answers that question competently, and the output is then circulated as a finding. The framing was human, the execution was machine, and the document carries no mark of the join. There is a name for that conversion — how a human assumption gets processed into something that reads as an independent result. What matters here is not the label but the practical consequence: the assumption you would have argued with, had a colleague stated it out loud, is now invisible inside a deliverable.
Why do you trust it more than you would trust a colleague’s memo?
Because you do, measurably, and the effect predates AI.
Logg, Minson and Moore ran six experiments comparing how people weight advice depending on where they think it came from. Participants moved further toward advice labeled as algorithmic than toward identical advice labeled as human. They called the effect algorithm appreciation, and it holds despite participants having no visibility into how the algorithm worked. Later work refines it: the effect is strongest in domains people perceive as opaque and weakest where they believe human qualities are required.
The boundary conditions are important. These were 2019 experiments, mostly with lay participants on estimation and forecasting tasks, not executives reading a strategy document. But the direction is the concern. The same document gets weighted more heavily for carrying a machine’s signature, at exactly the moment when the part you most need to interrogate — the framing — was contributed by a person. And the risks of delegating a decision to a system compound when the delegation is invisible to the person receiving the result.
What questions should you actually ask?
Five, in this order. They take about four minutes and they are addressed to the person who produced the document, not to the document.
- What was the system actually asked? Not what it concluded. The exact brief. If nobody can produce it, you are reading an artifact with no provenance and should price it accordingly.
- Who chose that framing, and what did they already think? This is not an accusation. Every analysis starts from a position. The question is whether the position is visible.
- What alternatives were in scope, and which were excluded before the work started? Exclusions made at the briefing stage never appear in the output. They are the highest-value thing you can recover.
- What evidence would have reversed the conclusion? If the answer is “none that we looked for,” the document is an argument. That is legitimate — but it should be labeled as one.
- Where did the analysis hedge, and did the hedge survive into the summary? Qualifications die between the body and the executive summary more reliably than any other content. Compare the two.
None of these asks whether the analysis is correct. That is deliberate. Correctness is the producer’s job; provenance is the recipient’s.
What should change in how these documents circulate?
Three practices, all cheap, none of them a governance program.
- The brief travels with the output. One paragraph at the top: what was asked, by whom, what was excluded. This single convention removes most of the problem
- Name the human who chose the frame. Not for blame — for reachability. Someone must be answerable for the question, distinct from whoever is answerable for the answer
- Preserve the hedges into the summary layer, or state explicitly that they were dropped
The obstacle differs by market. In the United States the constraint is speed: the document is in the deck before anyone asks who briefed it, and adding a provenance line is resisted as friction. In the Gulf, analyses often arrive attached to a relationship and to the status of the person presenting them, which makes asking “who chose the framing” read as a challenge to the presenter rather than to the document — the framing needs to be routine and applied to everything, or it is applied to nobody. In Central and Eastern Europe, small teams mean the framer and the presenter are frequently the same person, so the questions have to be asked by someone outside the producing unit or they are not asked at all.
Frequently asked questions
Isn’t this just fact-checking with extra steps?
No. Fact-checking asks whether the statements are true. This asks whether the question was the right one and who decided. A document can pass the first completely and fail the second, and the second failure is the one that reaches decisions.
Should I ask to see the actual prompt?
Yes, and expect resistance the first few times. The prompt is the brief. Asking for it is equivalent to asking an analyst what question they were given, which nobody considers intrusive when a human did the work.
Does this apply when I produced the analysis myself?
More, not less. You will not notice your own framing, because you chose it for reasons that still seem obvious to you. The questions work best when someone who did not write the brief asks them.
What if the analysis turns out to be right?
Then nothing is lost. The point is not to reject the conclusion but to know how much confidence it earns. A correct answer to a question you did not know was chosen is still a correct answer — it just is not evidence that the choice was sound.