The argument
I finish the research before I write the first prompt
Redha Alayesh is a marketing consultant in Riyadh, Saudi Arabia. Two numbers in his own build documents for this site were wrong, and reading them again would never have caught either one. This is the argument for the order.
You cannot review what you have no standard for. Two numbers in my own build documents for this website were wrong, and I had read that document more than once. Reading it again was never going to catch them, because reviewing is a comparison, and I had nothing to compare against.
Most professional work now starts in the same place. You describe the task, something complete and well organised comes back in about four minutes, and then you decide what to do with it. The careless publish it. The careful read it, ask for two or three changes, and publish that. The careful version fails for the same reason the careless one does. It just takes longer to find out.
The first of the two said that lists and tables extract 43% more accurately than the same content written as prose. It has no source. It sits in the methods section of one preprint (arXiv:2603.29979), attributed to "our experiments", and no such experiment is reported anywhere in that paper. Extraction accuracy is never defined and is not one of the paper's metrics. The claim printed beside it carries a citation. This one carries none.
The second was a real finding restated at the wrong size. A citation rate that moved from 45.0% to 52.8% went into my document as "17.3%". That is the relative lift, and almost every reader takes a figure written that way as percentage points. The actual gap is 7.8 points.
Both were in the file as settled fact. One of them carried a p-value. They came out on 19 August 2026, and what removed them was not another reading. It was opening the paper.
Why does reading it again not catch the problem?
Reviewing a draft is a comparison, and a comparison needs a second term. If you have not gone and got that second term in advance, the only thing left to compare against is your own sense of whether the text reads well. Reading well is the single property these systems are guaranteed to deliver, because it is the property they were built for. So the signal you are actually reading carries no information about whether the work is right.
This has been measured, and on the right people. Shakked Noy and Whitney Zhang ran a randomised experiment on 453 college-educated professionals, published in Science: marketers, consultants, grant writers, data analysts, HR staff, managers. 33% of the group given ChatGPT reported submitting its first output with no edits at all.
Another 53% said they had edited it, and were active for an average of 3.3 minutes after pasting the text in, most of them for between zero and two. How long they spent editing had no relationship to the grade they got. And the people who used ChatGPT did not score higher than the model's raw output, handed to the same evaluators.
The careful group and the raw output graded the same.
Here is what that looks like with a number in it.
The most repeated statistic in Arabic web typography says that properly typeset Arabic pages increased average session duration by 38% and conversion by 22%, according to a 2025 Smashing Magazine report. There is no such report.
I traced it this week. The page carrying the claim links to Smashing Magazine's typography category as its source. The most recent article in that category is dated 22 August 2024. Smashing Magazine's articles about Arabic are dated between 2010 and 2014, and they are about type design and calligraphy. Not one of them contains a session figure or a conversion figure.
Now read the fabricated sentence again, as a reviewer would. It has a source, a year, a publisher with a real reputation, and two figures that are neither large enough to look suspicious nor small enough to look pointless. Nothing about it looks wrong. It is not wrong to look at. It is wrong to check, and those are two different operations. The check took four minutes.
When the research behind this site's writing system was built, 269 claims were pulled out of the dossiers and sent to independent checkers with one instruction: find the primary source, or mark it unfindable. 70 came back corrected, unfindable, or folklore. Roughly one in four numbers in circulation about writing for the web did not survive being looked up. The model reported the web accurately. That is the problem.
What does the research phase actually produce?
The research phase produces three things, and not one of them is a document.
The standard: what this artefact is supposed to be. The processes, the methodologies, the frameworks, and which of them is an auditable standard somebody can be held to, as against a diagram a consultancy is selling. Those two look identical on a slide and they are not the same object.
The process: the order the thing is made in, which is usually not the order it is presented in.
The benchmark: named cases, what they did, and what specific result followed. This is the one people skip, and it is the one that tells you what to expect and what not to expect. Without it you have a quality bar and no idea whether the bar is reachable from where you are standing.
Hold those three and you can look at a draft and name which part of it is wrong, why, and what it should have said instead. That is the entire output of those hours.
What can you give a model that nobody else can?
There are two inputs a model cannot reach from anywhere except you, and both have to be handed over on purpose.
The first is the instruction to ask. Tell it to come back with clarifying questions before it produces anything, and then answer them. Every question it asks is an assumption it was otherwise going to make silently, in your name, inside a document you are going to put your own credibility behind.
The second is everything that was never published. The meeting recording. What the client said about their own market, in their own words, in a room. The customer interview. The real reason the last agency was fired, which is never in the brief. Every published thing is already inside your competitor's prompt too. The only asymmetric input you have is the one nobody put on the internet.
The Boston Consulting Group experiment that everyone quotes for the productivity gains has a second half that almost nobody quotes. 373 of its consultants were given one business case built so that the right answer was not in the tidy spreadsheet they were handed. It was in a file of interview notes with people inside the company. Consultants working without AI reached the correct recommendation about 84.5% of the time. The group given GPT-4 reached it 70.6% of the time. The group given GPT-4 and training in how to prompt it reached it 60% of the time.
All of them were faster: 6 and 11 minutes quicker to the wrong answer. And human graders, blind to the correct solution, scored the wrong AI-assisted recommendations higher for coherence and for persuasiveness.
That is one task, deliberately built to defeat AI, and nobody has replicated it. Read it as a demonstration and not as a rate. What it demonstrates is the shape of the problem exactly: the answer was sitting in the human-only material, and better prompting made the failure worse rather than better.
What if you do not have 30 hours?
The hours are the part worth arguing about, and the objection comes in two strengths.
The weak one is that nobody has 10 to 30 hours for a task the client is paying one hour for. That is a real constraint. It is also a pricing problem, and pricing problems are not solved by changing the method.
The strong one is better, and I hear it from people whose work I respect: the model already knows the frameworks. Ask it for the standards inside the same prompt, get them in forty seconds, and judge the draft against those. Same second term, two orders of magnitude cheaper.
Asking one system to supply the standard and then to produce the work against it is asking it to set the exam and sit it. If the frame is invented, the check is invented with it, and it will be invented consistently, which is the worst available failure: it removes the disagreement that was your only alert. The 43% figure in my own document arrived in exactly this shape. Something told me what good looked like, and then I graded against it.
There is also a measured reason to arrive with the whole requirement rather than discover it as you go. Philippe Laban and colleagues at Microsoft Research and Salesforce Research took fully specified instructions, cut them into pieces, and released one piece per conversational turn, which is how a working session actually runs. Across six generation tasks and fifteen models, average performance fell 39% against giving the same instruction whole in the first turn.
Almost none of that loss was capability. Best-case scores barely moved, while the gap between a model's best run and its worst more than doubled. Push the same pieces back into one turn and 95.1% of the performance returns, which is what shows the damage comes from the drip rather than from anything left out. It was an Outstanding Paper at ICLR 2026 and the effect has since been reproduced on current models.
Iterating your way towards the requirement is the expensive route to it. The two or three revisions at the end are still worth doing. They are worth doing on top of a full specification rather than instead of one.
Then there is the question of whether you would notice any of this going wrong. METR randomised 246 real issues across 16 experienced developers, in repositories they had maintained for about five years, each issue assigned to allow or forbid AI tools. Allowing AI increased completion time by 19%. The developers had forecast 24% faster before starting, and afterwards, having been measured, still estimated they had been 20% faster.
It is a preprint, 16 people, software work on mature code, and METR has since changed the design and treats the number as historical. Do not read it as AI making people slower. Read it as people being unable to tell.
The strongest published finding against everything above is a preregistered meta-analysis of 106 experiments and 370 effect sizes, in Nature Human Behaviour. On average, humans and AI working together performed significantly worse than the better of the two working alone.
Read its moderators before using that for me or against me. The losses sat in decision tasks and the gains sat in content creation. And the combination won when the human's own baseline was above the AI's, and lost when it was below. That is this argument in one sentence, written by somebody else: the research is how you get your baseline above the model's before you start.
On the hours, the honest answer is that the research is not spent per task. It is spent per domain. Twenty hours on what a marketing plan is actually supposed to contain is paid once, and every plan after it costs the implementation only. Lose the client and you keep the standard.
Why does this work in an industry you have never worked in?
The method carries into an industry you have never worked in because most of what domain experience gives you is a stored standard and a stored set of cases, and that is precisely the part somebody has already written down. What experience does not give you is this company's own situation, which is why the second step exists and why the meeting recording earns its place ahead of the framework.
Here is where the method stops being true. It does not give you taste. It does not give you judgment about people. It will not tell you which of two defensible recommendations a particular board will actually approve, or which sentence a founder will hear as an insult. Where the work turns on something nobody has published and you have not personally sat through, this method has nothing to offer, and it will produce a confident document about it anyway.
I would argue that almost everything I have produced this way sits in the top 1% of what gets made anywhere. That is an argument and not a measurement, and there is no scoreboard I can point you at. Judge it by whether the method holds when you run it yourself.
I might be wrong about the number of hours. I am not wrong about the order.
The last thing the research buys is the one I would not trade. When somebody asks why, in the room, I can name the methodology, the standard behind it, the cases it rests on, the result those cases produced, and the place where our situation differs from theirs. I have worked with 40+ marketing departments, and that meeting always has the same shape. The document gets checked for ten minutes. The person who wrote it gets checked for the rest of the hour.
A draft is cheap now. Being able to answer for it never was.