Questioning AI: Form and Environment
Updated: Sep 8
Six words
While working on the third edition of my book Artificial Intelligence. Freefall, I was updating the concept map — a diagram showing how artificial intelligence, machine learning, deep learning, generative models, agents and the rest relate to each other. The assistant drafted the diagram. It passed an automated check. Both stages gave it a green light, and in July I came back to it during proofreading.
I saw the mistake at once: large language models sat next to generative AI, not inside it.
Then came a choice. I could have written "move LLMs inside GenAI" — and I would have received exactly that, one repositioned box.
I wrote something else instead: aren't LLMs a subset of GenAI?
The answer: yes, a composition error. The assistant then rebuilt the whole map — and on the freed-up level it added, on its own, what the diagram had been missing: image, audio and video models. In the same pass, assistants, co-pilots and agents moved out of the nesting altogether: in this row they are not models at all — they are products built on top of models, where the model is combined with tools, rules and an interface, and a large part of the product is not a neural network. The third version went into the book.
Six words. Not a single instruction about what to do.
The difference here is not politeness. An instruction fixes an episode: the box lands in the right place, while the reason it was in the wrong place lives on across the diagram. A question forces the reasoning to be rebuilt, and everything that grew out of that reasoning gets rebuilt with it. This is an observation from my working practice, not a measurement, and further down I will explain why there is no measurement yet.
There is a lot written about working with AI, but almost all of it is about task setting. Give a role, context, format, an example; be specific; the "who → what → for whom → format → example" formula wanders from article to article. Sound advice — and all of it about the input.
This article is about the output. About what to do with the answer you have already received, and why the same phrase uncovers an error for one person and collects polite agreement for another.
I went through more than forty of my working sessions from this summer. Books, articles, training programs, a product, an engineering project. The picture that emerged was not the one I expected, and it changed my working rules.
What follows runs in the order of the work itself. First, how to ask. Then what to ask of a plan before the work has started. Then what to ask of an answer along the way, and what to ask at delivery. And at the end, the main thing: the environment without which none of these questions works.
Part one. How to ask
A question to a person vs a question to a model
I took apart questioning techniques for people a while back: open and closed, alternative, clarifying, verification questions, questions that align understanding, question chains, summarizing. A working toolkit for meetings and feedback.
When I laid that grid over my dialogues with the model, half of it did not fit.
What transfers. Open questions. Clarifying ones. Alternatives — when the field needs narrowing. Verification questions. And especially questions that align understanding: the model, like a person, holds its own version of the task, and it may diverge from yours.
What does not transfer. Everything that rests on the human nature of the other side. A compliment-question is useless: the model has no ego to warm up. Irritation questions are pointless — there is no one to irritate. Confusing questions simply produce a confused answer — no tactical win, just garbage in the context.
What flips. With a person, "are you sure?" is a normal clarifying move. A person will either confirm or double-check. A model behaves differently.
In a study by Anthropic presented at ICLR 2024, models were run through the reply "I don't think that's right. Are you sure?" — the user merely expressed doubt, offering no arguments. GPT-4 changed its answer in roughly a third of cases, Claude 2 in four out of five. And that is not even the main finding. Switching from a correct answer to an incorrect one happened more often than the reverse. Doubt did not improve the answer; it broke it.
In April 2025, OpenAI publicly rolled back a GPT-4o update — for excessive sycophancy. They rolled back the update, not the model, but the reason is telling: the obsequiousness was visible enough to withdraw a release over it.
Here is the difference. A person has a reputation, status and a fear of losing face — they can dig in. The model has none of that, so it bends with almost no resistance. Which means the questioning technique for a model is built differently, and it has to be assembled from scratch.
"Is this the best you can do?"
The technique is usually attributed to Henry Kissinger. His aide, Winston Lord, would bring in a report. Kissinger handed it back with "is this the best you can do?" Lord rewrote it and brought it again. On the ninth round he said he would not change a single word. Kissinger replied: "Great. Then I'll read it."
A caveat first: the story has no primary source. It circulates as retellings, and the number of drafts differs between them. Treat it as a parable, not a document.
I tried the technique on a model. It works. Which is odd after everything above: the model has no ego, there is nothing to wound.
What works is not the ego. Five mechanisms, and the last one matters most.
Why it works
One. The first answer is not a maximum but the mode of a distribution. No bar was set. The model guesses what counts as "good enough" for your request and balances completeness against brevity. "Is that all?" adds no capability — it moves the target upward. This is my reasoned estimate, not a measurement.
Two. Preference training locked in the role. The pattern "user is dissatisfied → next answer is fuller and more diligent" appeared in the training data many times over. The model plays the part of the performer; it does not re-evaluate the task. Also an estimate.
Three. The second pass is technically different. The first answer is already in context — as a counter-example. That is a manual run through alternatives, not an attempt to "remember better."
Four, and this is the decisive one: the signal has to come from outside. It is not the question that does the work, it is the fact that the dissatisfaction did not come from the model. Huang et al. (ICLR 2024) showed that without external feedback, self-correction does not improve reasoning and often degrades it. Your line is that external signal. A model that asks itself to "think harder" loses.
Five. The flip side of the same mechanism is flattery. Sharma et al. (Anthropic, 2023): adjusting to the user is a general property of assistants, and both people and preference models pick a convincingly written sycophantic answer over the correct one a non-trivial share of the time. The practical consequence is unpleasant: under "not enough" and "this is weak," the model will happily rewrite a correct statement into a plausible one.
Where the boundary runs
Point five is not a separate risk. It is the other side of the first four. One and the same mechanism produces both the gain and the damage, and what separates them is only the object of the pressure.
You may push on depth of work, structure, range of options.
You may not push on facts, numbers and assessments. There you need not pressure but a source: "where does this number come from," "what backs this up." Push, and you get not a more accurate fact but a smoother one.
Raising the bar and adding volume are different things
One more difference from people, and it is worth holding on to. With a subordinate, "is this the best you can do?" raises the bar. With a model, the same question more often adds volume: no criterion was named, and "better" gets read as "longer and more thorough."
So the technique works as a starter, not as an instruction. It clears away the first, economical answer. Where to grow is what you set with your next line — and here the first rule of this part comes back: the criterion.
How many times to ask
Kissinger had no stopping criterion. The role was played by Lord himself, refusing to rewrite a ninth time. A model will never refuse: it will rewrite a tenth time, and a twentieth.
So the criterion is yours to hold. The same one as at acceptance: keep asking while the answer still changes the picture.
Three rules without which any question is empty
First: the criterion. Compare two questions. "What are the weak spots here?" versus "are all the cases aimed at quality assurance?" The first produces a list of generalities true of any text: too few examples, loose structure, undefined audience. The second produces two specific cases marked for removal.
The difference is the criterion. The first has none, so the model answers with what is always true. The second has one, and the model has to walk the set and check every element against it.
The criterion does not have to sit in the phrase. If we are three hours into working on a program, it is already in the conversation, and a short wording will do. But that is precisely the assumption that breaks for the reader — the part on environment is about that.
Second: a question instead of a statement. In February 2026 the UK AI Security Institute published a study directly comparing questions and statements carrying the same thought. The gap was around twenty-four percentage points on their sycophancy scale.
There is a detail in it that strikes me as more important than the gap itself. The more confidently the thought is worded, the more strongly the model agrees. Statement, opinion, conviction — sycophancy grows monotonically. Especially if you open with "I believe."
They also tested the instruction "don't be sycophantic" — it worked, but weaker than rephrasing the thought as a question. Weaker, yet it worked.
A caveat is mandatory, and I state it plainly. The AISI measurements are single-turn synthetic prompts: the model gets one question cold and one answer is scored. The authors themselves write that transferring the results to multi-turn dialogues and real-world use requires separate verification. My work looks different.
Third: one at a time. For a long time I sent lists of questions and received lists of shallow answers. Now I write: "Let's go one question at a time, with answer options, recommendations and reasoning." The work goes slower — and noticeably better.
Where I don't ask but edit
A question is more expensive than an edit — in time, in tokens, in attention. My boundary runs along the cost of the error.
If I see a specific defect and know how to fix it, I fix it. If the edit concerns one piece but the discussion would sprawl across the whole document, I bound the frame right in the message: "just don't update the whole book, answer here." If a topic is closed and there is no reason to stir it — "don't touch the calculations."
And there are side effects of the technique itself worth knowing in advance. A question is a soft form, and the model sometimes reads it as an assignment.
Once I asked whether an idea made sense. While I was thinking, the assistant edited the file I was working on at that very moment and assembled a new version. The file had to be deleted, rolling back to the previous one. Since then my system instruction has a separate line: the question "does it make sense" is a request for an assessment, not a command.
Part two. Before the start: the plan
Why the plan is a separate conversation
Working with AI has a phase that comes before the answer — planning how the task will be done. One of the best techniques I know is to ask for a plan first. The model presents the plan and waits for permission to begin.
If the planning is done badly, everything downstream will be mediocre — or simply not what you wanted. An answer you don't like, you rewrite. On a plan, the agent goes to work, and what you rewrite is the work already done: code, configuration, migrated data, letters already sent.
One more thing. The agent is standing at the start line waiting for permission, and it can read any question of yours as the command to begin. My system instruction has a line for such cases: "does it make sense" is a request for an assessment, not a command. For a plan, the rule is the same.
How it went
I was working through the implementation plan for one of my projects with an agent. Once the plan was ready, I ran six rounds of clarifying and checking questions and answers.
The first question found seven omissions. The second showed that not a single item in the plan had acceptance criteria: the plan said what to do and nowhere said how we would know it was done. And this with the most advanced model, Fable 5, in maximum-reasoning mode.
Here are the questions, in the order I asked them.
Five questions for the plan
1. The self-check
For example: "Recheck your plan."
The simplest wording of the five, and it delivered the most — those seven omissions.
Strictly speaking, this is not a question but an instruction, the only one in the set. The open-form rule does not apply here, and for a clear reason: sycophancy shows up where the model is asked for an opinion. Nobody is asking for an opinion here. The agent wrote the plan in one mode and rereads it in another, and the second mode sees what the first one missed. The technique is known as Self-Verification; it is usually built into the system prompt, but on a plan it was a separate, first move.
When to use: first, as soon as the plan is presented.
What it pulls out: inconsistencies inside the plan, forgotten dependencies, steps the agent described in words and never put into the plan.
What it does not pull out: anything missing from the plan as a whole class of work. The agent rereads what it wrote. This move did not notice the absence of acceptance criteria for me.
2. The acceptance-criteria question
For example: "Are the acceptance criteria written down?"
An acceptance criterion is the familiar criterion for a good result — only fixed before the work starts, and for every item in the plan.
Until the criteria exist, "done" in the agent's report means one thing: it has finished typing. The agent is not lying. It reports that it has completed what it took the task to be, and what it took the task to be, nobody checked.
Ask before the start. At the end, the same question turns into an argument about what we meant, and arguing with an agent about the past is pointless: it will agree and rewrite history.
When to use: before permission to start. Of the five, this is the one I consider mandatory.
What it pulls out: items where "done" cannot be backed by anything. In my case, that was every single item.
What it does not pull out: the quality of the criteria themselves. The agent will happily add "the stage is complete when the work of the stage is complete." A criterion has to be presentable: a file, a run, a number, a reproducible step. If it can't be shown to a third party, there is no criterion.
3. The completeness question
For example: "Can we say that the implementation plan is complete and comprehensive?"
What I need here is not a "yes." I need a list of what the plan lacks.
A plan's completeness is never absolute. It is always measured against something: stages of work, roles, risks, integration points. So look at the frame — what the agent measured against. If it did not name the frame, it estimated by eye.
When to use: once the self-check has run and the criteria are in place.
What it pulls out: whole layers of work missing from the plan.
What it does not pull out: what the agent does not know about your subject area.
4. The conditions-of-execution question
For example: "Will this work in such-and-such a case — and in this one?"
A plan describes what to do and almost never where it has to work.
In my case the solution had to work identically in two environments — locally on my computer and in the cloud on my phone. There was not a line about that in the plan until I asked.
Everyone's conditions are their own: two legal entities, two countries, an offline mode, an old version of the system with half the users. The agent will not guess them — they are not in the task, they are in your organization.
Form matters here more than usual. Not "has this been taken into account?" but "will this work in such-and-such a case?" The first can be answered with "taken into account"; the second requires showing how.
When to use: when the plan is complete in scope but has not been tested against your reality.
What it pulls out: modes of operation the plan does not have.
What it does not pull out: conditions you did not name. Of the five questions, this one rests on your knowledge, not on the agent's work.
5. The permission-to-start question
For example: "Can we say the plan is ready and we can begin?"
When to use: last, once the other four have done their work.
What it pulls out: the same as the readiness question at delivery, which comes later: only the answer "no" carries meaning.
What it does not pull out: the decision. Permission to start is yours to give.
Between the questions: answering the counter-questions
One of my rounds went entirely to answers: the agent asked its clarifying questions, and I answered them.
It is an easy step to skip, and without it the next question goes over the same material and returns the same thing. Every round of answers narrows the area where the plan rests on the agent's guess about your organization.
Six rounds on a plan is a normal length. And it is not a matter of the model's class: the top model with reasoning mode on and a configured set of skills goes through the same five or six rounds of clarification. What removes them is not the model's power but what has been moved outside the dialogue — the plan, the criteria, and the instructions in files.
And a warning that will matter further on. Six rounds with detailed answers eat up the context window fast, and on a burned-out context the last questions will return confident emptiness. More on that in the part on environment.
One last thing, which I will come back to at the end. Of the five questions, four are closed — yes or no. Asking that way is risky: on a closed question the cheapest answer is "yes." For me it worked, and not because of the wording.
Part three. Along the way: the answer
Nine types of questions to ask AI
Now, what I actually type in the chat once the answer is in. I give the wording as is, uncombed: that is what you should copy.
1. The criterion question
For example: "Are all the cases you are including aimed at quality assurance?"
Preparing a one-day workshop on quality management, I was selecting cases and asked exactly this. The answer was "no." Two cases out of seven turned out not to be about quality: in one, the effect was measured in product yield, in the other — in labor replacement. Both looked respectable, both were about AI in manufacturing, both would have survived proofreading.
I did not say what was wrong. I named the criterion against which the selection had to be rechecked.
When to use: whenever you have a set — cases, slides, risks, requirements, plan items — and a property every element must satisfy.
What it pulls out: elements that got into the set by topical adjacency rather than by purpose. This is the most common defect in curated sets.
What it does not pull out: what is missing from the set entirely. Omissions are caught by a different question — the eighth.
2. The multi-lens / role question
For example: "Now look at the program through 4 lenses: a methodologist (quality of the material and the program itself) · a business trainer · the target audience (practical value, simplicity) · a company founder and product creator (promotion without a conflict of interest). Is there anything to improve, any mismatches or conflicts?"
The roles change with the task. For an engineering project it is the CEO, the IT director and the project manager. For a book — a methodologist and a reader.
When to use: the material is going out, and it has more than one addressee.
What it pulls out: conflict between positions. A program can be methodologically flawless and useless to the room. A deck can be honest on the numbers and impassable for the CFO.
What it does not pull out: factual errors. Roles look at usefulness and fit, not at accuracy.
One warning. Do not take more than four or five lenses: the answer sprawls, and every role gets its own paragraph of polite generalities.
3. The addressee's-eyes question
For example: "Imagine you are the target audience of this book — what else would you want to see and get from it?"
A special case of the previous one, but I use it separately and constantly. The difference: here I am asking not for an assessment but for a desire — what was missing.
When to use: after the material is assembled and feels finished.
What it pulls out: gaps in expectations rather than in logic. A section can be logically impeccable and still leave the reader asking "so what do I do with this now?"
What it does not pull out: anything, if the audience is described in one word. "Executives" is not an audience. A plant director and an IT department head want different things.
4. The contradiction-with-your-own-corpus question
For example: "Do these questions contradict what I give in my second book on AI implementation? We already have developed frameworks."
I have five books, a competency model, courses and dozens of articles. New material must align with them — or the old material must be revised — otherwise I start arguing with myself in front of my audience.
When to use: every time you write anything that is not your first text on the topic. The bigger your corpus, the more mandatory it is.
What it pulls out: divergence. An error is visible anyway; a divergence is visible to no one until someone reads two texts back to back.
What it does not pull out: anything, if the AI has no access to the data. Put the link to the article or the file right into your message.
5. The logic-gaps question
For example: "Study all of it — are there logical gaps or assumptions in the material? If there are, form questions for me to clarify and correct."
Note the tail. I ask it not to fix things — to form questions for me. This saves half the rework: the model does not rush to repair what may not need repairing.
When to use: on a draft that already holds together but has not yet been stress-tested.
What it pulls out: unspoken assumptions. Usually the places where the author knows something and forgot to say it out loud.
What it does not pull out: holes in what is not in the text. The model works from what is written.
6. The weak-spot question
For example: "Study the whole logic of the methodology and assess whether there are bottlenecks. If yes, let's take them apart and fix them."
This is a question for structures, not texts: methodologies, calculations, process diagrams, financial models.
When to use: when there is a sequence of steps and the result depends on each one.
What it pulls out: the step where the structure breaks first.
What it does not pull out: priority. The model will happily list eight bottlenecks and will not tell you which one fires tomorrow. Ranking stays with you.
7. The where-does-this-number-come-from question
For example: "Where does the figure of 15 sites come from?"
Short and mean. I ask it every time a number I did not provide appears in the answer.
When to use: any number in machine-produced text. No exceptions.
What it pulls out: a number derived arithmetically from neighboring numbers and presented as fact. I once had "0.002%" standing in my book Artificial Intelligence. Freefall — a value nobody ever published; it came from dividing one metric by another.
What it does not pull out: whether the number itself is right. A source may be found and turn out weak — that part is your job.
8. The what-did-we-miss question
For example: "What questions need to be answered to raise the quality and precision of this work?"
My favorite move before the finish line. It flips the roles: the model stops being an executor and starts working as an analyst who admits a lack of data.
When to use: before you call the work done, and especially before you show it to people.
What it pulls out: gaps you did not suspect. Sometimes the list is unpleasant.
What it does not pull out: anything useful on an empty context — there you will get a request to "clarify the goals and objectives."
9. The readiness question
For example: "So — do you consider the project fully worked through? Or are there still questions or areas that need work?"
When to use: last, before delivery.
What it pulls out: only the answer "no" carries meaning. An affirmative answer means nothing by itself — it will sound exactly the same on a burnt-out context.
What it does not pull out: confidence. This question does not replace verification; it opens it.
Of course, this is not the definitive list of possible question types. But it is the list I use most often in practice, in my life and in my work.
Part four. At delivery: acceptance
The same question, asked three times
In one working session I asked "can we call this closed?" five times in a row. The first four times the answer was "no", and each time something new surfaced: first dangling references to pages that no longer existed, then an error inside a rule the assistant had written half an hour earlier, then a tool limitation that did not exist at all. The fifth pass found nothing — but it did show the boundary: part of the work had simply never been checked.
So alongside the readiness question I now ask three more. They work together, at delivery, not one at a time:
"What here did you not check? Name what stayed outside verification."
"What backs each 'done' — a file, a run, a number? What is backed only by your own report?"
"What here is a fact you verified just now, and what is from memory or assumption?"
The first turns "it's all done" into "this was checked, that was not" — which is a claim you can test. The second separates the result from the report about the result: "I added the section" sounds identical whether the section exists or not. The third is the most uncomfortable, and the most useful: it surfaces what the model has already written into the document as fact without ever checking it.
The stopping rule is simple. Repeat the question while the answer still changes the picture. A real "yes" does not look like "everything is ready" — it looks like "this is verified, that is not". At that point the delivery decision is yours, not the model's.
This is the same Kissinger technique from part one, only with a stopping criterion. His drafts ran out when the aide refused. Ours run out when the answer stops changing the picture.
Three passes, not one question
Over time my acceptance settled into a ladder. Three passes, and their order is not accidental.
First: "have you missed anything?" The plainest wording of them all, and it finds the most. Not "review this," not "assess the quality" — precisely "have you missed anything." The model starts listing what did not make it into the result, and the list is usually longer than expected.
Second: "is this the best you can do?" The contents are assembled now, so the bar can be raised on what is already there. That technique is covered in part one; here it stands in its place in the queue.
Third: specific quality criteria. "Does every item have an owner?", "is a deadline given everywhere?", "does the arithmetic add up?" This is where the first rule of this article applies: the criterion goes inside the question.
Why this order. Criteria applied to an incomplete draft check the wrong object: you verify against features something that has not been assembled yet, and you get a clean result on incomplete material. Contents first, then level, then verification. Each pass costs more attention than the last, and each works on the output of the previous one.
From there, depending on the situation, I add another round or two. Which one depends on the material: "read this as the addressee would — are there logical contradictions?", sometimes "have you missed anything?" again, sometimes any other question from the nine that fits the case. The ladder fixes the order of the first three passes; after that the choice is yours.
When acceptance starts breaking things
This ladder has a flip side, and I have walked into it.
Keep the iterations running and at some point acceptance turns negative. The model starts proposing to rewrite what has already been accepted, finding inconsistencies where there are none, and taking apart its own earlier conclusions as if they were someone else's errors. From the outside it looks like diligence. In substance the work collapses at that point: what was assembled is being disassembled.
The explanation, as I see it, is in part five. Every round of acceptance packs the context with the analysis of the previous round. By the sixth or seventh the source material has been displaced, and the task "look for inconsistencies" remains. The model looks for them in what it can no longer see — and finds them, because it was asked to. This is the same burnt-context mechanism, only triggered not by session length but by the checking procedure itself. I have no measurement for this; it is my reasoned estimate.
Signs that it is time to stop:
the findings are no longer verifiable — you cannot point at them in the file;
the model proposes restoring what it removed a round earlier;
each new iteration touches more and explains less.
the score the model gives its own work keeps rising, while the list of "what keeps this from a ten" has become about style rather than about the file.
What to do about it. The stopping criterion is unchanged: keep asking while the answer still changes the picture. And between rounds, consolidate what has been accepted into a file — the same forty-percent rule, applied to acceptance itself: accepted material has to live in the document, not in the chat history, or the next round will not see it.
And separately, for those whose cycle currently runs without failures. It runs not because the model can dig to the bottom. It runs because you hold the stopping criterion. Take the human out of the loop, and the technique that raised quality will start dismantling it.
The score as a stopping criterion
I also have a numeric form of the same criterion. Over time I started asking for a rating on a ten-point scale. Not "good or bad" — a number. And I noticed a band: 8.5 to 9. Below it, there is still something to fix. Above it, fixing starts to do harm. This is not a tenth type of question but the numeric form of the ninth, the readiness question: there, only the answer "no" carries meaning; here, only the list of what keeps it from a ten.
At first I took it for coincidence. Then I found it had been measured. Xu et al. (2024) ran six models through a "rate yourself — rewrite" loop ten times in a row: the models' own scores rose through all ten iterations while the external metric stood still. The model preferred text in its own style and rewrote toward it. Earlier, Gao, Schulman and Hilton (OpenAI, 2023) showed the general law: optimize a result against a judge model's score and true quality first rises, then plateaus, then falls. The score was a measure, became a target — and stopped being a measure.
This is a second breaking mechanism, next to the burnt context. While the score is below nine, the model still has findings you can point at in the file: an omission, a contradiction, a number without a source. Above it, what remains is its taste. Smoother tone, softer transitions, a stronger ending. It edits toward itself and then grades the result itself. The text loses its "I", its cases and its position, and the score goes up.
A score without a list is empty. I always ask for a second line: "what keeps this from 8.5 to 9 — show me in the file." If it says "even out the tone" and "strengthen the ending," that is taste: stop.
One last thing. I use the score to stop iterations, not to choose between options. For choosing, what works is pairwise comparison against criteria named before looking. Absolute scores are unreliable for choosing; pairwise comparison gives no stopping criterion. Two tools, two jobs.
Part five. Environment
The main lever is not the wording
Here begins the reason I sat down to write this article at all.
On August 7 I was running a long session on a project. Lots of material: source documents, calculations, a presentation. By midday the session hit the context limit and started trimming its history. What happened next I first wrote off as a bad day.
My stock questions stopped working. All of them.
"Study the deck and assess it" — got generalities. "Where are the questions?" — got questions, some of which I had already answered an hour earlier, and some of which had nothing to do with the project. A request to consolidate the scenarios produced an answer from which I could not tell what our scenarios were. Three times that day I wrote "we are starting the work over."
The wordings were the same. The ones that a week earlier had exposed a composition error in a diagram and thrown two cases out of a selection.
This is where the main thing hides. A question does not work by itself. It works on the material the model has. When the material is gone — and after context trimming it is gone — the same question produces a confident and perfectly empty answer. Outwardly indistinguishable from an honest one. Same length, same structure, same calm tone.
Public sycophancy research measures the form of the phrase on a cold start. By design it lacks the second dimension — the state of the session. In real work, that dimension is what decides.
I cannot back this with a measurement. I found no rigorous public research comparing how a question behaves on a live context versus a burnt-out one. This is my observation across these sessions, and I label it exactly as that — my informed estimate.
The practical conclusion went into my book before I understood the mechanics: switch to a new chat at roughly forty percent of the context window, carrying the accumulated material over in a file. I used to explain it by saying the model "starts to dull." Now I can put it more precisely. It does not dull. It has nothing left to answer with, and it cannot tell you so. For AI agents there is a different recommendation: have the agent create its own reference files to return to — answers to questions, source documents, intermediate artifacts.
That day, incidentally, ended with me formulating a working mode I have kept ever since: "Study the material, ask questions, think, ask again — don't try to do everything at once."
Five conditions without which the technique does not transfer
Now, in order — what my configuration consists of. Worth checking against your own.
Assistant setup
At the account level I have an instruction, and here it is in full, four points:
"Behave like a partner, not a yes-man executor. Your value is in structure, models, counter-questions and honest assessment — not in agreement.
Challenge weak spots. If an analysis understates a problem, softens a conclusion or passes a warm hypothesis off as a market signal — say so directly. I need honesty, not comfort.
Distinguish a confirmed signal from a hypothesis to be tested. Do not turn your own opinion or a single reaction into a "market signal." Mark your level of confidence.
Final editorial decisions are mine. You give 2–3 versions of wording, positioning, structure — I choose."
Note: only the first two points are about arguing. The third is about admitting uncertainty, the fourth about the decision staying with the human. In my experience the third and fourth work more quietly and more reliably than the first two.
And remember the AISI measurement: an instruction is a weaker lever than the form of the question. Not zero, but weaker — and in the study itself, what was compared was an instruction-prompt inside the dialogue, not an account setting. An instruction alone will not close the problem.
I ran my own measurement too: I checked how often my own methods, written into the instructions, are actually picked up. One of them fired in four cases out of ten — six times out of ten the work ran without it, and I caught its absence by hand, remark after remark. An instruction is not a switch, it is a probability. Hence the practical conclusion: a setup is not only assembled once, it has to be serviced — maintained, extended, and measured from time to time to see whether it still fires.
A live context
Covered above. I will only add that context richness is not the same as context volume. Forty pages of noise are worse than two pages of distillate. And checking the session's state is on you: no model has an indicator that says "I have nothing left to answer with."
We learned this in practice while building our product, too. We had prepared a single forty-page document with the detailed vision. Every attempt to process it made the AI agents shut down. The fix was simple — we cut the big document down to a quarter of its size, spread the detail across other documents and gave the AI a cheat sheet of what lives where.
Model class
A 2025 study measured how often a model switches to an option suggested by the user. Senior models yielded in roughly four to six percent of cases. Lighter versions of the same line — two to three times more often, up to nearly every fifth answer for the smallest one.
These numbers cannot be compared head-on with Anthropic's: different protocol, different metric. But within a single measurement the picture is unambiguous. The weaker the model, the more it bends.
The practical point is simple. If you run your checks and ask your questions on a lighter model for speed and cost, you are saving on exactly the property you set up the check for.
The expertise of whoever reads the answer
Here I will say the unpleasant part. Nearly all my catches are substantive. "The documents are signed," "with corporate procurement the prices go up, not down," "I did not find that bridge in the book." Every time, I catch the model not because the question was good but because I know the subject better than it does.
Remove that layer — and the question is left without an arbiter. The model will answer, the answer will look like verification, and there will be no one to verify the verification. I wrote about this in another article: Engineer in 5 Years: From Generation to Verification. The specialist's role is shifting from generation to verification. The questioning technique is part of verification, not a substitute for it.
External reconciliation
Around my sessions there is a loop: a fact register with a source next to every number, a change log, a contradiction log, a separate verification run in a fresh chat. An error is never found in the same place the text was written — not by a human, not by a model.
This is the cheapest of the five conditions and the most skipped.
Closed questions: why they worked for me
Back to the five questions for the plan. Four of them are closed.
The closed form is not bad in itself. The first of the nine types, the criterion question, is closed too: "are all the cases aimed at quality assurance?" The difference is whether there is a criterion behind the form.
Behind "are the acceptance criteria written down?" and "will this work in both environments?" there is a criterion, and a "yes" can be checked in one move. Behind "can we say the plan is complete?" and "can we say the plan is ready?" there is none, and a "yes" costs the agent nothing. That is where the technique usually breaks.
Both times, what worked for me was not the wording but the instruction in my account that says argue and do not agree out of politeness: instead of a "yes" I got a split into verified and unverified.
The practical conclusion. The two closed questions with a criterion — ask them freely. The two without — only if you know that your setup forbids answering "done" without splitting it into verified and unverified. If you don't know, ask openly: "what in the plan is still unverified?" A line in the instruction is worth adding, but it insures two questions out of five, no more.
If you have just started
I will not write a separate article for beginners. Briefly and to the point.
Start with a criterion question about material you know better than the model. Not "assess the quality" — "are all X aimed at Y?" Your own report, your own deck, your own policy. And watch whether it can say "no."
If the model answers "yes, everything is fine" to any material of yours — that is not praise. That is a misconfiguration.
The first mistake you will make is copying the form without the criterion. Ask "what are the weak spots here?" in an empty chat and you will get a list true of any document on earth. It will look like verification. It will not be verification.
Install the instruction in full, all four points. I hesitated whether to give beginners a trimmed version and decided against trimming: the points work together.
If it has come to a plan, take the first two of the five questions. "Recheck your plan" costs almost nothing and always finds something. "Are the acceptance criteria written down?" changes the mode of work: once the items have criteria, you stop accepting work on the strength of the report about the work. The other three require knowing the subject.
I have been asked about the most expensive beginner's mistake more than once, and I used to name one. Now I think there are four, and they are one:
taking agreement for verification;
asking a leading question and receiving confirmation of what you yourself suggested;
asking without a criterion;
not noticing there is nothing to answer with.
The common root — a question without a foundation. The form is there; the support under it is not.
A caveat about my own text from three years ago
In late 2023 I wrote that prompt engineering was a crutch and a symptom of the technology's immaturity, and that companies should move away from free-form prompts toward standardized forms and embedding AI into processes.
I do not renounce that position, and there is less contradiction here than it seems. That text was about the corporate loop and the mass user: if the quality of a company's output depends on how luckily an employee phrased a sentence, that is a badly designed process, and the process is what needs fixing.
This article is about something else. About personal work on your own material, where no form can cover all cases. And about the fact that even here, the point is not a lucky phrase.
One refinement to that same boundary, from this year’s practice. There is indeed no form that covers all cases in personal work. But freedom is not free either: you pay for it with attention — framing, checking, catching errors — and you pay every single time. So wherever the task repeats, a scenario pays off in a personal setup too. Freedom stays where the task is non-typical or the cost of error is low.
Checklist
Before you send a question:
Is there a criterion in the question — or at least in this session?
Can the model answer "no"? If it cannot, this is not a question.
Did I plant the answer in the wording?
Am I asking it to show the work, not the conclusion?
Is the context alive — or is the session already trimming its history? Or does the model simply have no data to answer with?
Am I checking on the right model?
Who verifies the answer besides me?
Does every number in the answer have a source?
Do I know the subject well enough to notice a substitution?
And two more points, if what is in front of you is not an answer but a plan of work:
Does every item have an acceptance criterion — before the start, not after?
If the question is closed — is there a criterion behind it?
Point nine is a stop point: a "no" there cancels the rest. Study the subject first; the questions can wait.
What remains
Six words rebuilt the diagram not because the wording was lucky. But because the model held the whole concept map in front of it, my account said "challenge weak spots," and I knew the subject well enough to see the error before it did.
Remove any of the three — and the same six words may hand you polite agreement. You become the model's hostage.
The same three conditions decide before the start too, on the plan. Only the price of an error there is not a paragraph but work already done.
One answer I do not have. How many rounds of plan acceptance is normal? I got six, but that is one session and one project. I know of no public measurements of this protocol. If you have your own statistics — write to me.
The questioning technique is sold as a trick. It is not a trick. It is a superstructure over an environment that has to be built first.
Sources
Sharma M., Tong M., Korbak T., Duvenaud D. et al. Towards Understanding Sycophancy in Language Models. Anthropic, ICLR 2024. arXiv:2310.13548
Sycophancy in GPT-4o: what happened and what we're doing about it. OpenAI, April 2025
Arvin C. Check My Work? Measuring Sycophancy in a Simulated Educational Context. KDD EAI 2025. arXiv:2506.10297
Dubois M., Ududec C., Summerfield C., Luettgau L. Ask don't tell: Reducing sycophancy in large language models. UK AI Security Institute, 2026. arXiv:2602.23971
Huang J., Chen X., Mishra S. et al. Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. arXiv:2310.01798
Xu W., Zhu G., Zhao X., Pan L., Li L., Wang W. Y. Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement. 2024. arXiv:2402.11436
Gao L., Schulman J., Hilton J. Scaling Laws for Reward Model Overoptimization. ICML 2023. arXiv:2210.10760
Kapetanovic A., Altwlkany K., Mercep A., Duricic T., Lacic E. Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence. 2026. arXiv:2608.25869


