Getting Consistent, Reliable Answers
“I think the potential of what the Internet is going to do to society, both good and bad, is unimaginable. I think we’re actually on the cusp of something exhilarating and terrifying.” — David Bowie, BBC Newsnight interview, 1999
epigraph-ch32-tuning-dial.png
A vintage radio tuning dial with the needle centered precisely on a station mark.
A vintage radio tuning dial in 8-bit pixel art, approximately 48x36 pixels scaled up. A horizontal dial window with tick marks and frequency numbers too small to read as digits (rendered as short flat color bars, not legible text), and a single thin needle pointer centered exactly on one tick mark --- precision, not ambiguity, is the point of the image. Palette: Recycled for the dial face, Black for tick marks and needle, Red for the single tuned-station mark the needle points to, Charcoal for the dial's outer casing edge. Colored in ballpoint pen and pencil per the inline-graphics style. **True pixel grid --- non-negotiable.** Actual 8-bit pixel art, not a blocky illustration. Render on a literal uniform grid of axis-aligned square pixels, as if captured from an NES framebuffer and enlarged with nearest-neighbor scaling. Every pixel the same size, hard edges between colors, no anti-aliasing, no smooth curves, no sub-pixel detail. The viewer should be able to count pixels along an edge. **No meta-elements --- non-negotiable.** Only the dial. No color swatches, palettes, legends, keys, hex codes, callouts, arrows pointing at the needle, sidebar text, or any UI explaining the colors or technique. No caption text inside the image. **Background: pure white, `#FFFFFF`, flat.** Not gray, not off-white, not cream, not paper texture. The build removes white to create transparency, so any gray will show as a halo in the ePub. **Watch out for:** - NO readable frequency numbers --- tick marks only, this is not a labeled diagram - NO hand or figure turning the dial --- object-focused - NO full radio body/speaker, just the dial window and its immediate casing edge - NO speech bubble, thought bubble, or dialogue text
Your Agent does not always give identical answers to identical questions. This is by design — the same way two conversations with the same knowledgeable friend might go differently depending on how the question was framed. For most tasks in this book, that variability is fine. For others — when you need a checklist, a step-by-step process, or a conservative interpretation of something with real stakes — you want to reduce variability and increase reliability.
Here is how.
Ask for Numbered Steps When Sequence Matters
If the order of operations is important — a legal process, a repair procedure, a medical preparation checklist — ask explicitly for numbered steps. Your Agent will default to numbered steps only if you ask for them.
Please give me numbered steps, in the order I should do them, with no steps combined.
Ask Your Agent to Show Its Work
For factual claims with real stakes, ask your Agent to identify the basis for each main point and flag any areas where it is less certain. This is not about doubting your Agent — it is about knowing where to verify independently.
For each main point in your answer, please tell me briefly what it is based on, and flag anything I should independently verify with a professional.
Ask for the Most Conservative Interpretation
For legal, medical, and financial questions:
Please give me the most cautious, conservative interpretation. I would rather over-prepare than be caught off guard.
Ask for a Checklist Instead of Advice
When you need to act on something — not just understand it — ask your Agent to convert its answer into a yes/no or done/not-done checklist.
Can you turn that into a checklist I can print out and work through item by item?
Ask Your Agent to Review Its Own Answer
For important questions, ask your Agent to take a second pass before you act on the answer:
Now review what you just told me. What did you get wrong, oversimplify, or leave out? What would change if my situation were slightly different?
Your Agent will often catch its own oversights — a missed exception, an assumption it made about your circumstances, a detail it glossed over for brevity. This is not a sign the first answer was unreliable. It is the same principle as rereading a contract before signing: the second read catches what the first read missed. For high-stakes questions — legal, medical, financial — a self-review followed by the calibration question later in this chapter gives you two layers of error-checking before you act.
Check the Spec Before You Trust the Answer
Your Agent can give a perfectly coherent, well-reasoned answer to the wrong version of your question. The answer looks right because it is right — for someone else’s situation. Before you act on a plan, an interpretation, or a recommendation, ask your Agent to tell you what it thinks you asked:
Before I act on this — restate what you understood my situation to be, what you think I am trying to achieve, and what constraints you were working within.
Read the read-back against what you actually said. You are looking for three things:
Goal drift. You asked for options. It gave you a plan. You wanted to understand a bill. It jumped straight to disputing it. The shape of the answer does not match the shape of the question. This happens because your Agent optimizes for helpfulness, and “helpful” often means “action-oriented” even when you asked for understanding.
Dropped constraints. You said “I want to avoid legal escalation” and the first recommendation is to file a formal complaint with a regulatory agency. You said “keep this under $500” and the proposed solution costs $1,200. A constraint you stated got treated as a preference rather than a boundary.
Invented facts. Your Agent assumed your state, your income bracket, your timeline, or your relationship to the other party. The assumption is plausible — that is what makes it dangerous. It filled in a detail you did not provide, and built its answer on top of it.
When you spot one of these, you do not need to start over. Correct the specific misunderstanding:
You assumed I am in California — I am in Texas. And I asked for options, not a step-by-step plan. Can you redo this with those corrections?
The read-back checks whether your Agent understood your situation. The calibration question, covered later in this chapter, checks whether its answer is reliable. These are two different failure modes — wrong understanding versus uncertain knowledge — and they need two different checks.
Recognizing When the Spec Is Right
The failure modes tell you when the spec is wrong. But the harder question is the opposite: when is it right enough to act on?
A correct spec is not a perfect spec. Perfection is a stalling tactic. A correct spec has four properties you can check:
It uses its own words, not yours. When your Agent restates your situation, it should paraphrase, not parrot. If the read-back is your exact sentences rearranged, your Agent may be reflecting your words without processing them. A real understanding sounds slightly different from how you said it — the way a good listener says “so what you’re dealing with is…” rather than reading your message back to you.
It names your specifics, not generic categories. “You’re dealing with a landlord dispute” is a category. “Your landlord in Texas has withheld your $1,200 deposit for 60 days citing carpet damage you photographed on move-in” is your situation. If the read-back could apply to anyone with a vaguely similar problem, it has not locked onto yours.
It admits what it does not know. A spec that fills in every detail without asking is guessing. Your Agent should flag assumptions: “I’m assuming you haven’t already contacted the landlord in writing — is that correct?” The absence of any hedging is not confidence. It is a sign that gaps are being papered over.
You could explain the plan to someone else. Not the Agent’s plan — your plan, in your words, after reading the Agent’s output. If you cannot summarize what you are about to do and why, you do not understand the recommendation well enough to act on it. This is not a test of the Agent. It is a test of your readiness.
When all four properties hold and the three failure modes are absent, the spec is correct. Act on it — or take it to the professional. The goal is not to keep refining until the answer is perfect. The goal is to refine until you are prepared, and then to move.
When to Start a New Conversation
If you have been going back and forth for 30 messages and the Agent starts repeating itself, losing track of earlier points, or giving vague answers — the conversation is not broken. It is full. Your Agent’s working memory holds a lot, but not everything, and the middle of a long conversation gets less reliable attention than the beginning and end.
The fix is not to push harder. It is to start fresh — but not from scratch. Ask your Agent to summarize what you have established so far:
Summarize the key facts, decisions, and open questions from this conversation. I am going to start a new conversation and I want to carry forward only what matters.
Copy the summary. Open a new conversation. Paste it in with your next question. You now have a clean workspace with all the context that matters and none of the noise that doesn’t. Two focused conversations will outperform one sprawling one every time.
This is the same principle the billionaire’s chief of staff applied to the morning briefing: read the whole file, extract the page that matters, hand over one clean sheet. You are the chief of staff. Your Agent is the principal. Brief it well.
The Calibration Question
After any important answer, before you close the conversation:
How confident are you in this answer, and what are the two or three things I should independently verify before acting on it?
This single question will consistently improve the quality of the guidance you receive. Your Agent will tell you where its answer is reliable and where it is extrapolating. That boundary is exactly where you should apply your own judgment — or pick up the phone and call a professional. It is the same line Chapter 31 draws from the other direction: where AI’s edge ends and the human’s begins.
-
[1] Wei, J. et al. (2022). “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.” NeurIPS 2022. The finding holds most strongly for sufficiently large models on multi-step arithmetic, commonsense, and symbolic reasoning problems.↩
-
[2] Dietvorst, B.J., Simmons, J.P., & Massey, C. (2015). “Algorithm aversion: People erroneously avoid algorithms after seeing them err.” Journal of Experimental Psychology: General, 144(1), 114–126. Dietvorst, B.J., Simmons, J.P., & Massey, C. (2018). “Overcoming algorithm aversion: People will use imperfect algorithms if they can (even slightly) modify them.” Management Science, 64(3), 1155–1170.↩
-
[3] Cummings, M.L. (2004). “Automation bias in intelligent time critical decision support systems.” AIAA 1st Intelligent Systems Technical Conference.↩