Skip to main content
The journal / Field notes

If you’re hiring me to agree with your chatbot, save your money

AI can help you challenge a consultant. Its agreement still needs evidence. Why good technical advice should survive questions from both sides.

AIConsultingProduct Development
Cover for If you’re hiring me to agree with your chatbot, save your money

If you’re hiring me to agree with your chatbot, save your money.

I run into a version of this daily in consulting: a founder arrives with a detailed implementation plan, asks for my opinion, then feeds my objections into an AI conversation that has already helped them build the case for the plan. The discussion starts to feel like an audition for whether I’ll approve the answer they already have.

I understand the appeal. AI makes an unfamiliar subject easier to approach. It can help you ask better questions, compare proposals and catch a consultant talking nonsense. Use it for all of that. But if the only acceptable outcome is agreement, you’ve turned a useful tool into a way to avoid the advice you’re paying for.

The problem starts with what you ask it

Imagine asking an assistant to defend a particular approach, explain why it is efficient and help respond to objections. It may do that job very well. That still leaves the approach itself untested.

Now imagine taking a developer’s concern about that plan and pasting in a summary that leaves out the awkward constraint. Perhaps the system must work offline. Perhaps an external service cannot provide the data the plan assumes. The model can produce a coherent answer to the version of the problem it received. You can both come away satisfied while the original engineering problem remains.

The vendor’s reputation doesn’t fill in missing context. A response inside your chat is not a technical review of your repository, contracts, integrations and operating conditions unless those things were actually examined. It should earn your confidence through evidence, just as my advice should.

Agreement is a known failure mode, not an intelligence test

Researchers call excessive agreement with a user’s views sycophancy. In research published in 2023, five assistants showed this behavior across several tasks. The researchers also found that human preferences could reward answers that matched a user’s beliefs, sometimes at the expense of correctness.

That is not evidence that every current model always agrees, or that every disagreement proves a human is right. In an April 2026 analysis of personal-guidance conversations, the model provider reported that most assessed conversations did not show sycophancy, while rates differed substantially by topic. It also reported improvements after further training. Context and model behavior matter.

A Stanford study published in March 2026 found that people receiving overly agreeable advice about interpersonal conflicts became more convinced they were right and preferred those responses. That study concerned personal dilemmas, not software architecture. My concern about consulting is an inference from a similar feedback loop, supported by what I encounter in my own work—not a measured claim that AI makes founders worse at engineering.

It feels good to have an articulate answer in your corner. The useful question is whether the answer survives the parts of the problem that make your preferred solution inconvenient.

Sometimes your AI feature is really a set of rules

In one app I developed, the requirements for its conversational feature kept accumulating instructions to return particular statements when particular questions came up. I’m leaving out the identifying details because the design issue matters more than the client.

There is a legitimate requirement underneath that request: the owner wants control over what the product says. Approved wording can be necessary. The problem is assuming that adding another instruction to a language model is the same thing as implementing a reliable rule—or teaching the model a better understanding of the subject.

A prompt changes the instructions and context used to generate an answer. It does not, by itself, update the model’s learned parameters. If the requirement is that a sentence must appear exactly as written, the application can store that sentence and return it directly once the relevant condition is established. Asking a model to reproduce it adds a step that needs checking.

AI can still be valuable in that design. It might help interpret an open-ended question, retrieve relevant material or handle a conversation that genuinely needs flexible wording. But if a model chooses which fixed reply applies, that choice remains uncertain and needs its own evaluation. Exact stored wording does not guarantee that you selected the right answer.

Consider a hypothetical support assistant with a fixed cancellation policy. Someone asking how to cancel and someone asking whether cancellation already happened may share most of the same words. A growing pile of keyword-like instructions can obscure that distinction. Define the intended behavior, test ambiguous cases and decide when the system should ask a clarifying question or hand the conversation over.

I am happy to build a controlled response system. I am happy to build an AI conversation. I want us to agree on which parts require each, because that decision changes what we implement and how we judge whether it works.

Make the disagreement testable

You should question a consultant who dismisses your plan because AI helped write it. Good ideas do not become bad because of where they came from. Equally, a polished explanation is not a reason to skip review.

Give the reviewer the actual goal, constraints and proposal. Ask them to identify a specific failure, explain when it would occur and propose a proportionate way to test it. If I say your approach is unreliable, I should be able to do better than pointing at my years of experience.

Use AI to improve that conversation. Give it the objection with its context intact. Ask which assumptions the two approaches depend on, what evidence is missing and what result would change the recommendation. Ask it to make the strongest case against my suggestion too. A fresh conversation may reduce the influence of an earlier framing, but it is not a guarantee of independence or correctness.

Then run the relevant check. That might be a small prototype, a supplier documentation check, a permissions test or a set of representative questions with agreed expected behavior. Choose the evidence that answers the disputed claim. A second chatbot agreeing with the first is not a substitute for it.

You are allowed to choose the trade-off

Sometimes a client understands a limitation and accepts it to meet a budget or deadline. That is a business decision. My job is to explain the consequence clearly, offer sensible alternatives and document what we have agreed—not to win every argument.

But paying someone for judgment and then treating every disagreement as incompetence wastes the part of their work you most need. You can get enthusiastic agreement cheaply. Finding a problem before it becomes expensive is the service worth buying.

That is the purpose of Rosecraft’s technical consulting: examine the proposal, explain the trade-offs and help you decide what to build. Bring the AI-generated plan. Bring the questions. Just leave room for an answer neither of us expected.

Ready to hear what you need to hear—not just what you want? Drop us a message at rosecraft.studio. Tell us what you’re building and which decision you want challenged.

Keep the conversation going

Share this article

Corey Rosamond, Founder and Principal Engineer of Rosecraft Studios

Corey Rosamond

Founder & Principal Engineer

Learn more
Occasional notes

Stay in the loop.

Get notified when we publish new insights on web development and engineering.