Skip to content
Kulicki.tech / Blog
Back

The Real Problem With AI: It Tells You What You Want to Hear

Why agreeable AI responses can make weak ideas feel validated, plus practical prompts that force useful pushback before important decisions.

When was the last time an AI challenged the core assumption behind your idea?

If you can’t remember, it’s not because all of your ideas have been flawless. It’s because the tool is built to be agreeable, and we rarely push hard enough to find out what it actually thinks. Wrong facts and strange errors are easy to catch. The bigger problem is that AI can make a weak idea feel fully reasonable when it isn’t.

This text is for people who use AI daily to brainstorm ideas, review plans, write strategy, or validate decisions. If you already know how to prompt a model, this will show you why most of your prompts are not doing what you think they are doing.

Agreement isn’t the same as proof

When an AI says an approach makes sense, what has it actually confirmed? Usually, almost nothing.

It doesn’t mean customers care. It doesn’t mean the market is ready. It doesn’t mean the product survives contact with real users. Often, it simply means the model’s conclusion follows logic you framed. In other words, the model is generating the statistically most probable continuation based on the context you provided, not independently verifying whether your assumptions are true.

That’s easy to miss because the answer sounds coherent and well-structured. But a perfectly built argument on top of weak assumptions is still standing on shaky foundations.

Let’s say you write:

We’ve noticed that mid-sized retailers struggle to manage product data across sales channels. We want to build an AI platform to fix that. Help me strengthen the business case.

You’ve already decided the problem, the audience, and the solution. The model will build on that foundation, not question whether it should exist at first place.

This isn’t necessarily a bad thing. In fact, it’s often exactly what you want when you’re refining an idea rather than questioning it. The important part is recognizing the limitation, so you don’t end up confidently walking deeper into a maze built on faulty assumptions.

An AI-generated argument built on unverified assumptions

You can steer a model until it agrees

A conversation with AI can slide from exploration into negotiation. The model pushes back, you add context. Still unconvinced, you rephrase. Eventually the answer turns positive.

Sometimes that means the model finally understood. Other times you just narrowed the conversation until disagreement stopped being likely.

People do this to each other too: we highlight facts that support us and find people who get our vision. AI does it faster, because it’s endlessly patient and works with whatever framing you give it last.

Research on sycophancy in language models backs this up: AI assistants often favor answers that match a person’s existing views over more accurate ones, and people often prefer confident, well-written answers over correct ones. The model’s behavior is only half the story. We reward it for telling us what we want to hear.

That builds a comfortable loop: you share assumptions, the model builds a case around them, the strength of the answer boosts your confidence, you come back more certain, the model matches that certainty. Eventually the conversation stops testing your idea and starts dressing it up for a party it was never invited to.

Very simple example

Is the AI lying to you on purpose?

That’s usually the wrong question to ask. A model doesn’t need to plot against you to mislead you. A system trained to sound useful and win approval can become too accommodating without any intention at all.

Researchers define AI deception as systematically creating false beliefs to reach some goal other than the truth. That’s a real and serious area of study, but most of what I’m describing is more subtle. The model simply produces an answer that follows the direction of the conversation.

Ask someone how they’re doing and the likely answer is fine, but that doesn’t make it true. Models are very good at producing the expected script (they’ve read a lot of scripts). The useful answer is sometimes the one that breaks it.

How is going

Your confidence goes up, but accuracy goes down

A 2026 preprint based on five experiments with 3,132 participants, also covered by TNW and The Register, tested difficult questions chosen specifically because the AI gave the wrong advice.

Without AI, 44% of participants admitted they did not know correct answer and 27% answered correctly. With AI advice available, only 3% admitted uncertainty, however accuracy fell to 9%, and confidence jumped from 30% to 76%. Accuracy dropped by two-thirds while confidence more than doubled. Some participants who would have answered correctly on their own changed their answer after consulting the model and became wrong.

The AI did not merely give them a bad answer. It replaced uncertainty with borrowed confidence. This is exactly the point I am trying to make. You need to be especially cautious about that pattern. Fluent response can make a weak assumption stop feeling like something to test and start feeling like something you know.

Add friction on purpose

The default use of AI: here’s my idea, help me improve it. Fine for execution, but weak for testing whether the idea even deserves to exist.

For decisions with real stakes, I split the job in two. First, attack the premise. Only then, improve whatever survives. One of my favourite prompts to make the LLM fight the yes-man instinct is:

Adopt the role of a Meta-Cognitive Reasoning Expert.

For every complex problem:

1. DECOMPOSE: Break into sub-problems
2. SOLVE: Address each with explicit confidence (0.0-1.0)
3. VERIFY: Check logic, facts, completeness, bias
4. SYNTHESIZE: Combine using weighted confidence
5. REFLECT: If confidence <0.8, identify weakness and retry

For simple questions, skip to direct answer.

Always output:
- Clear answer
- Confidence level
- Key caveats

The confidence score isn’t a scientific measurement, a model can be confidently wrong about how confident it should be. Still, the structure asks for weak points instead of a polished conclusion, and it changes how I read the output: not a verdict, rather an argument I still need to check.

Try this prompt yourself and see how magic happens.

Getting out of the yes-man loop

A few adjustments change how useful your AI conversations actually are:

  • What would have to be true for this answer to be wrong?
  • Which part of my prompt are you accepting without checking it?
  • Give me the strongest case against this idea.
  • Separate what’s verified, what’s a reasonable guess, and what’s just assumed.
  • What am I not telling you that could change your recommendation?
  • How would a skeptical customer, investor, engineer, or competitor respond?
  • Which part of your answer sounds convincing but has the weakest support?
  • If this plan fails in six months, what probably caused it?

That last habit is the one people skip most often, and it is the one that pays off the most. The first answer is usually the smooth one. The second or third pass is where the friction starts, and friction is where you find out if the plan actually holds.

Let it make you uncomfortable

I don’t need AI to challenge every word. I need it to push back when a bad assumption could send the whole project in the wrong direction.

Building friction into every AI conversation from scratch is slow. I put together a package of 15 prompts designed to make AI challenge your assumptions instead of confirming them, covering plan reviews, strategy checks, product decisions, and technical calls.

Grab it, drop the prompts into your next AI conversation, and see how much of your current plan survives the pushback.