Ask.Sort.Check.
Blog · 5 September 2026 · 7 min read

AI Answers Questions — The Trick Is Knowing Which Ones to Ask

AI answers questions whether or not it knows the answer. A practical sort of the four question types, which ones it gets right, and the test to apply before you ask.

The interesting question was never whether AI answers questions. It answers all of them. Type anything into a chatbot and something comes back, formatted, calm and complete. The interesting question is which of those answers you can act on.

People tend to sort this by subject — good at writing, bad at maths, and so on. That sorting does not work very well. The same tool that nails one marketing question invents a statistic on the next one. The split that actually predicts reliability is not the subject. It is where the answer had to come from.

Why it answers even when it should not

These systems have no step where they decide whether they know something. Producing a plausible answer is the entire operation, which is the short version of how AI actually works.

There is also a reason the behaviour has been slow to improve. A 2025 paper from OpenAI and Georgia Tech, Why Language Models Hallucinate, argues that the way these models are graded rewards guessing: a benchmark that scores a wrong answer the same as "I do not know" makes confident guessing the winning strategy. The models are behaving exactly as their scoring taught them to. Saying nothing is penalised, so nothing is what you never get.

Four kinds of question

Sort your question into one of these before you ask it. The sorting takes a couple of seconds and it tells you how much checking the answer needs.

1. The answer is in the message you sent

Summarise this. Extract the dates. Rewrite this shorter. Find the inconsistency between these two paragraphs. Turn these notes into an agenda. Here the model is not retrieving anything — it is transforming text you supplied, and the source of truth is sitting right there in the conversation.

This is the most reliable category by a distance, and it is where most of the genuine time savings live. It is also the least glamorous, which is why people skip past it looking for something more impressive.

2. Stable, widely documented general knowledge

How does compound interest work. What is the difference between gross and net margin. Explain what an API is. Draft a standard non-disclosure clause. These are things written about thousands of times in consistent terms, so the plausible answer and the correct answer converge.

Usually right, and useful for getting oriented in an unfamiliar area. The failure mode is specifics: the general explanation is sound and the particular figure, date or citation inside it is invented. Take the shape of the answer, verify any hard detail you plan to quote.

3. Recent, fast-moving or contested

This is where it goes wrong most often, and there is now good evidence for how often. In October 2025 the European Broadcasting Union and the BBC published News Integrity in AI Assistants, in which journalists at 22 public media organisations across 18 countries assessed more than 3,000 answers from four major assistants. Almost half of the answers had at least one significant issue, around a third had serious sourcing problems, and roughly a fifth contained major accuracy errors including fabricated or outdated detail.

That held across every language, territory and platform tested, so it is not a quirk of one product. The pattern in the failures is worth internalising: the assistants struggled most with stories that were still moving, timelines with several actors, and questions where fact and opinion needed separating. If your question has any of those properties, treat what comes back as a lead to check rather than an answer.

4. Anything specific to you

What does our refund policy say. How many of our customers churned last quarter. Is this clause standard in our contracts. The model has never seen your policy, your numbers or your contracts, and — per the paper above — it will not stop to mention that. It will produce a fluent, reasonable, entirely fictional version.

The fix is not a better question. It is pasting the actual document in, which converts a category-four question into a category-one question. That single move is most of what writing a prompt that works first try is about, and it is also why the "pre-trained" in what GPT stands for matters day to day.

Anyone working with figures runs into this weekly, which is why using AI as an accountant keeps your own numbers on the paste-it-in side of the line rather than the ask-it side.

The test, in one line

If you provided it, you are on solid ground. If the answer would have to come from the open internet, expect the general shape to be right and the specifics to need checking. If it would have to come from inside your organisation, it is going to be invented.

The same subject, four ways

Say you are preparing for a supplier renegotiation. Four questions, one subject, four different levels of trust.

  • "Here is the current contract. Summarise the termination and price-review clauses." Category one. Reliable, and you can confirm it against the document in a minute.
  • "What are the standard levers in a supplier renegotiation?" Category two. A sound checklist to work from, no specific claim worth quoting.
  • "What has happened to freight rates on this route this quarter?" Category three. Do not use the number. Use it to know what to look up.
  • "What did we pay this supplier last year?" Category four. It will answer. The answer is fiction unless you paste the invoices.

Nothing about the tool changed between those four. Only the location of the truth did.

What about the ones that search?

Assistants that search the web before answering genuinely improve category three, because there is now a link to open. But the EBU study covered tools that search, and sourcing was still the weakest area — a citation can be attached to a claim the source does not actually make.

So the habit holds: if there is a link, open it and confirm the source says what the answer claims. That is a two-minute job and it is doable outside your own field, which is the whole point of checking an AI answer when you are not the expert.

Where this leaves you

None of this is an argument against using the tools. It is an argument for aiming them. The people getting real value are not asking better questions in some mysterious sense — they are asking more category-one questions, because they have got into the habit of supplying the source material with the request.

The quickest way to find those in your own week is to look for work that is repetitive and text-shaped, which is the exercise in finding the repetitive part of your job. It is also worth knowing the wider set of things these tools handle badly, laid out in what AI is actually bad at, so that a bad answer is something you recognised rather than something that caught you.

Coursium teaches this as a habit rather than a fact — short lessons, a quiz that checks it stuck, and a practice task so you have actually sorted a few questions yourself. Stay ahead of AI by learning the tools on your phone.

Coursium

Stay ahead of AI — learn the tools on your phone.

Get the app