The answer was vague because the question was

Nearly every "Copilot is useless" story I hear turns out, on inspection, to be a four-word prompt and a disappointed face. Someone typed "summarise this project" into a chat box, got back something that reads like a press release written by a committee, and concluded the tool is a gimmick.

The tool is genuinely uneven. It has real failure modes and I have written about several of them. But the complaint almost always arrives before the prompt was ever specific enough to test it, and that is a different problem with a different fix.

Everything else in Microsoft 365 works on one click

This instinct is completely reasonable, and it comes from twenty years of training. Every other feature in the ribbon does exactly one thing, immediately, with no explanation required. You click Bold and the text goes bold. You do not have to describe your intent.

Then there is search. Search taught us that three words is the correct amount of effort, and that if the results are bad you rephrase and try again rather than writing an essay. Copilot presents as a text box in the same corner of the same apps, so people bring search-box input and expect assistant-grade output. Nothing in the interface tells you otherwise.

Where the vague prompt actually breaks

No source, so it answers from nowhere.

"Summarise our Q3 performance" does not name a file. Copilot grounds its answer by retrieving what it guesses is relevant from your Microsoft 365 content before the model ever sees your request. A vague request produces weak retrieval, and weak retrieval produces confident prose about nothing in particular. Name the workbook, the thread, or the document and the same question returns something you can actually check.

No expectation, so the shape is wrong.

Ask for "a summary" and you have specified nothing about length, format, or audience. You get five paragraphs when you wanted three bullets to paste into an email, and the complaint becomes "it is too wordy" - which is a formatting instruction you never gave. Format is not something Copilot infers from your mood.

No context, so the register is wrong.

The same set of facts written for your manager, for a client, and for the person who has to action it on Monday are three different pieces of writing. If you do not say which one you want, you get the average of all three, which serves none of them.

The silent grounding failure, which genuinely is Microsoft's fault.

Copilot can only ground its answer in content you already have permission to open. When retrieval comes back empty, it does not announce that it found nothing. It writes a plausible general answer instead. That is a failure dressed as a success, and it is the one case where "Copilot doesn't work" is fair comment - though it still looks identical to a bad prompt from where you are sitting.

The four blocks, used as a diagnostic

Microsoft's own guidance is goal, context, source, expectations. Treated as a ritual it feels like homework. Treated as a diagnostic it is quick: when an answer disappoints, one of those four is missing and it is usually obvious which.

Wandering answer means no goal. Wrong tone means no context. Generic content means no source. Wrong format means no expectations. You do not need all four every time - a specific goal alone gets you most of the way. You need them the moment the first answer is bad, because rephrasing the same vague question louder is not a debugging strategy.

Worked examples of that in practice: One Copilot prompt that summarises a 20-tab workbook shows what naming the source does to the output, and Copilot in Excel: the moment it clicks for a beginner covers the point where the mental model shifts.

Frequently Asked Questions

Why should I have to write a paragraph to get an answer?

You do not. You have to write one specific sentence. "Summarise the risks in the Q3 forecast workbook as five bullets for my manager" is one sentence and it contains all four blocks. The paragraph-length prompts you see in listicles are mostly padding.

So Copilot never actually fails?

It fails constantly. It invents function names, it mangles inconsistently structured tables, and it silently produces general answers when grounding comes up empty. The argument here is narrower: those are diagnosable, specific failures, and you cannot reach them while the prompt is still the weakest part of the setup.

When is "Copilot doesn't work" a fair complaint?

When you gave it a specific goal, named the source, stated the format, and it still returned something wrong or invented. That is a real defect worth reporting. It is also a much rarer event than the volume of complaints suggests.

What is the fastest way to rescue a bad answer?

Do not start over. Reply in the same conversation with the missing block: "use the September file, not the summary" or "three bullets, not prose". Copilot keeps the thread's context, so you are correcting rather than re-asking.

Does this apply to Copilot in Excel and Word too, or just Chat?

All of them. The surfaces differ in what they can reach - the app-embedded versions default to the open document, which quietly supplies the source block for you - but goal, context, source, and expectations is the same spine everywhere.

How do I tell a grounding failure from a bad prompt?

Ask it to cite what it used. If it names files you recognise and the answer is still generic, the prompt was thin. If it cannot name anything, or it names something you would not expect, retrieval failed and the specifics of your prompt were never the issue.

Copilot arriving as a workplace expectation rather than a feature is its own problem, and it is the reason so many people are quietly relieved to conclude the thing does not work.

It probably does. Read your prompt back and ask whether you could have answered it yourself.