Does AI Actually Work? The Questions We Get Asked Every Week
Plain answers to the eleven questions people actually ask us about AI — what it's good at, what it's bad at, whether it's coming for your job, whether it's safe to paste company information into it, and whether waiting a year helps.
Why this exists
We get asked about AI constantly. At dinner, on client calls, by people with no intention of hiring us. The questions are remarkably consistent, and they’re almost never the ones the industry writes about.
So here they are, with straight answers, roughly in the order people ask them.
If you’re about to sign a proposal for an AI project and want the longer version — how to tell whether the project will work, and how to buy one without getting taken — that’s a different piece. This one is for everybody else.
One disclosure, because it shapes the answers: we build software, and some of it involves AI. We also talk people out of AI projects fairly often, which is a worse way to run a business and the reason we can write this honestly.
1. So what do you actually think of AI?
That it’s a useful tool wrapped in an enormous amount of noise, and that most arguments about it are arguments about the noise.
“Is AI good?” is a bit like asking whether spreadsheets are good. It’s a category of tool. It does some things well, one thing badly in an unusual way, and it’s being sold far harder than it’s being explained. The two popular positions — that it changes everything, that it’s a parlour trick — are both describing the sales pitch rather than the thing.
Our actual view: it’s genuinely useful, and it has one dangerous habit — when it’s wrong, it looks exactly the same as when it’s right. Everything else on this page follows from that.
2. Does it actually work?
Yes, for a narrower set of jobs than you’re being sold.
The best single predictor is whether anyone can tell a good answer from a bad one. If a person who knows the subject reads it, or a test runs against it, or the numbers have to add up — then AI is genuinely useful and often much faster than doing it yourself. If the answer goes straight into a decision with nobody checking, it’s a liability, and the fact that it sounds convincing is not evidence of anything.
That’s the whole rule. If you only take one thing off this page, take that one.
3. What is it genuinely good at?
From using it every day, not from a sales deck:
- Reworking writing you already have. Shortening it, restructuring it, turning notes into paragraphs or paragraphs into a table. You have the original, so you can see immediately whether it got it right.
- Rough first drafts. The version you were dreading starting. Fixing something mediocre is easier than facing a blank page.
- Explaining things you don’t understand yet. A contract clause, an error message, a standard nobody has read since 2011. It’s a patient explainer that never makes you feel stupid for asking twice.
- Generating options. Twenty names, ten angles, five ways to approach something. You’re going to throw out nineteen of them, and throwing out nineteen bad ideas is easy work.
What those have in common: someone who knows the subject can spot a bad result in seconds. That’s not a flaw in the technology. It’s the condition under which it’s worth using.
4. What is it bad at?
- Knowing when it doesn’t know. It can’t reliably tell the difference between having the facts and being able to produce something that looks like the facts.
- Arithmetic, and counting carefully through a long document where one detail in the middle matters. It gets a lot of this right, which is exactly what makes the misses expensive.
- Anything specific to your business. Your prices, your customers, your stock, your history. It knows none of it unless somebody has connected it up — and that connecting-up is ordinary, unglamorous software work. It’s most of the actual effort in a typical “AI project,” and it looks nothing like AI.
- Being consistent. Ask it the same thing twice and you can get two different answers. Fine for a draft email. Not fine when another system has to read the result.
- Anything where being slightly wrong looks identical to being right. This is the one that matters, and it’s why every good use above has someone checking.
5. Why does it sound so sure of itself when it’s wrong?
Because it’s built to produce writing that reads well, and writing that reads well and writing that’s true overlap a great deal — but not completely.
There’s no separate part of it keeping track of how confident it should be. A wrong answer arrives sounding exactly like a right one. Something that hedged honestly would be far more useful and far less impressive, and nobody has really solved that.
6. Is it going to take my job?
Our honest read: it changes what jobs involve considerably faster than it removes them, and it isn’t hitting everything evenly.
The work most at risk is work that mainly consists of producing documents nobody reads carefully. If that describes a task, it was already fragile, and this isn’t the first thing to threaten it.
The work least at risk is work where somebody has to be answerable for a decision. Not “write the report” but “sign the report.” Judgement when the answer isn’t clear, responsibility when it goes wrong, relationships, and knowing which of twenty suggestions survives contact with an actual customer.
The pattern in our own field: the people getting real value out of it are, without exception, the ones who could have done the work themselves and are using it to go faster. It amplifies whatever you bring. If you bring nothing, it amplifies that too.
Where we’d be genuinely worried: if your plan involves spending three years doing the junior version of a job before you’re trusted with the real one, that ladder is being pulled up faster than anyone is discussing. That’s a real problem and we don’t have a good answer to it.
7. Is it safe to put company information into these tools?
The most-asked question, and the answer is annoyingly conditional: it depends on which product and which plan. The defaults on a free personal account are not the defaults on a paid business one, even from the same company.
What we’d tell any owner:
- Read the terms for the exact product and plan your people are using. Not the company’s general privacy page — the terms attached to that plan. Two things to look for: how long they keep what you type, and whether they may use it to improve their own product. Those answers differ between a free account and a paid business agreement.
- Approve one tool and say so out loud. If there’s no approved option, people use whatever’s on their phone with a personal account, and you’ll never know it happened. One tool you’ve actually read the terms for beats a policy nobody follows.
- Assume anything typed into a personal free account is out of your hands. Not because anyone’s acting in bad faith — because you have no agreement, no record, and no way to answer a question about it later.
- If you’re in a regulated business, this is a compliance question, not an IT preference. Client information, health information, anything you’re under a duty to keep confidential — that gets a written decision with somebody’s name on it. If your compliance officer hasn’t been asked, they’re going to be asked at a much worse moment.
None of this means don’t use it. It means the answer to “where does this go” should exist before three hundred client records go there.
8. Which one should we use?
For most business work the main ones are close enough that the choice matters far less than how you use them. People agonise over this and it’s rarely where the value is.
Decide on the things that genuinely differ:
- Where your information goes — see the previous answer.
- Whether it connects to what you already own. A decent tool wired into your documents and your email will beat a slightly better one sitting in a separate browser tab, every time, because the second one stops getting opened by Thursday.
- Price at the volume you’ll actually use, which for most small teams is not the deciding factor people expect.
The published comparison scores are close to useless for this. They measure performance on standard test sets that aren’t your work, and the gaps between the leaders are smaller than the gap between using it well and using it badly.
9. Do we need to learn prompt engineering?
Mostly no, and the phrase has aged badly.
There’s no secret wording. What actually helps is being able to describe a job clearly, give it the background it doesn’t have, and check what comes back — which is the same skill as briefing a capable new contractor who has never seen your business. People who are good at handing work to other people turn out to be good at this almost immediately.
The useful habits are dull ones: say what the finished thing should look like, show an example, include the background rather than assuming it’s known, and ask it to show its working when you plan to check it. That’s the whole course.
10. Should we wait? It’ll be better in a year.
It will be better in a year. Waiting still doesn’t help, for reasons that have nothing to do with the technology.
The work that makes AI useful inside a business is this: tidying up the information it would need to read, writing down how a job actually gets done rather than how the manual says it does, and deciding who’s responsible for checking the output. That work is identical no matter which product you eventually use, and it takes months. Do it now and you’re ready when the tools improve. Wait, and you start the same clock later, having gained nothing.
What we would genuinely wait on is signing a large, long commitment to one company’s platform for a problem you haven’t defined yet. That’s not caution about AI. That’s just how you’re supposed to buy things.
11. We tried it and it didn’t work. Was that us or the tool?
Usually neither — usually the job you gave it. Three questions, in order:
Could anyone tell a good answer from a bad one? If not, it was never going to work, whatever you’d used.
Was the information any good? This is the most common cause by a distance and the least welcome answer. Point AI at a system where the same customer exists four times under three spellings and you’ll get confident nonsense back. That’s not the tool’s fault. Most AI projects that quietly die, die here.
Was the job small enough to describe? “Summarise this document” works. “Handle our support inbox” is a hundred different jobs with no edges, and you find out it went wrong when a customer tells you.
If all three are fine and it still didn’t work, then it was the tool or the way it was built, and that’s worth another look. In our experience it’s that third case maybe one time in five.
The short version
It works, for fewer things than you’re being told, and the limits are predictable rather than mysterious: it works where somebody or something checks the answer.
It isn’t coming for the jobs where a person has to be answerable for something. It is coming for the tasks that were already producing paperwork nobody read.
And the most useful thing you can do this quarter isn’t picking a vendor. It’s finding out what state your own information is actually in — because whatever you eventually build will inherit every problem down there, and you’d rather find them now than in front of a customer.