Guide · 8 min read

How to choose an
AI automation agency

We are an AI agency writing about how to pick one, so read this with that in mind. It is also the honest version, including the questions we would rather clients did not ask us unprepared.

00Why this matters

Every agency can build a demo. Far fewer can ship something you still use next year.

The AI services market rewards confidence. An agency that promises a specific accuracy figure before seeing your data will usually win the pitch against one that says "we would need to test that on your documents first", even though the second answer is the only honest one.

That gap is the whole problem. You are being asked to judge technical delivery capability in a field where the vocabulary changes every six months and everyone's website looks equally credible.

So judge the things that are actually visible from outside: how they scope, how they talk about failure, who does the work, and what you are left holding when it ends.

GUIDEThe playbook

What to look for, ask, and refuse.

01Ask

Six questions that reveal more than a case study

Case studies are marketing, and by definition they only cover the projects that worked. These questions get at delivery reality instead.

  • Tell me about a project that went badly. The answer matters less than the willingness. Anyone who has shipped AI has had a pilot underperform. An agency that cannot name one is either new or not being straight with you.
  • Who will actually do the work. Ask for names and whether you will speak to them directly. The gap between the people in the pitch and the people on the keyboard is where a lot of disappointment lives.
  • What does the first phase deliver, exactly. You want a sentence a non-technical colleague could check: 'in four weeks, invoices arriving by email will be read and entered, with anything unclear flagged to Maria.'
  • How will we know if it is working. A good answer includes measuring the current process before the build starts. If nobody proposes a baseline, nobody intends to prove the result.
  • What happens to my data. Which provider processes it, whether it leaves your region, whether it is retained or used for training, and how access follows your existing permissions. Vagueness here is disqualifying.
  • What do I own at the end. The code, the configuration, the data, and documentation good enough for another team to continue. Get it in writing before the first invoice, not after.
02Avoid

Warning signs worth walking away over

Some of these look like strengths in a sales conversation. That is exactly why they are worth naming.

  • Accuracy promised before seeing your data. Performance depends on your document mix, your edge cases and your rules. A number quoted before a pilot is a sales figure, not an engineering estimate.
  • Agreement with every idea you raise. Part of what you are paying for is being told which of your ideas is not worth building yet. An agency that never pushes back is selling capacity, not judgment.
  • No mention of exceptions or human review. Every serious AI workflow has cases the system should not decide alone. A proposal that only describes the happy path has not been thought through.
  • A large commitment before anything works. Long programmes signed upfront transfer all the risk to you at the point where uncertainty is highest. Insist on a small first phase.
  • Deliberate lock-in. Proprietary wrappers you cannot inspect, no access to your own configuration, hosting only they can operate. Convenience and dependency are not the same thing.
03Compare

How to compare proposals that look nothing alike

Proposals differ mostly in how much vocabulary they contain. Strip that out and put every one on the same three axes.

What exists at the end of phase one

Write each proposal's deliverable in one plain sentence. Anything you cannot reduce to a sentence is not yet a scope, and you will be negotiating it later at your own cost.

The total cost of finding out

Not the programme price. The amount you would spend before you know whether the approach works on your data. Lower is better even at a higher day rate.

The exit if the pilot disappoints

Ask directly what happens if the numbers come back weak. The answer separates partners from vendors, and it costs nothing to ask.

Who carries the integration risk

Most overruns come from the systems around the AI rather than the AI. Check which proposal has actually looked at your stack and which has assumed it will be easy.

04Fit

Specialist agency, generalist partner, or your own team

There is no universal right answer. The question is where the difficulty in your particular project sits.

A specialist agency makes sense when
  • You are unsure what to build, not just how to build it
  • The work depends on designing around what AI gets wrong
  • You need it delivered in weeks and cannot hire in that time
  • Compliance or explainability obligations are in play
Look elsewhere when
  • The hard part is your legacy systems, and an existing partner already knows them
  • You have engineers with capacity and the problem is well understood
  • The requirement is genuinely a product you can buy off the shelf
  • The real blocker is organisational rather than technical
05Contract

The five clauses worth reading twice

None of this needs a lawyer to spot. It just needs someone to actually read the statement of work rather than the summary email.

  • Ownership of code and data. Explicitly yours, including anything derived from your data. Watch for language granting broad rights to reuse your content.
  • A defined scope for phase one. Named deliverables and named exclusions. The exclusions matter more, because that is where change requests come from.
  • Model and provider changes. AI providers deprecate models. The contract should say who is responsible for keeping the system working when that happens, and at whose cost.
  • Support after launch. What is covered, for how long, and what a fix costs afterwards. A warranty period plus optional retainer is a fair shape.
  • Hand-off obligations. Documentation and a walkthrough with your team as a deliverable, not a favour. If it is not in the scope it will not happen.
CTATalk to Brains

Ask us these questions.

Genuinely. The first call is free, and we would rather be judged on how we answer than on how our website looks. If we are not the right fit for what you need, we will say so. How we work or get in touch.

Guide FAQ

Common questions about choosing an agency.

What should I ask an AI automation agency before signing?

Ask them to walk you through a project that went badly and what they changed afterwards. Ask who exactly will do the work and whether you will speak to them. Ask what the first phase costs and what you own at the end of it. The answers to those three tell you more than any case study.

Should I pick a specialist AI agency or my existing software partner?

It depends on where the risk sits. If the hard part is your systems and data, an existing partner who already knows them can be faster. If the hard part is deciding what to build and designing around what AI gets wrong, that judgment is what a specialist is for.

How much should a first AI project cost?

Less than you would be comfortable losing. A well-run agency will propose a small fixed first phase rather than a long programme, precisely because neither side can honestly estimate the whole thing on day one. Treat a large upfront commitment as a warning sign.

What are the warning signs of a bad AI agency?

Guaranteed accuracy figures before seeing your data. Reluctance to name who will do the work. A proposal that never mentions the cases the AI should not decide. No answer on where your data goes or whether it is used for training. And agreeing enthusiastically with every idea you raise.

Who should own the code and the models at the end?

You should own the code, the configuration and the data, and it should be written on stable, widely used technology so someone else could pick it up. Get that in writing. An agency whose value depends on you being unable to leave is not a partner.

How do I compare proposals that look completely different?

Normalise them onto three axes: what is delivered at the end of the first phase, what it costs, and what happens if the pilot shows the approach does not work. Proposals become comparable quickly once you strip out the differing amounts of vocabulary.