← Alle Beiträge

Choosing an AI agency or AI partner: the questions that expose a weak vendor

Choose an AI agency by what it can show on your own task, not by its demo. Before you sign, ask ten questions: which model and why, how they test, what an agent that does nothing scores on those tests, what the agent may do without a person, where your data goes, what one finished task costs, whether that cost was measured on Greek, how customers will know they are talking to an AI, what happens when something fails at night, and who owns the result. A strong vendor answers each one with a number or a document. A weak one answers with the name of a framework.

Why is choosing an AI partner harder than choosing a web agency?

Because the demo proves almost nothing. A website either loads or it does not, and you can see it. An AI system that works in a meeting can fail quietly on real data, and the failure shows up weeks later as a wrong answer to a customer or a wrong number in a report. The scores vendors quote are the second trap. In my own cost study, an agent that did nothing scored 22.8% on a real benchmark's graders, because some checks rewarded not doing things (what an agent costs). If a vendor cannot tell you what doing nothing scores on their own tests, their accuracy figure has no floor under it.

What should you ask an AI agency before you sign?

  1. Which model do you use, and what happens when the provider retires it? Providers retire models on their own schedule, so the contract needs a clause for it, as covered in who pays when AI gets it wrong.
  2. What does your test set look like, and what does an agent that does nothing score on it? A test set with no baseline cannot tell improvement from luck. How a proper one is built is in LLM evals.
  3. Can I see the results on twenty of my own cases before I commit? Twenty real cases say more than any case study, and the failures are the useful part. What a pilot has to prove is in from AI pilot to production.
  4. What may the agent do on its own, and what needs a person's approval? The answer should be a list of actions, not a promise, as described in human in the loop.
  5. Where does my data go, and is it used for training? Ask for the provider's terms and a processing agreement. OpenAI, for example, states that data sent through its API is not used for training by default (OpenAI, your data). The Greek angle is in GDPR and AI.
  6. What does one finished task cost, including retries and review time? A price per month without a volume and a task definition cannot be compared with anything, as explained in what an AI agent costs.
  7. Was that cost measured on Greek text? On OpenAI's current vocabulary Greek needs 2.06 times the tokens of English, so an estimate made in English is roughly half the real bill (Greek in AI models).
  8. How will customers know they are talking to an AI? Article 50 of the EU AI Act requires it, at the latest at the first interaction (AI Act, Article 50).
  9. What happens when an API times out at 3am? Retries, alerts and a safe stop are designed in, or they are missing. What that takes is in what an AI agent actually does at 3am.
  10. Who owns the prompts, the code and the test set at the end? The test set is the part most often forgotten, and it is the part that lets another team take over.

Which answers should worry you?

  • "Our framework makes it accurate." In my study of six agent harnesses on four models, moving between models explained 88.0% of the spread in behaviour scores and the harness design explained 6.9% (how much of an agent is the model). A framework is a way of working, not evidence. Results on your own cases are evidence.
  • "It is 95% accurate." Ask on which questions, measured when, against what baseline. A percentage without a test set is a slogan.
  • "The agent can do anything a person can." In the same study, 16 sessions deleted files nobody had asked them to touch, all of them running on capable hosted models. Limits belong in the design, as described in AI agent permissions.
  • "First we fine-tune a model for you." Most projects need clearer instructions and better retrieval long before they need training, as argued in RAG, fine-tuning or a better prompt.
  • No questions back. A vendor who quotes before asking about your volume, your data and what a mistake costs you is pricing a demo, not your project.

What should the contract say?

Four things protect you. Acceptance is measured on an agreed test set built from your own cases, with the target written as a number. The supplier may change models only while that number still holds. Your data is covered by a processing agreement that names every provider it passes through. And the prompts, code and test set are delivered to you, so the system can outlive the relationship. The liability side, including what happens when an answer is wrong, is covered in who pays when AI gets it wrong.

How does a good engagement start?

Small and measured. Every engagement I take starts with a discovery call, a written blueprint that prices the work, and a pilot on the client's own data before anything larger is signed. That order is what lets both sides commit to a number: nobody can promise accuracy on data they have never seen. The systems themselves are described under AI agents for business, and several run as working demos on synthetic data that you can open before a first call.

Frequently asked questions

Should I choose a large consultancy or a specialist?

Choose whoever answers the ten questions above with numbers for your task. Size buys continuity and process; a specialist usually buys speed and a direct line to the person who builds. Either can be right, and either can fail the questions.

How long should a pilot take?

Long enough to run your twenty cases, read every failure and fix the ones that matter. For one narrow, well-defined task that is usually a matter of weeks, not months. A pilot that has no end date and no target number is not a pilot.

Is a cheaper vendor always riskier?

No. The risk sits in the questions left unanswered, not in the price. A cheap proposal that states the model, the test set, the baseline and the limits is safer than an expensive one that states none of them.

If you are comparing AI proposals right now, send me the proposals without confidential data, and I will tell you which of these questions they leave open.