← All posts

How to Choose an AI Development Partner: 10 Questions That Separate Builders From Demo Artists

A founder's checklist for vetting AI development agencies and freelancers — the questions that expose demo-ware, and the answers a real production team gives.

The AI gold rush created a strange market: thousands of “AI development” vendors, most of whom have shipped impressive demos and almost nothing that survives contact with real users.

We’ve inherited enough half-built projects to know the pattern. Here are the questions that expose it early — and what good answers sound like.

The 10 questions

1. “Show me something in production. Who uses it daily?” Demos prove interest in AI. Production proves engineering. Ask specifically about systems that real users depend on — like a support assistant handling 1,000+ conversations a week.

2. “What happens when your AI doesn’t know the answer?” The single best filter. A demo artist says “it always answers.” A production builder talks about grounding, confidence thresholds, and designed “I don’t know” behavior — because that’s what compliance teams sign off on.

3. “How do you prevent hallucinations?” Wrong answer: “we use a good prompt.” Right answer: architecture — retrieval-augmented generation, source-grounded responses, guardrails that block unsupported claims.

4. “What’s your escalation path to humans?” Any AI touching customers or money needs a designed human lane. If the vendor hasn’t thought about it, they haven’t shipped to real users.

5. “What did you deliberately NOT automate on your last project?” Real builders have judgment about where AI moves the needle and where it doesn’t. Someone who automates everything indiscriminately is selling hours, not outcomes.

6. “How will we measure whether this worked?” Expect numbers: deflection rate, response time, processing time, hours saved. Our lead-qualification client measures in minutes instead of hours. “Improved efficiency” without a metric is a red flag.

7. “What’s your stack, and why?” You’re not testing the stack — you’re testing the why. “LangChain because we know it deeply” is fine. Buzzword salad without reasons is not.

8. “What happens after launch?” Models drift, docs change, edge cases surface. Ask about monitoring, evaluation, and knowledge updates. “We hand it over and leave” means you own a system nobody understands.

9. “What will this cost to RUN, not just build?” LLM API costs at scale are real. A production team estimates per-conversation or per-task costs up front and designs around them (caching, model routing, small-model fallbacks).

10. “Tell me about a project where AI was the wrong answer.” Anyone who can’t name one is either inexperienced or selling. Some processes need a plain workflow, a database query, or a better form — not a model.

The pattern behind the questions

Every question probes the same thing: has this team operated AI systems after the launch party? Building the happy path is easy now. The craft lives in the unhappy paths — wrong answers, edge cases, drift, cost curves.

What working with a production-first team looks like

Every engagement we run starts with clarity: what’s the problem, what does success look like, and what’s the fastest path there. No fluff, no scope bloat, no surprises. Sometimes the first deliverable is a one-page architecture doc that says “you don’t need an agent for this — here’s the simpler thing.”

That’s the bar. Hold every vendor to it, including us.


Vetting partners for an AI project right now? Send us the brief — worst case, you leave the call with sharper questions for whoever you hire.