Skip to content
Amula AI
AI19 July 20266 min read

Vetting an AI partner: six red flags in regulated finance

Vetting AI consultants comes down to a few specific questions. Six red flags that separate a real partner from a confident pitch in regulated finance.

By Rinor Recica

Choosing an AI partner is one of the few procurement decisions a regulated institution makes where the downside is not a wasted budget line but a governance liability. Anyone who touches a regulated process inherits its requirements — data residency, auditability, documented accountability — so vetting AI consultants matters more here than almost anywhere else. The deciding signal is rarely in the pitch deck; it shows up in the answers to a handful of specific questions. Six red flags separate a real partner from a confident performance, and each has a question that surfaces it.

The first red flag is vague case studies. The consultant talks in outcomes without mechanism — "improved efficiency by 40 percent" — with no baseline, no picture of the process before, no point where something actually changed. The question that surfaces it: "Walk me through one engagement end to end — what the process looked like before, exactly what you changed, and the single number that moved." A real practitioner answers in specifics; a bluffer retreats into adjectives.

The second is the absence of a measurement methodology. The consultant cannot tell you how they would prove the return. For a controller who signs off on numbers that go to investors, that is fatal. The question: "What baseline would you measure on day one, and how would you show me the improvement six months later?" If the answer is a feeling rather than a metric — cycle time, error rate, the number of manual touches — walk.

The third is a claim to serve every industry. Expertise everywhere is expertise nowhere. A partner as confident about a dental practice as about a fund management company understands the constraints of neither. In regulated finance the constraints are the point — data residency, auditability, sign-off — and a generalist who has never met them will discover them at your expense. The question: "What do you deliberately not take on?" A serious partner has a boundary and will name it.

The fourth is silence on failure modes. The consultant only ever describes the happy path and never what happens when the model is wrong. In a FINMA-regulated process the failure design is the design. The question: "When the model produces a wrong answer, what catches it before it reaches an investor, and how do we roll it back?" No answer means no plan — and a plan invented on the spot is its own red flag.

The fifth is an opaque team structure. The consultant is vague about who actually does the delivery and where your data sits while they work on it. In regulated finance that is precisely an outsourcing and data-residency question, not an administrative detail. The question: "Who specifically builds this, and where does our data live while they work on it?" You are entitled to a straight answer before a single file changes hands.

The sixth and deepest red flag: the consultant cannot explain a workflow end to end. Ask them to whiteboard one automation from input to signed-off output. A practitioner reaches for a real example — reporting, reconciliation, KYC extraction — and gets specific about the control points and who owns each one. A bluffer stays at the altitude of "the AI handles it." When reporting across 39+ funds at a leading Zurich investment foundation was automated, the useful conversations were always the specific ones — never the ones that stayed abstract.

Which leaves the signal behind the signals: the best partner tells you what not to automate. Vetting AI consultants comes down, in the end, to finding the one who narrows the scope rather than widening it — who starts with a workflow whose baseline is provable, measures it, and lets the proven case fund the next one. A consultant who answers every question with "yes, we can do that too" has already shown you the most important red flag of all. That is the same discipline that runs through everything we build.

See what your reporting could look like automated.