AI Partner Selection Is a Capability Decision, Not a Procurement Decision

AI Consulting vs Agencies vs Tool Vendors - Kamyar Shah, Fractional COO

AI consulting firms, agencies, and tool vendors are not three prices for one product. They are three products, and each assumes a different level of operating capability in the buyer. Consultants sell judgment, agencies sell scoped execution, and vendors sell a platform the buyer runs. Choosing among them is a diagnosis, not a comparison of quotes.

The Opening Question Most Buyers Ask Cannot Separate the Options

Evaluations usually begin with cost and delivery timeline. Any of the three models can produce an answer to that question, which is exactly why the answer fails to distinguish them. The variable that separates the models is what the organization can operate on its own once the engagement closes. Define that gap before requesting a single proposal.

A capability gap is written as a sentence about the company, not about the technology. It names what the team cannot presently do: define the problem, build the solution, or run the system afterward. Each of those three gaps points at a different partner. Write the sentence first and the shortlist assembles itself.

Three Billing Models Produce Three Different Products

A consulting firm bills judgment against an open scope, so it can redesign a process it was not originally hired to touch. An agency bills a fixed scope against a fixed fee, so it delivers what the statement of work names and stops there. A tool vendor bills a subscription, so the product roadmap sets the boundary of what is possible.

None of those constraints is a defect. Each one is the mechanism that makes the model economically viable, which is why no amount of negotiation moves it. An agency asked to behave like a consultancy will either refuse or reprice into one. Read the billing model as a description of the product and the boundaries become predictable.

The Anti-Pattern Is Buying the Familiar Shape

Companies default to the procurement pattern they already know. A firm that has always bought software buys a platform. A firm that has always hired agencies writes a project scope. The pattern is comfortable, and comfort is a poor proxy for fit.

This is the most common failure in the category, and it is a governance failure rather than a technical one. The evaluation committee is usually drawn from functions that have bought one shape repeatedly. Nobody in the room is incentivized to question the shape itself. Ask which model the organization is prepared to operate, not which model it has bought before.

Composure Is a Financial Control in This Category

The pressure to move quickly on AI produces compressed evaluations and reversed sequencing. Speed applied to the wrong partner model does not shorten the timeline. It relocates the delay to the maintenance phase, where it costs more and attracts less attention. Slow the selection and the deployment accelerates.

Calm here is a procedural commitment rather than a temperament. It means holding the evaluation open long enough to measure internal readiness, and declining to read a competitor announcement as information about the organization’s own constraints. Diagnosis precedes purchase in every other capital decision. Apply the same discipline to this one.

Absorptive Capacity Sets the Ceiling on What Any Partner Can Deliver

Cohen and Levinthal described absorptive capacity as an organization’s ability to recognize, assimilate, and apply external knowledge. Applied here, it is the practical question of whether a team can take a delivered system and improve it. A platform handed to a team with no prior exposure to data pipelines produces an expensive dashboard nobody adjusts.

Absorptive capacity is built through prior related work, which means it cannot be purchased in the same transaction that requires it. An organization with documented processes and an analyst who already builds reports has some. One where reporting is manual and undocumented has very little. Measure it first, because it caps the value of every option below it.

Transaction Cost Economics Names the Dependency Buyers Feel Later

Williamson framed make-or-buy decisions around the cost of governing a relationship rather than the price of the good itself. Every AI partner model carries a governance cost that appears after signature: renegotiation, coordination, and the effort of supervising work the buyer cannot independently evaluate. That cost rarely appears in any proposal.

Consulting relationships carry high governance cost and high adaptability, because an open scope requires active management and permits redirection. Platform relationships carry low governance cost and low adaptability. Agencies sit between the two on both dimensions. Choose the governance burden the organization can actually carry.

Asset Specificity Is the Variable Nobody Scores

Asset specificity measures how much of an investment loses value outside the relationship that created it. A custom workflow built on a vendor’s proprietary integration layer is highly specific, and that specificity is what converts a subscription into a dependency. The dependency is invisible at signature and expensive at renewal.

The rule that follows is straightforward: highly specific assets belong inside the organization, and generic ones can sit with a partner. Data models, decision logic, and process documentation should be owned. Compute, hosting, and commodity tooling need not be. Score specificity on the evaluation sheet next to price.

Total Cost of Ownership Runs Longer Than the Contract

The three-year cost of an AI deployment includes migration, retraining, and the redesign of every workflow the system touches. A platform priced below a consulting engagement can exceed it once the organization hires a second partner to make that platform operational. Anything shorter than three years measures the purchase rather than the ownership. Compare all three models on identical assumptions.

Break-even then follows capability rather than price. A system that runs parallel to operations and requires manual handoffs never compounds, regardless of the invoice. Organizations that match the model to internal capability generally reach payback inside twelve to eighteen months, and mismatched pairings extend that window past two years. Match the model to the capability and the timeline takes care of itself.

Score Partners on What They Leave Behind

Standard scorecards measure technical capability, price, and delivery record. None of those measures systems-building capacity, which is the ability to leave documentation, training, and process architecture behind. That single omission explains most disappointing engagements, because a partner scored only on delivery has no reason to invest in the handoff. Add the missing dimension to the sheet.

A partner who cannot show prior work product is a supplier of effort rather than a builder of capability. Ask for a redacted process document or a handoff plan from an earlier client. The request is reasonable, and the response is informative either way. Weight documentation as heavily as technical skill.

Structure the Pilot to Test the Partner, Not the Technology

A sixty to ninety day pilot with a defined kill decision is a low-cost option on a large commitment. The technology is rarely what fails, so read the pilot as evidence about the partner. Watch the clarity of communication, the discipline of documentation, and the willingness to train internal staff. Those signals do not appear in a proposal.

Watch specifically whether knowledge moves toward the internal team or accumulates with the vendor. Organizations that score the pilot this way consistently report that the partner question settles itself before the technical one does. Decide on that signal rather than on pilot output alone.

Tie Payment to Operational Outcomes Rather Than Deliverables

Milestone payments attached to documents reward document production. Payments attached to measured operating results reward the change the organization actually bought. A strategy deck is not an outcome, and neither is delivered code that no internal person can modify.

Writing the contract this way requires naming the operating metric before the engagement begins, which is itself a useful discipline. If the metric cannot be named, the problem is not defined well enough to buy a solution for it. Specify the metric, then specify the payment.

Knowledge Transfer Is a Deliverable, Not a Courtesy

Process documentation, training material, and a staged handoff plan belong in the statement of work with dates and named owners. Without that clause, tacit knowledge stays with the partner and the organization rents its own operations indefinitely. Goodwill is not a substitute for a contractual term.

The engagement ends when the internal team can run and improve the system, not when the system goes live. Companies that write a dated handoff into the agreement describe renewal conversations as negotiations between aligned parties rather than rescues. Name the handoff date and attach the final payment to it.

Which Model Fits Which Gap

If the problem is undefined and the sequencing is unclear, engage a consultant. When the problem is defined, the scope is stable, and execution bandwidth is the only constraint, engage an agency. Where technical staff and documented processes already exist and only the platform is missing, buy the tool and staff it internally.

Two of these conditions are often true at once, and that is not a reason to blend the models. Sequence them instead, closing the diagnostic gap before the execution gap. Buying all three at once produces three vendors and no ownership of the result.

Sequence the Three Models Instead of Picking One

Most mid-market organizations eventually need all three models, and the order determines whether the spending compounds or resets. Strategic fit here is a question of sequence rather than selection. Diagnosis precedes execution, and execution precedes platform ownership. Reversing that order produces the familiar outcome of a configured tool aimed at an undefined problem.

Sequencing also changes what each engagement is asked to deliver. The first buys clarity, the second a working process, and the third the ability to run it cheaply at scale. Organizations that hold the three engagements to one shared definition of the problem report that each partner arrives better briefed than the last. Sequence deliberately, and capability accumulates alongside the technology.

Somebody Inherits This System on Day One

Every AI deployment eventually becomes somebody’s daily responsibility, usually a person who was not in the selection meeting. Documentation, decision rights, and training are what protect that person from absorbing the organization’s ambiguity as personal stress. Undocumented systems transfer risk to whoever is standing closest.

A well-structured engagement moves capability to people rather than concentrating it in a contract. Teams that inherit a documented system, with owners named in advance and a shared definition of done, describe the transition as uneventful. Build the system so the team can carry it, because that is what makes the investment durable.

The partner question and the capability question are the same question asked at different distances. An organization that knows what it can operate will select correctly under almost any market condition. One that does not will keep buying whichever model was easiest to explain internally, and will keep paying twice for the same result. Capability compounds, and vendors do not.

The selection framework, with the cost and capability tables behind it, is set out in AI consulting vs agencies vs tool vendors.

Chief Operating Officer @COO