Why generalist tech DD misses AI risk
Most technical due diligence practices were built for a world where software risk lived in architecture, code quality, security, and team. They remain good at exactly that. But AI due diligence for private equity is a different exercise, because AI risk lives somewhere else: in evaluation practice, in data flows, in unit economics that move when a model provider reprices, and in defensibility claims that sound technical and are actually marketing. A reviewer who has never shipped an AI system to production reads straight past all four.
This is why funds have started adding an AI-native reviewer alongside their usual advisors on AI-heavy deals. Here is concretely what that reviewer catches that a classic tech DD does not.
What classic tech DD still does well
To be clear about the baseline: a competent generalist review will tell you whether the architecture holds the roadmap, whether the code is maintainable, whether security basics are in place, and whether the team survives a departure. On an AI-heavy target it will also, usually, confirm that the company "uses LLMs via API" and note the provider dependency. That single line is where the generalist read stops and where the actual AI risk begins, because everything that determines whether the AI works, and keeps working, is invisible at that altitude.
Eval practice: the tell a generalist walks past
The first thing an AI-native reviewer asks to see is the eval suite: the versioned set of real cases the team runs before every prompt, model, or pipeline change. Its presence, size, and freshness is the single strongest signal separating an AI system from an AI demo. A generalist reviewer rarely asks, because nothing in classic software maps to it; the closest analogue is a test suite, and the target always has one of those.
We have reviewed companies with impeccable engineering hygiene, clean repos, green CI, whose AI quality was governed by someone trying five prompts by hand before each release. That company will regress its product the next time a model version changes, and no line in a generalist report predicts it.
Data loops: the difference between an asset and a pipeline
Deal theses about AI companies lean hard on data: "proprietary data" appears in almost every AI-heavy memo we have been shown. The question a generalist read cannot answer is whether the data is a loop or a pile. A loop means production usage generates labeled outcomes that flow back into evals, retrieval, or training, making the product measurably better each quarter, and creating the compounding advantage the memo assumes. A pile means data sits in a warehouse and adorns the deck.
Tracing this requires following actual pipelines: where outcomes are captured, who labels what, which artifacts consume the data, on what cadence. When the trace comes up empty, the "data moat" line in the model should too. It is the operational half of evaluating an AI moat, and it changes valuations more often than any code finding.
Margin sensitivity to model pricing
An AI-heavy P&L carries a cost line that classic software never had: inference, priced by a third party who can change the number. The exposure runs both directions. A provider price increase, or a forced migration off a deprecated model, can compress gross margin in a quarter. A price drop can do the same by handing every competitor the capability the target charges a premium for. The reviewer’s job is to model both: cost per request fully loaded, margin at ten times volume, margin if model prices halve or double, and switching cost across providers. This is the margin question every AI deal should start with, and it needs someone who has actually operated these systems, because the retries, context growth, and orchestration overhead that dominate real inference bills never appear in the target’s own model.
Fake defensibility
Finally, the claims. "Proprietary models" that are a fine-tune any competitor could reproduce in a week. "Our AI" that is a system prompt over the same API everyone rents. Accuracy numbers from a benchmark the team built itself. None of this is fraud exactly; it is the ambient exaggeration of the category, and a reviewer who cannot reproduce the claims cannot deflate them. The test is practical: an AI-native reviewer can sit with the team, rerun the numbers, and tell a real capability from a wrapper in a day. A generalist writes down what the slides said.
When to add AI due diligence to a deal
The trigger is thesis weight, in line with when AI due diligence is worth running at all: if the valuation assumes the AI is the product, or the value-creation plan assumes AI will transform the target’s cost base, the four areas above carry the thesis and deserve a specialist read. The setup that works is complementary, never duplicative: the generalist covers the classic scope, the AI reviewer covers evals, data, margins, and defensibility, and both feed one risk register. Three to four days of focused work is enough, and it fits inside a deal window that is already open. That is the shape of our technical and AI due diligence: run by engineers who ship AI in production, alongside whoever already does your classic workstream.