The cheapest diagnostic in enterprise AI is two weeks of your own engineer failing at the workflow. Run it before you sign. The failure tells you what you are actually buying, and the setup doubles as the preparation.

Before you sign a forward deployed engineering engagement, run a spike: give one of your own engineers two weeks, the vendor's public documentation and SDKs, and one real workflow, and ask them to build it. The goal is not to succeed. The goal is to fail in a way that tells you something.

The exercise pays twice. The failure itself is an assessment, of the vendor's platform in unassisted hands and of your own organization's state. And everything the spike requires is exactly what an FDE will need on day one, so the setup forces preparation that otherwise happens on the FDE's clock, at the FDE's rate. FDEs do not spend their first weeks reading strategy documents. They spend them trying to get access and trying to figure out what "correct" looks like. Prepare in that order.

  1. Run the Entry Test.
  2. The Setup Is the Preparation.
  3. Assemble What "Correct" Looks Like.
  4. Read the Failure: Three Signals.
  5. The Documentation, Ranked.
  6. Define Done Before They Arrive.

1. Run the Entry Test

The buyer's guide ends with the exit test: point your eval set at a second model before renewal, and learn the true switching cost. The spike is its mirror, the entry test. Before the engagement starts, learn the true starting cost: what the vendor's platform can do without a vendor engineer driving it, and what your organization can bring to the table.

Two weeks is enough, because you are not building the production system. You are finding out where the attempt breaks. Where it breaks is the information.

2. The Setup Is the Preparation

To run the spike at all, you need three things. They are the same three things that stall every FDE engagement in its first weeks, which is the point.

Access comes first, because it is the actual blocker: sandbox credentials, API keys to the systems in play, a de-identified data extract, a security review cleared. Weeks of FDE time evaporate here. Better to burn week one of your own spike fighting for access than week three of a contract billed at software rates.

Second, an engineer with a quarter to half of their time. Whoever runs the spike is your natural candidate for the named owner the buyer's guide requires: the future orchestrator, with decision authority, who will own evals, model migrations, cost, and instruction hygiene after the vendor leaves. If nobody can be freed to run the spike, that is itself a result. Hold that thought for signal three below.

Third, one workflow, chosen. Not a list of candidates. Pick the one with real volume, a measurable outcome, and a business owner who wants it. The FDE will refine it; do not make them run the selection process.

3. Assemble What "Correct" Looks Like

Your engineer cannot judge the spike's output without knowing what a good result is. Assembling that knowledge is the single highest-leverage preparation you can do, and it is the artifact FDEs need most and buyers almost never have.

Start with worked examples of the task done well: fifty to two hundred real cases, each with the inputs, the correct output, and why it is correct. This is the seed of the eval set, which the buyer's guide argues is the real deliverable. Include the ugly ones and the exceptions.

Add the exception log: where the current process breaks, gets escalated, or gets quietly worked around.

Agents fail exactly where humans already struggle.

Then the decision rules that live in people's heads, written down badly if that is what it takes. A page of "we always X unless Y" beats a polished process document.

4. Read the Failure: Three Signals

Heavy reliance on FDEs can signal three different things, and only one of them is a product shortcoming. The spike tells you which one you are dealing with.

The product is immature. If your engineer failed on mechanics — documentation that is wrong, a sandbox that does not work, basic operations that require a support ticket — the FDEs are substituting for missing docs, SDKs, examples, and self-serve onboarding. This is a known pattern, and the original Palantir loop is explicitly designed to burn it down: the FDE does it by hand, core engineering productizes it, and FDE headcount per revenue dollar falls. The test is trajectory. If the vendor's FDE-to-revenue ratio is not falling year over year, and the things FDEs did last year are not self-serve this year, the product is not maturing and the FDEs are load-bearing. That is a red flag, because you would be depending on the vendor's staffing rather than its software.

The knowledge is tacit. If your engineer got the pieces working but could not make the workflow reliable, you have found the thing an FDE legitimately sells. It is not product knowledge; your engineer can read the API reference. It is capability intuition: what the model will and will not do reliably with your kind of data, which decomposition of a workflow holds up, where it fails silently. Documentation describes the interface, not the behavior, and the behavior shifts with every model release. That knowledge is tacit and perishable, which is why it lives in people and not in a docs site. The eval harness is the real documentation, and most customers do not have one until someone builds it.

The missing capability is yours. If the spike never ran because no engineer could be freed to run it, the diagnosis is not about the vendor. Most enterprise buyers are not choosing between an FDE and an in-house engineer coming up to speed. They are choosing between an FDE and nobody. The FDE is a substitute for the customer's missing capability, not the vendor's missing docs. That can be a legitimate purchase, but name it honestly: without someone to onboard, the handoff test from the buyer's guide has no one to hand off to, and what you are buying is a managed service at FDE prices.

5. The Documentation, Ranked

With the spike behind you, finish the paperwork in order of value per page.

Glossary and taxonomy: high value, cheap to produce, and it feeds directly into prompts and skills. Do it.

Process maps: valuable only for the chosen workflow, and only if they describe what actually happens rather than the official version. Have the person who does the job annotate them.

System architecture: one diagram of the systems in scope, who owns each, and how data moves between them. Not the enterprise-wide one.

Strategy documents: the lowest value per page. The FDE needs to know what matters and what is off limits, and that fits in a conversation.

6. Define Done Before They Arrive

Finally, agree on the metric, the target, and who signs off. Then agree on the maintenance plan in the same conversation: who owns the evals, who runs the model migration, what the FDE hands over. That is the agreement that prevents the rot the buyer's guide opens with, and it is much harder to negotiate at week twelve.

That leaves one piece of homework big enough for its own article: arriving with a point of view on where the build should live — in files, skills, plugins, or custom code. The FDE's default will be higher up that ladder than you want. It is the subject of the third piece in this series: Files First, Code Last.

The exit test tells you what it costs to leave. The entry test tells you what it costs to start. The buyers who run both never negotiate blind.