The short version
A pilot without a decision is a demo with a calendar invite.
Run an AI pilot around one bounded job, one accountable owner, a small set of representative cases, explicit safeguards, and a date when the team will decide to scale, revise, or stop. If you cannot name the decision, the pilot is not ready.
01 · The actual job
Use the pilot to reduce uncertainty—not to prolong it.
Many AI pilots begin with a model capability: summarize documents, draft replies, classify requests, search company knowledge. Those capabilities can be useful, but they do not yet identify the business decision. The better starting point is a recurring workflow where the team can name the current friction, the person who owns the outcome, and the consequence of getting it wrong.
A good pilot earns permission for the next investment. A weak pilot only creates more enthusiasm, more exceptions, and less clarity about what to do next.
This is a ForgedFuture point of view, not a claim that every experiment needs the same ceremony. The practical standard is proportionality: a low-risk internal drafting aid deserves a lighter process than a system that changes customer records, approves money, or influences a consequential decision.
02 · The pilot contract
Write five lines before anyone starts building.
The job
Name one observable workflow and its beginning and end. “Help sales” is too broad; “prepare an account brief from approved CRM and call notes” is a testable job.
The owner
Give one operating leader authority to judge usefulness, accept tradeoffs, and decide what happens after the pilot.
The evidence
Choose representative cases and a few measures: completion quality, time to complete, correction rate, escalation rate, or user follow-through.
The controls
State what the system may access or change, which actions require review, and how an operator can correct or stop work.
The decision date
Set a real date and the possible outcomes: scale the narrow use, revise the design, collect a specific missing fact, or stop.
NIST’s AI RMF Core frames AI risk work around governing, mapping context, measuring, and managing. Its functions are not a prescribed checklist, but they make a useful prompt: does this pilot have enough context, accountability, measurement, and response capacity for its actual risk?
03 · What to observe
Watch the work after the output appears.
Model outputs are only the first layer of evidence. The operating question is whether a person can use the result correctly, quickly, and consistently inside the real workflow. That means observing where information is missing, where a reviewer changes the result, where the process stalls, and where the system asks a person to make a judgment it cannot safely make.
Start with a deliberately small case set that includes ordinary work, known edge cases, and a few examples that should be escalated. Keep a lightweight record of failures and corrections. The goal is not a perfect score; it is a legible pattern that tells the team whether the design is improving the job.
OpenAI’s practical guide to building agents recommends establishing evaluations and planning for human intervention, especially around high-risk or repeated failures. Anthropic’s evaluation guidance similarly emphasizes making agent behavior visible before it affects users. Both perspectives support an evidence loop that is designed before scale, not retrofitted after an incident.
04 · The decision gate
End the pilot with a choice, not a vague next step.
| What the team learned | Responsible next move |
|---|---|
| The outcome is useful, repeatable, and controlled. | Scale only the proven workflow. Add monitoring and an operating owner before adding adjacent work. |
| The job is valuable, but the system fails for understandable reasons. | Revise the context, tool access, workflow boundary, or review design, then rerun a defined test. |
| The output is interesting, but it does not change the operating outcome. | Stop. Preserve the learning and redirect effort to a better candidate workflow. |
| The team cannot judge the result or safely intervene. | Do not scale. Add measurement, controls, and ownership—or select a lower-risk use case. |
The important constraint: “Continue exploring” is not a decision. If the team needs more evidence, specify exactly what is missing, who will obtain it, and when the decision gate returns.
Outside perspectives
Useful references for building the next pilot.
These sources are included to help readers test the approach against published practice, not as a substitute for understanding the particular workflow at hand.
- NIST AI RMF CoreA voluntary risk-management framework organized around govern, map, measure, and manage.↗
- A practical guide to building agentsOpenAI guidance on identifying candidate workflows, establishing evaluations, and introducing human intervention.↗
- Demystifying evals for AI agentsAnthropic’s account of making multi-step agent behavior measurable before it reaches users.↗

