PRACTICAL AI

Start an AI experiment with one business task

Define the work, check the result and measure the complete task.

Begin with a recurring task whose output someone can check. A draft based on public information can be a manageable first experiment: the input is available, the expected result can be described and a person can compare the output with the source.

Choose the task before choosing a model. “Use AI in the business” is too broad to test. “Prepare a short draft summary of a public product guide” gives the experiment a useful boundary.

Write a pilot card

Question Illustrative answer
What task are we testing? Draft a short summary of the current public product guide.
What input may we use? The public guide and a written summary format.
What counts as acceptable? Product statements match the guide; important limitations remain visible.
Who checks the output? The person responsible for the guide.
What will we measure? Time to a usable final version, corrections and missed information.
What makes us stop? Unsupported claims persist or review takes longer than the ordinary process.

This fictional example is a starting point. Replace the answers with the task and responsibilities in your organization.

Establish the ordinary process

Complete a few comparable examples without AI and record the time required. Note the checking work as well as the drafting. Then run the experiment on a small selection that includes a straightforward item and an item with exceptions or awkward wording.

Do not choose only the easiest example. You need to understand where the process becomes difficult before deciding whether it is useful.

Check the final result

Generated text can sound confident while being wrong. NIST describes that risk in its generative-AI profile. Compare claims with their source and make the responsible reviewer visible in the process.

Keep a simple error list: invented statement, missed limitation, incorrect number or unclear wording. Patterns are more useful than a general impression that the draft looked good.

Choose a task with a clear boundary

Compare two possible starting points. “Improve customer service” contains many decisions and depends on the behavior of several people. “Draft a summary of a public returns policy for a colleague to check” has an identifiable source and output. The second task is easier to test because a reviewer can explain what a correct answer must preserve.

A manageable pilot also has a clear end. Decide how many examples you will examine and what question the test is supposed to answer. Five carefully chosen examples can expose a problem worth investigating; they cannot establish that a method will work across every customer or document. Keep the conclusion proportional to the test.

Prepare a small evaluation set

Select examples before looking at generated results. That reduces the temptation to keep only the outputs that make the experiment appear successful. Include variation that matters to the task rather than collecting a large set of nearly identical inputs.

Example type What it helps you examine
A short, straightforward guide Whether the basic format is followed.
A guide containing an exception Whether the exception survives the summary.
A document with similar product names Whether names and features become mixed.
A document missing an answer Whether the draft admits the gap or invents information.

Write the important checks beside each example. A summary of a returns policy might need to preserve both the ordinary time limit and the exception for a particular item. A fluent paragraph that loses the exception should not pass merely because it reads well.

Give the reviewer a repeatable method

For every draft, compare the statements with the input, mark omissions, identify unsupported additions and record the corrections. Use the same review method for each example. A vague “looks fine” from one reviewer and a line-by-line check from another are not comparable evaluations.

Keep the result separate from the explanation. “The draft added a feature absent from the source” is an observation. “The instructions may have encouraged a sales tone” is a possible explanation to test. Changing one element at a time makes it easier to see whether a revision helped.

Work through a fictional timing example

Suppose the ordinary process takes eighteen minutes to produce a checked summary. In a pilot, input preparation takes three minutes, generation takes one, review takes seven and corrections take four. The complete task takes fifteen minutes. The observed difference for that example is three minutes, not the seventeen-minute difference between manual drafting and generation alone.

Now suppose an exception-heavy guide takes twenty-two minutes because the reviewer has to rebuild the summary. Record that result too. A useful conclusion might be to continue testing short standard guides while keeping complex ones in the ordinary process.

Questions before the next test

Should the first pilot use live customer decisions? Start with work that can be reviewed before it affects someone. The choice depends on the consequences of an error and the organization’s actual controls.

What if the output improves after repeated prompting? Include those attempts in the time and effort record. They are part of the process being evaluated.

Is a small time saving enough? Judge it alongside quality, maintenance and review demands. A saving is useful only if the organization can use the resulting process reliably.

Make the next decision

Compare the full time to a usable result, the quality of that result and the effort required to correct it. Continue with a narrow use, revise the task or stop the experiment based on what you learned.

NIST’s AI Risk Management Framework provides broader voluntary guidance for considering risks throughout AI use. For this first test, keep the decision small enough that you can explain why it is worth continuing and what evidence would change your mind.

Sources & further reading

Compare a demonstration with a real pilot