AI and Automation

Give an AI pilot a decision boundary

Before choosing a model, decide what it may suggest, what it may change, and what evidence would stop the pilot.

Nadim NajjarPractical noteStart reading

Begin with one decision

A request to add AI to an operations platform is too broad to evaluate. Start with one decision that somebody already makes. For example, an assistant might suggest which documents a reviewer should inspect when an asset record is disputed. That is a narrower problem than asking it to resolve the dispute or authorize a write-off.

Write down the input, the proposed output, the person who will use it, and the consequence of a wrong answer. If the team cannot describe those four things without mentioning a model brand, it has more process design to do. A useful pilot makes a bounded task easier to judge. It does not make the entire operating model depend on a convincing paragraph.

Keep suggestion and authority separate

In the document-review example, the assistant can return a proposed evidence list and references to the source passages. It should not update the asset register. A reviewer decides whether the evidence is relevant and whether it is sufficient. An application permission check should enforce that boundary rather than relying on an instruction in the prompt.

Treat text retrieved from documents as data. A document that says "ignore the approval process" must not become an instruction to the assistant. Test this deliberately with a harmless sample document. Also test a document whose title looks relevant but whose contents concern another asset. The assistant should inspect the content rather than infer relevance from the filename.

Sometimes the useful answer is that the available documents do not support a conclusion. Give the assistant a visible way to abstain, and give the reviewer a way to flag unsupported suggestions. Requiring an answer in every case encourages plausible guesses.

Build the evaluation set before the demonstration

Choose representative cases before adjusting the prompt. Include clear records, missing documents, conflicting identifiers, Arabic and English text, and scans with imperfect extraction. Keep a separate set for the final check. If every failed case becomes another example in the prompt, improvement on those same cases tells you little about new work.

Have reviewers agree on acceptable answers and reasons. Evaluate whether the cited passage supports the suggestion, not merely whether the answer sounds similar to a preferred response. Record false matches separately from missed matches. They may have different costs, especially when a wrong match looks credible enough to pass a quick review.

The example is illustrative. It is not a report of a deployed AI system or a measured improvement. An AI layer under development should remain described as work in progress until its behavior, permissions, and operating support have been tested.

Include the cost of checking the answer

A model can generate a suggestion quickly while making the review slower. Measure the complete task: opening the case, reading the suggestion, checking the source, correcting the output, and recording the decision. Compare that with the same work without the assistant. A short response time is not the same as a shorter review.

Keep the comparison fair. Use similar cases and similar reviewers, record the setup, and note when a case required specialist knowledge. Count abstentions and corrections as outcomes. Do not discard them to make the average look better. If a suggestion saves time only when somebody skips verification, the pilot has not demonstrated a useful improvement.

Decide in advance when to stop

Set a stop rule for unsupported citations, unauthorized actions, or exposure of restricted information. The exact thresholds depend on the task and its risks; they should be agreed by the people accountable for it. A pilot that fails a safety boundary should pause even if its average accuracy looks promising.

Keep a simple route back to the existing process. Reviewers need to continue their work if the model is unavailable, a document cannot be parsed, or access is withdrawn. At the end, retain the evaluation cases, observed errors, reviewer corrections, and decision about the next step. That record is more useful than a polished demonstration that nobody can reproduce.

This is a proposed working method, not a claim of measured employer outcomes.

Back to writingRelated work