Operations

The handover is part of the automation

A workflow is not finished when it runs once. It needs an owner, a recovery path, and a test that somebody else can repeat.

Nadim NajjarPractical noteStart reading

Design for the next person on duty

An automated workflow can work correctly in a demonstration and still be difficult to operate. The person who built it knows which input to use, which warning to ignore, and where to look when it stops. The person on duty next week may know none of that. Operational handover should be part of the design, not a document written after the builder moves on.

This matters even for a small workflow. Consider a fictional process that reads accepted transfer records and prepares a daily reconciliation file. The task sounds routine. Yet a partial run, a duplicate input, or a missing destination can leave the operator uncertain about whether the file is complete and whether running the job again will duplicate the output.

Describe the unit of work

Specify what one successful run means. Does it process a complete day, one batch, or a single accepted transfer? Record an identifier for that unit and make it visible in the result. A green status without a batch identifier gives the operator little to verify.

Track input count, accepted records, rejected records, and output count where those values are meaningful. The counts should reconcile according to a documented rule. If some records are skipped, show the reason and preserve a reference to them. "Completed" should not silently mean that the job processed whatever it could and abandoned the rest.

Choose how the workflow handles an input that changes during execution. It can use a fixed snapshot or a clear cutoff, for example. Without that choice, two apparently identical runs may process different records and leave the operator chasing a discrepancy that the process itself introduced.

Make retry behavior explicit

Assume that somebody will press run twice. Give the workflow a way to recognize work it has already accepted. A retry should either resume from a known point or replace the same output safely. It must not create a second financial or operational event merely because the first response arrived late.

Test a failure after the output is written but before the success message appears. The operator will probably believe the run failed. If retrying creates another output with another identity, both may later be accepted as valid. Write down how to detect this state and how to recover without guessing.

Some actions cannot be reversed automatically. Separate them from preparation steps and require the appropriate approval. Document exactly when the workflow crosses that boundary. The person operating it should know whether they are generating a proposal or applying a change.

Write a runbook that can be followed

A useful runbook identifies the owner, the required permissions, the input location, the expected output, and the checks that establish success. It should describe the common failure states in language the operator can recognize. A screenshot of a success screen is helpful only if the document also explains which fields must match.

Include a stop condition. For instance, a mismatch between accepted input count and output count should prevent the reconciliation file from being treated as complete until somebody investigates. Explain who can authorize an exception and where that decision is recorded. Do not make the operator invent policy while handling an error.

Avoid copying passwords or access tokens into the runbook. Point to the approved access process instead. Also distinguish a service outage from a data exception: calling the infrastructure team will not resolve a disputed asset identifier.

Test the handover before closing the work

Ask someone who did not build the workflow to run it using the documentation. Use a small sample, then introduce a missing input, a duplicate batch, and a partial failure. Watch where they need a verbal explanation. Those moments tell you what the workflow or runbook still assumes.

The test should leave a record of the version, sample inputs, observed result, and corrections needed. That is not a claim of production reliability; it is a reproducible check of specific behaviors. Close the work only when ownership and support are agreed. A workflow that depends on its original builder answering every question has not been handed over yet.

This is a proposed working method, not a claim of measured employer outcomes.

Back to writingRelated work