Skip to content
Working with the tools, not about them

Workflows

Most of the work in a document process happens before the tool sees anything

Getting source material into a state where it can be processed at all is usually the larger half of the job, and it is the half that gets left out of every estimate.

By Naina Sethi3 min read

Editorial note. Independent reporting and analysis. Nothing here is sponsored or paid for. How we work.

The demonstration used a clean file

A process that works beautifully on one tidy document meets reality in the form of a shared folder. Scans of varying quality, some of them photographs of screens. Files with three generations of naming convention. Spreadsheets where the header is on row four and somebody has merged cells to make a title. Duplicates with slightly different contents and no way to tell which is current.

None of that is exotic. It is the normal condition of accumulated organisational material, and it is the input the process will actually receive.

The estimate that gets quoted comes from the clean case, which is why processes of this kind so often deliver a working demonstration and then take three months to reach production.

Preparation is a set of specific, boring tasks

The work divides into recognisable pieces. Converting formats so everything arrives in one shape. Splitting bundles where twelve documents were scanned into a single file. Deciding which of four near-identical versions counts. Getting text out of images, and identifying which pages cannot be read at all.

Each of these is a task with a known solution, and most of them are better done by ordinary software than by anything that generates text. Conversion, splitting, deduplication and text extraction from scans are all mature problems, and using a language step for them is slower, more expensive and less predictable than using the tool built for the job.

The exception is the messy middle: material that is technically readable and structurally chaotic, where a judgement is needed about what a section is. That is where the flexible tool earns its place, and it is a much smaller portion of the work than the first pass suggests.

Decide what counts as a unit before anything runs

The question that causes the most rework is the simplest one. What is one item? A file, a document, a page, a case, a customer, a contract with its amendments? The answer decides the shape of everything downstream and it is usually assumed rather than settled.

Get it wrong and the symptoms appear late: totals that do not reconcile, a report where one case is counted three times, an output set that cannot be joined back to the source. By then the run is finished and the fix is a rerun.

Settle it by looking at the awkward examples rather than the typical one. The bundle containing two contracts, the case with a second file added later, the amended version. Those decide the definition; the straightforward examples never do.

Sort the unreadable material into its own pile

Every real collection contains items that cannot be processed: illegible scans, corrupted files, documents in a language nobody expected, handwriting, forms so damaged that the fields are gone. The worst outcome is for these to pass silently through the process and emerge as items with empty or invented fields.

So the first pass should be a triage that separates what can be processed from what cannot, with the second pile counted and kept rather than discarded. Knowing that eleven per cent of the collection is unreadable is a finding; discovering it later as a set of odd results is a mess.

That pile is also where the human effort belongs. A person working through the exceptions is a sensible use of an afternoon, and it is a far better use than the same person checking output from material that was never legible.

Count the preparation when you estimate

The practical instruction is to time the preparation on a real sample and put that figure in the estimate as its own line. If forty documents took a day to get into shape, four thousand will not take ten days by simple multiplication, and the deviation is worth thinking about before committing to a date.

It is also worth asking whether the preparation is a one-off or a recurring cost. A historical archive is cleaned once. A monthly intake of new material has to be cleaned every month, and a process that ignores that has an ongoing manual step nobody budgeted for.

Occasionally this arithmetic kills the project, and finding that out during a sample is a good outcome rather than a bad one. A process whose preparation cost exceeds the value of the output is a process that should not be built, and that conclusion is much cheaper to reach in week one.

Common questions

Why not use a flexible tool for the preparation as well?

Because format conversion, splitting, deduplication and text extraction from images are mature problems with dedicated software that is faster, cheaper and more predictable. Reserve the flexible step for material that is readable but structurally chaotic, which is a smaller share than it first appears.

What single decision causes the most rework?

What counts as one item — a file, a page, a case, a contract with its amendments. It is usually assumed rather than settled, and the symptoms appear late as totals that do not reconcile or output that cannot be joined back to the source.

What should happen to material that cannot be processed?

It should be separated in a triage pass, counted, and kept. Letting it through produces items with empty or invented fields, and the exceptions pile is where human attention is genuinely worth spending.

Workflowspreparationdocumentsestimationprocess
Naina Sethi
Consumer editor, Prompt After Prompt

Naina covers prompt craft, writing with ai, images & audio and the questions readers actually send in and is happiest when a piece answers the question completely.