Knowledgedocuments

Automatic document classification: less sorting in the daily incoming workload

Five document types, practical review steps and a complete worked example show when automatic classification reduces work for your office team.

An invoice arrives alongside a delivery note, an order confirmation, a credit note and a general letter. Before anyone works with them, the files are opened, named, classified and passed on. That preparation can repeat with every delivery of documents, even though the business already knows the document types it receives.

Automatic classification aims to reduce this recurring sorting work. Its business value appears when your team can start the real task sooner: understanding a transaction, identifying a delivery or reviewing an invoice. The useful measure is therefore the complete journey from an incoming document to a workable basis for the next activity.

Seeing the correct word “invoice” on a screen is not enough. Document type, extracted information and business decision need to fit together. This article develops a before-and-after process for five document types and works through a hypothetical batch to establish which manual tasks actually disappear. The examples describe a working method; they are neither a promised recognition rate nor a catalogue of automatically supported specialist fields.

Choose document types according to the work they support

Start with the documents your team handles regularly. The first classification should answer a practical question: does this document lead to a different next activity? An invoice is read differently from a delivery note. Two invoices with different logos, layouts and file names can still belong to the same document type.

Select a manageable sample from the actual incoming workload. Review it with the person who processes the documents afterwards. Which types do they already distinguish, and where do they repeatedly hesitate? A covering letter with an attached invoice illustrates that the accompanying message and the actual invoice can serve different purposes. A file name alone does not resolve that distinction.

Avoid beginning with dozens of rarely used subcategories. An additional category earns its place when it improves retrieval, checking or further work. Otherwise, sorting time simply becomes discussion about which obscure label is correct. Describe each chosen type with a short operational purpose and an example the team recognises. That gives colleagues a shared basis for reviewing results rather than a taxonomy they need to memorise.

Visit the documents module page to see how webRichtung analyses and classifies documents for retrieval, and create your account. Start with the document types for which your office currently repeats the same preparatory steps each day.

Separate recognition, review and decision

Classification first answers what kind of document you have. Analysis can provide additional information that helps with retrieval and association. The business decision answers a further question: what should your organisation do with this document? These three levels are connected, but they should not be treated as interchangeable.

A file recognised as an invoice has not thereby passed a substantive review. An extracted amount alone does not establish whether the invoiced service matches the order. Equally, identifying a delivery note does not establish that everything was received. The responsible colleague still needs to assess the information in the context of the actual transaction.

Before the trial, decide which information matters for each next step. Document type alone may improve a search. Associating a document with a transaction may require a name, a reference in the content or a date. These are review questions for your team, rather than claims that every mentioned item is available as a separate automatically extracted field.

Keeping the levels separate also makes the benefit measurable. When the system handles recurring classification, people can concentrate on contradictions and information that affects a decision. The aim is a shorter, understandable review. Entering every item manually for a second time would consume much of the potential reduction in work and make the process harder to assess fairly.

How the process changes for five document types

For this comparison, assume an office receives five familiar types of document. Previously, a colleague opens every file, identifies its type, enters search information in a working overview and prepares it for substantive processing. In the proposed new process, she uses automatic classification as the starting point, checks the necessary information against the document and handles unresolved cases individually.

For an invoice, she previously reads the heading and gathers the details needed for filing. In the new process, she checks that the recognised type matches the content and that the date and amount used are the intended values. The substantive invoice review then continues within the responsible team’s existing process. The improvement is that she does not rebuild the preparatory classification from scratch.

For a delivery note, she previously searches the text to establish which delivery the file concerns. Afterwards, a correct type gives her a quicker starting point for retrieval and processing. She reads the relevant delivery reference and associates the document with the existing transaction. Whether the goods arrived in full still depends on the actual delivery information available. The document class does not replace that finding.

For an order confirmation, she previously establishes whether the document is a new quotation or confirmation of an existing order. Afterwards, she checks the classification and compares the contents with the related transaction. Differences that require further discussion deserve particular attention. Document type helps with retrieval; it does not automatically confirm that all commercial conditions agree.

For a credit note, the value lies in distinguishing it clearly from an invoice. Previously, a similarly designed file might accidentally enter the same working group. Afterwards, the colleague explicitly checks the content and the meaning of the amount. She establishes the link with the original transaction where the documents provide that information. This is a team review step, not a promise of automatic offsetting.

A general letter may not fit one of the frequent, clearly defined document types. Previously, someone might create a new folder on the spot or keep it in a personal collection. In the new process, the document remains available in the shared incoming workflow while its purpose is clarified. The colleague reads whether it creates any task at all. One unusual letter does not justify a new permanent category for the whole business.

Check extracted information in its business context

A document can contain several plausible dates. The letter date, a delivery date and a reference to an earlier transaction are different information. If colleagues later search by date, they need to understand which context they expect. Review whether the information has the right meaning for the intended purpose, as well as whether the characters have been read correctly.

Amounts need the same attention. A document can contain subtotals, overall amounts and references to earlier figures. For a meaningful trial, choose specific retrieval tasks: “find the document for this transaction” or “show me the document with this total”. Compare the analysis with the original document. You will see whether the information supports real work rather than merely looking plausible on screen.

Focus substantive review on information that changes the next activity. A minor spelling error in accompanying text has a different consequence from an amount about to inform another decision. Your business sets that weighting. It is a way to prioritise actual work, rather than an instruction to disregard missing or incorrect information.

When something is wrong, record briefly what was unclear and how it was resolved. “Date does not match the search purpose” is more useful than “AI wrong”. The specific finding shows whether the source document, analysis or your own working rule needs attention. A vague error list can produce general distrust without improving the process that colleagues use every day.

Handle exceptions without creating a second filing system

An exception needs an understandable route to resolution. It should not disappear into a private folder because its automatic classification was unsuitable. Agree who takes unclear documents and how their processing state remains visible in your existing workflow. Describe the open question so the next colleague does not have to start again.

Distinguish the causes. An unreadable source may require a better file. An incorrect document type requires a classification correction. A statement that is unclear in business terms requires someone familiar with the transaction. These cases create different work. Grouping them all under “not recognised” makes the trial’s results much less useful for deciding what to change.

A PDF containing several documents also deserves attention. First establish what the file actually includes and which parts must be understood separately for subsequent work. Do not assume automatic splitting unless you have verified it for your process. The practical question remains whether another colleague can reliably find and associate every document they need.

When appropriate, turn recurring exceptions into a shared rule. If two colleagues regularly classify the same type differently, start by agreeing on its definition. Adding categories does not automatically resolve an unclear concept. A rare individual case, on the other hand, may be handled adequately with a brief explanation rather than a permanent addition to the classification scheme.

Work through a complete document batch

Our hypothetical trial batch contains 50 documents: 20 invoices, ten delivery notes, eight order confirmations, four credit notes and eight general letters. The total is 50. This distribution illustrates different recurring tasks. It is neither a typical industry mix nor a statement about recognition performance or document volumes at webRichtung.

For the previous process, assume two minutes per document to open it, identify the type, record search information and prepare its association. That gives 100 minutes for the batch. Subsequent substantive processing is excluded because the comparison concerns preparation only. That later work remains a separate activity in both versions of the process.

For the new process, assume ten minutes to prepare and place the entire batch into the intended input route. Reviewing 40 straightforward documents takes half a minute each, or 20 minutes. Ten exception cases require three minutes each for review and clarification, adding 30 minutes. The total is 60 minutes: ten plus 20 plus 30. Every document in the example has been included in one of the two review groups.

The preparation being compared falls from 100 to 60 minutes, a reduction of 40 minutes per batch. The ten exceptions are already included. These are freely chosen assumptions, not a time guarantee. Replace them with measured handling times in your trial. Record waiting time without active work separately so it does not silently become counted as staff time.

Assess whether the reduction justifies the setup effort

Processing one batch faster does not automatically cover implementation. Suppose preparing the working rules and briefing the team takes 120 minutes once in this example. If each comparable batch continues to save 40 minutes of preparation, the initial staff time is mathematically recovered after three batches. Software costs and other expenses are not included in that calculation.

Also calculate a version with more clarification work. If 20 documents rather than ten are exceptions, still taking three minutes each, 30 straightforward cases remain at half a minute each. With ten minutes of preparation, the total is 85 minutes: 60 plus 15 plus ten. Compared with the previous 100 minutes, the saving is now 15 minutes. Recovering the assumed 120 minutes of setup time would take eight comparable batches.

That difference helps you choose where to begin. Automatic classification may already be useful for frequent invoices while a mixed historical collection creates too many unusual cases. You can improve the current incoming flow first and introduce additional types gradually. The decision follows the real work mix, rather than a desire to process the largest possible number of files immediately.

Check retrieval as well. If classification becomes faster but colleagues spend longer searching afterwards, part of the calculation is missing. Ask another person to find selected documents using their content and the agreed search information. Shorter preparation combined with useful subsequent work shows whether the new process actually benefits the office.

Use documents for a clearly defined first input flow

webRichtung documents supports browser upload, email import and batch processing. Its module page describes automatic analysis and classification, search criteria including content, type, date and amount, and the ability to inspect the recognised classification at the document and reassign it when needed. This fits an initial use case focused on making recurring documents usable sooner.

Choose a manageable real batch and agree on the five or fewer document types relevant to the trial. Before starting, write down the retrieval task for each type and who will resolve open cases. Then let the team work with the analysed documents. Successful input alone is not enough to establish that the information is ready for the next person.

After the trial, retain rules that reduce sorting work while keeping substantive decisions understandable. Address specific error causes and expand only when the next document type offers a clear benefit. The result is a practical incoming workflow whose value can be judged by the transactions colleagues process and the manual preparation they no longer need to repeat.

Explore documents with automatic classification and archive search, and create your account on the module page. Begin with a batch that lets your team directly compare manual preparation with focused review, including the exceptions that arise in its own work.

Frequently asked questions

Should we introduce all document types at once?

Start with frequent types whose distinction supports a specific search or next task. Rare exceptions can initially be handled individually.

Has a recognised invoice already passed substantive review?

No. The document class identifies the kind of document. Whether its contents match the transaction remains a separate review.

How do we include exceptions when measuring value?

Add preparation, review of straightforward documents and resolution of every exception. Compare the same activities with the previous process.

Are the example times a promised saving?

No. The 50 documents and all times are hypothetical. Your measured work, setup effort and other costs determine the business result.

webRichtung Documents

Find information in your documents

Search receipts and documents by content. Find the information you need even when you do not know the file name.