Local AI on a CPU: How a 270M-Parameter Model Makes Invoice Processing Practical

Every small and medium-sized business receives invoices and receipts by email, and someone has to find them, read them, file them and type the numbers into a spreadsheet. It is exactly the kind of work AI should take over. The catch: financial documents are sensitive, and sending every email to a cloud AI service is not something every business is comfortable with.

Transit Ledger is our answer: an AI pipeline that identifies invoices and receipts in a mailbox, extracts the key details, files the documents and updates a spreadsheet. A version already runs inside Gmail. The version we are building runs entirely locally, on an ordinary CPU server. This post explains how we made that practical, and what a small model taught us about real-world data.

Key takeaways

  • Local AI on a CPU is practical if you don’t send everything to your largest model.
  • A fine-tuned 270M-parameter model classifies a document in about 0.4 seconds on a CPU (601 documents in 252 seconds).
  • The 4B-parameter extraction model can take minutes per document, so it only ever sees invoices and receipts.
  • Extracted figures are checked in code; anything missing or inconsistent is set to REVIEW rather than filed.
  • The first classifier learned a shortcut. Real-world data exposed it, and continuous fine-tuning fixed it.

Why run it locally, on a CPU?

Most small and medium-sized businesses already have a server. Very few have a GPU. Running the whole pipeline on the hardware a business already owns means:

  • Financial documents never leave the building. No invoice, receipt or email is sent to an external AI service.
  • No per-document AI charges. The cost is the server the business already runs.
  • No new infrastructure to buy or manage. A CPU server is familiar territory for any IT provider.

The trade-off is speed. On a CPU, language models are slow, and the more capable the model, the slower it gets.

The constraint: minutes per document

Extracting the details from an invoice needs a capable model. We use Google’s Gemma 4 E4B, a 4-billion-parameter model, which extracts seven fields: company, invoice number, transaction date, services, total amount, tax amount and amount before tax. On a CPU it can take minutes to process a single document.

That is fine for the invoices themselves. It is not fine for everything else in a mailbox. Most email is not an invoice, and attachments include signatures, newsletters, contracts and photos. As an illustration: at two minutes per document, running 1,000 documents through the extraction model would take more than 33 hours. Our classifier gets through the same 1,000 in about seven minutes.

The solution: a small model as the gatekeeper

The pipeline only uses the expensive model where it adds value:

  1. Read. The pipeline reads new emails and their attachments.
  2. OCR. Attachments go through PaddleOCR, which extracts the text only. Everything downstream works on text, which keeps it fast.
  3. Classify. A fine-tuned Gemma 3 270M model decides one thing: is this an invoice, a receipt or something else? “Other” is discarded immediately, with no further processing.
  4. Extract. Only invoices and receipts go to the 4B model, which extracts the seven fields.
  5. Validate. Ordinary code checks the extracted figures (below).
  6. File. An AI agent files the document and updates the spreadsheet with its details.
~0.4 s
Per classification on a CPU: 601 documents in 252 seconds
~8,500 / hour
Documents the classifier can screen at that rate
Minutes
Per document for the 4B extraction model, which only sees invoices and receipts

A 270M-parameter model is tiny by today’s standards. That is the point: it answers one narrow question, fast enough that screening a whole mailbox costs almost nothing, and it decides what the large model ever needs to see.

Don’t trust the model’s output: check it in code

Extraction is where a generative model is most likely to be confidently wrong: a misread digit, a tax amount taken from the wrong line, a missing invoice number. So nothing is filed on the model’s word alone. Code checks the result first, and the simplest check is one of the most effective:

amount before tax + tax amount = total amount

If a field is missing, incomplete or fails that check, the entry gets a REVIEW status instead of being filed silently. The person responsible looks at a handful of flagged entries, rather than auditing every one to be sure.

The lesson: the first classifier learned a shortcut

The first version of the classifier worked well on its initial training set. Once it saw more real-world email, it started to misclassify.

The cause was the training data. Certain keywords in the first training round had become far too important to the model’s decision, and numbers that looked like totals were pushing documents towards “invoice”. The model had learned a shortcut, “documents with totals are invoices”, rather than what actually makes something an invoice or receipt. That shortcut worked on the first dataset and failed on the variety of a real mailbox.

The fix was not a bigger model. It was better data, and a process rather than a one-off: collect the real-world documents the model gets wrong, add them to the training set, fine-tune again, and repeat as new kinds of documents appear. In practice, the most valuable training examples are the look-alikes, documents that share features with invoices but aren’t, because they are exactly where a shortcut breaks.

A model that passes its first test is the start of the work, not the end. Plan for the retraining loop from day one: a way to capture mistakes, a growing test set drawn from real data, and a routine for fine-tuning and re-testing.

What this means if you’re planning AI on your own hardware

  • Size each model to its job. A tiny model for the high-volume screening decision, a larger one only for the step that needs it.
  • Discard early. The cheapest document to process is the one you stop processing after 0.4 seconds.
  • Keep checks in code. Arithmetic and required fields don’t need AI, and they catch the errors AI makes.
  • Make uncertainty visible. A REVIEW status is better than a silently wrong spreadsheet.
  • Budget for retraining. Real-world data will surprise the first version of any model.
How we can help

Enterprise AI integration

Transit Ledger works with the email, folders and spreadsheets a business already uses, on a server it already owns. We bring AI into your existing databases, applications and cloud or on-premise infrastructure in the same way, without replacing systems that already work.

  • An integration design that fits your current systems and data flows
  • Local, cloud or hybrid models, chosen for your cost, speed and data-privacy requirements
  • Validation in code at every point where AI output enters your systems
  • Review routes for anything the AI can’t process with confidence

Considering local AI for your business data?

We design AI pipelines that run on the infrastructure you already have, and we test them on your real documents before they go live.

Discuss your document workflow →