AI agents and “agentic AI” are among the fastest-growing topics in enterprise technology. Most demos follow the same shape: every input goes to the largest language model available, and the model decides what to do. That works in a demo. In production it is slow, expensive and hard to trust.
Across the AI systems we have built and taken into production at IT Integrated Business Solutions, from fraud detection to spam filtering to invoice processing, the same lesson keeps repeating: the LLM should be the last step in the pipeline, not the first. This article introduces the pattern. Over the coming weeks, a post on each project will show how it works in practice, with real numbers.
Key takeaways
- Use a cascade. Cheap deterministic checks first, small fast models next, a large model only for the cases nothing else can decide, and people for what remains uncertain.
- In our spam filter, over 78% of messages are decided before any LLM call, up from 55%, and every LLM verdict makes the cheap stages smarter.
- In our invoice pipeline, a 270M-parameter model classifies documents in about 0.4 seconds on a CPU, so a larger model that takes minutes per document only runs where it is needed.
- In a fraud-scoring engine, calibrating against a business-defined target raised auto-approval precision to 99.5% and cut false approvals by 78%, with no LLM involved at all.
- Models that work on the first dataset break on real-world data. Plan for continuous testing and retraining.
The pattern: a cascade, not a single model
An AI workflow that runs reliably in production usually looks less like “ask the model” and more like a series of filters, each more capable and more expensive than the one before it:
Deterministic checks
Schema validation, blacklists, lookups, arithmetic, business rules.
Microseconds · near-freeSmall, fast model
A classifier, embedding match or scoring engine that answers one narrow question.
Sub-second · very cheapLarge model
An LLM for the cases the earlier stages could not decide confidently.
Seconds to minutes · costlyHuman review
For outcomes that remain uncertain, or where an error would be expensive.
Slowest · most expensiveEach stage only sees what the previous stage could not resolve. The goal is to push as much volume as possible to the left.
Four projects show how this works, and why it matters for cost, speed and trust.
Transit Ledger: a small model as the gatekeeper
Transit Ledger finds invoices and receipts in a business’s email, extracts the details, files the documents and updates a spreadsheet, running entirely locally on a CPU server. The model that extracts the details can take minutes per document on a CPU, so it cannot see every email. A fine-tuned 270M-parameter model decides first, in about 0.4 seconds per document, and everything that isn’t an invoice or receipt is discarded. Extracted totals are then checked in code, and anything that doesn’t add up is flagged for review. The first version of the classifier also taught us that a model which passes its first test can still fail on real-world data.
SPAM Escudo: over 78% of decisions made before the LLM
SPAM Escudo filters contact-form spam through five stages: schema validation, a honeypot, rules, embeddings and finally an LLM. The cheap stages only ever say “spam”; anything that might be genuine goes to the LLM. Every spam verdict from the LLM adds the message’s links, phone numbers and email addresses to the rules, and a weekly AI-assisted review of the messages that still reach the LLM finds new ways to decide them earlier.
The result: on the same 956 real messages, pre-LLM decisions rose from 54.9% to 78.1%, and past 90% with the full pre-LLM stack. That is 78% fewer LLM calls, with no genuine message wrongly flagged.
SPAM Escudo case study →Cerberus ID: confidence thresholds without an LLM
Aletheia Systems’ Cerberus ID, for which IT-ISS is the technology architecture partner, scores college registrations to detect “ghost student” financial-aid fraud. It uses a calibrated scoring engine, not an LLM. The business set the requirement first, at least 95% precision for automatic approval, and calibration against 10,806 real registrations delivered 99.5%, with 78% fewer false approvals. The uncertain middle goes to review or to an automated ID and liveness check.
Cerberus ID case study →Where AI agents fit: give them validated tools
Agents become useful when they can take actions through tools, and the Model Context Protocol (MCP) has become a common way to provide them. In a proof of concept at Runner Technologies, an MCP server let an agent validate addresses through a dedicated validation service. Don’t ask the model to judge whether an address is real; give it a tool that knows.
A checklist for putting AI into production
- Start with deterministic checks. Schema validation, lookups, blacklists and arithmetic are fast, free and explainable. Let them handle everything they can.
- Gate expensive models with cheap ones. A small, fine-tuned model answering one narrow question can decide what the large model ever needs to see.
- Define thresholds from the cost of an error before you tune anything, then calibrate against real data to meet them.
- Validate model output in code. If the numbers must add up, check that they do.
- Route uncertainty somewhere useful: a review status, a secondary check or a person. Never let it pass silently.
- Close the loop. Every expensive decision should make the cheap stages better, so costs fall over time.
- Test on real-world data, and keep testing. The first dataset will not contain the cases that break your model.
- Choose the technology per stage, not per project. Rules, scoring engines, small local models and cloud LLMs all have a place.
Planning to put AI into production?
We help organisations design AI workflows that work inside their existing systems, and we benchmark them on real data before they go live. Read about our approach to enterprise AI in production, or bring us a workflow you’d like to test.
Benchmark your use case →