Skip to main content

Glossary

What Is Inference

Inference is the process of using a trained AI model to generate predictions or outputs from new input data. Plain-English definition.

Definition

Inference is the stage where a trained AI model is put to work. It receives new input data and produces a prediction, classification, or generated output based on what it learned during training. Training is the learning phase; inference is the doing phase. When you type a question into a chatbot and receive an answer, the model is performing inference. When a fraud detection system analyses a transaction in real time and flags it as suspicious, that is also inference. It is the moment the investment in training pays off as useful output.

Why It Matters

From a business perspective, inference is where AI delivers its value. Training happens once (or periodically), but inference happens every time the model is used. At scale, that can mean thousands of requests per minute. This has direct cost implications: inference requires computing resources, and the more complex the model and the higher the volume of requests, the more it costs to run. Understanding inference helps businesses plan for operational costs, latency requirements, and infrastructure decisions. A model that is brilliant in testing but too slow or expensive to run in production is not viable. Balancing inference speed, accuracy, and cost is a key part of deploying AI successfully. See our AI development services for how we approach this in practice.

Example

An insurance company deploys a model that analyses claim photographs to estimate repair costs. During training, the model studied hundreds of thousands of annotated damage photos. Now in production, a claims adjuster uploads a photo of a dented car panel. The model analyses the image and returns an estimated repair cost within two seconds. That is inference at work: the model applies what it learned to a specific, real-world input and produces an actionable result that would have taken a human assessor fifteen minutes.

For related definitions, browse the glossary or read about large language models, which are one of the most widely used inference systems in business today.

Portrait of Alexander De Sousa, founder of Digital Royalty
Founder-led
“I’ve put everything I know into how this company works — the standards, the method, the care on every project. It runs through the whole team, and I hold us all to it.”

Alexander De Sousa · Founder LinkedIn

Featured on BBC Radio Solent

Get started

Tell us what you need

A few quick questions, then a straight answer from a real person — usually within a few hours.

Tell us what you're working on

Whether it's a new site, a platform, or a process that shouldn't be manual any more — we'll tell you honestly if we can help.