Skip to content

Glossary

What is inference?

Inference is the step where a trained AI model is used to produce an output, such as an answer, a label or a prediction, from new input.

What it means

AI models go through two stages. Training is when a model learns from large amounts of data, which is expensive and done by the provider. Inference is every time the trained model is used: each question answered, each document read, each image described.

When you pay for an AI API, you are mostly paying for inference, usually priced by tokens in and tokens out.

Why it matters for a business

Inference cost and speed decide whether an AI feature is practical at your volume. A model that is perfect but slow or expensive per request may not work for live chat or thousands of documents a day.

A business example

Things to watch

  • Estimate monthly cost from real request sizes.

  • Measure response time where people are waiting.

  • Use smaller models for simple, high-volume steps.

  • Cache repeated requests where possible.

Keep exploring

FAQ

Questions about inference

Have a question that is not here? Ask us directly.

Start a project

Tell us what you want to build. We will show you a faster path.

Send a short brief. We reply with questions, a suggested plan and an estimate you can compare with other offers.

Your privacy choices

We use necessary storage to run this site. With your permission we also use Google Analytics to see which pages help people, and load maps from Google. You can change this at any time. Read the cookie policy.