AI Glossary / InferenceInferenceRunning a trained model to get outputs (as opposed to training it). Where all production LLM cost and latency lives.Related termsLatency (TTFT / total)← Back to the full glossary