CHEAPEST LLM APIS

The cheapest LLM APIs, ranked by real cost.

GPT-5 nano is currently the cheapest verified-price LLM API at $0.138 per 1M blended tokens ($0.05 in / $0.40 out). Rankings below use a blended 3:1 input:output rate and re-verify against official provider pricing every day, so this list is always current.

Pricing verified

Ranking: cost per 1M tokens

#ModelProviderInput / 1MOutput / 1MBlended (3:1)10K in / 2K out request
1GPT-5 nanoOpenAI$0.05$0.40$0.138$0.0013
2GPT-4.1 nanoOpenAI$0.10$0.40$0.175$0.0018
3Gemini 2.5 Flash-LiteGoogle$0.10$0.40$0.175$0.0018
4DeepSeek V4 FlashDeepSeek$0.14$0.28$0.175$0.00196
5GPT-4o miniOpenAI$0.15$0.60$0.262$0.0027
6GPT-5.6 LunaOpenAI$0.20$1.20$0.45$0.0044
7GPT-5.4 nanoOpenAI$0.20$1.25$0.463$0.0045
8Grok Code Fast 1xAI$0.20$1.50$0.525$0.005
9DeepSeek V4 ProDeepSeek$0.435$0.87$0.544$0.00609
10Gemini 3.1 Flash-LiteGoogle$0.25$1.50$0.563$0.0055
11GPT-5 miniOpenAI$0.25$2.00$0.688$0.0065
12GPT-4.1 miniOpenAI$0.40$1.60$0.70$0.0072
13Mistral Large 3Mistral$0.50$1.50$0.75$0.008
14Gemini 3.5 Flash-LiteGoogle$0.30$2.50$0.85$0.008
15Gemini 2.5 FlashGoogle$0.30$2.50$0.85$0.008

Methodology. Blended cost = 0.75 x input rate + 0.25 x output rate per 1M tokens (a 3:1 input:output mix, typical for chat and RAG). Prices are official provider list rates for the standard API tier, verified 2026-08-11; batch discounts, caching, and tiered long-context rates are not applied. Context windows and tiers are on each model page and the full pricing table.

Keep exploring

Compare any two models head-to-head in model comparisons, see which models fit the most text in best long-context LLMs, or model your own workload in the token cost analyzer. Price cuts show up daily in the pricing change tracker.

FAQ

Cheapest LLM API FAQ

What is the cheapest LLM API right now?

As of 2026-08-11, GPT-5 nano from OpenAI is the cheapest verified-price model in our catalog at $0.05 per 1M input tokens and $0.40 per 1M output tokens — a blended $0.138 per 1M tokens on a typical 3:1 input:output mix.

How is "cheapest" calculated here?

Models are ranked by blended cost per 1M tokens assuming a 3:1 input:output token ratio, the typical shape of chat and RAG workloads. Both input and output rates come from official provider pricing, re-verified every day. Rankings shift automatically when providers change prices.

Is a cheaper LLM always the better choice?

No. Cheap tiers trade capability for cost: they are ideal for classification, extraction, summarization, and high-volume simple tasks, while harder reasoning usually justifies a premium tier. Many teams route easy requests to a cheap model and escalate hard ones, cutting bills substantially.

How much does the cheapest model cost per request?

A request with 10,000 input tokens and 2,000 output tokens costs about $0.0013 on GPT-5 nano. At one million such requests per month, that is roughly $1,300.