Direct answer: use GPT-6 Luna for high-volume tasks that code can check, and DeepSeek V4.1 Flash for agents with tools. Use Gemini 3.8 Flash for images, audio, and video, and ClaudeClaude ProThe paid subscription for the Claude AI assistant from Anthropic. It unlocks connectors to outside tools.Open the glossary Haiku 4.5 for short answers that must be careful. Luna, DeepSeek, and Gemini sit within 4 points on the index score, but their cost per task spans 18x.

Main condition: this split holds when you measure the cost per successful task and turn on reasoning for arithmetic. In our test, Luna, DeepSeek, and Haiku failed invoice math 0 of 5 times on default settings and passed 3 of 3 at high effort. Limit: the Gemini 3.8 Flash price doubles on 1 January 2027, and the DeepSeek privacy policy says it stores personal data in China.

On 23 September 2026 we read the official pricing pages of OpenAI, DeepSeek, Google, and Anthropic. We also read 4 model pages and 9 evaluation pages on Artificial Analysis. On the same day we tested all 4 models with the same prompts and dummy data: 40 calls across 5 tasks, then 21 follow-up calls.

The prices look alike, the real cost does not

These 4 models often appear on "cheap model" lists. Output prices per 1M tokens sit between $0.50 and $5. That 10x gap looks small until you multiply it by millions of messages per month.

From the Claude family, we picked Haiku 4.5 because its price sits closest to the other 3 models. Claude Sonnet 5 costs $2 input and $10 output, so Sonnet 5 belongs to the price class above.

The second problem is less visible. A model that writes many reasoning tokens can cost more per task than a model with a higher token price. So we compare 3 things: price per token, cost per task from an independent test, and our own test.

For the class above, read our comparison of GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5. This article covers the volume class: the models you use for thousands of tasks per day.

How each vendor bills

All 4 vendors bill per token, but their extra rules differ. These extra rules make 2 models with the same price produce different bills.

GPT-6 Luna

The GPT-6 Luna model page lists $0.10 input, $0.01 cached input, $0.125 cache write, and $0.50 output per 1M tokens. A prompt above 272,000 input tokens costs 2x for input and 1.5x for output on the full request. Batch and Flex cost 50% of the standard price.

Luna has 6 effort levels: none, low, medium, high, xhigh, and max. The default is medium. The context window is 1,050,000 tokens, maximum output is 128,000 tokens, and the knowledge cutoff is 18 May 2026.

DeepSeek V4.1 Flash

The DeepSeek pricing page uses the model name deepseek-flash for DeepSeek-V4.1-Flash. The peak price is $0.30 input and $1.20 output. The off-peak price is half: $0.15 and $0.60.

Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. In Western Indonesia Time (WIB), that is 08.00 to 11.00 and 13.00 to 17.00. So DeepSeek peak hours fall exactly on Indonesian office hours.

A 24-hour timeline in WIB: DeepSeek peak hours run 08.00 to 11.00 and 13.00 to 17.00 on weekdays, and off-peak runs all day on weekends and Chinese public holidays
Amber blocks are peak hours at 2x the price. Move batches to the night, to 11.00 to 13.00, or to the weekend.

The old name deepseek-v4-flash is still accepted, but that model is retired. DeepSeek serves those requests with V4.1 Flash at the Flash price. The DeepSeek APIAPIThe official door 2 systems use to exchange data, without anybody copying it by hand.Open the glossary accepts the OpenAI and Anthropic formats, so you only change the base URL and the model name.

Gemini 3.8 Flash

The Gemini API pricing page lists $0.75 input and $3.75 output through 31 December 2026. From 1 January 2027, the price becomes $1.50 and $7.50. Batch and Flex cost 50%, and Priority costs 1.8x.

The Gemini 3.8 Flash model page lists text, image, video, audio, and PDF input, with limits of 1,048,576 input tokens and 65,536 output tokens. Thinking supports low, medium, and high. The minimal level returns an error.

Gemini has a free tier for 3.8 Flash. But Google states that it uses free-tier content to improve its products. Grounding with Google Search gives 5,000 free requests per month, then $14 per 1,000 requests.

Claude Haiku 4.5

The Claude pricing page lists Haiku 4.5 at $1 input and $5 output per 1M tokens. A cache read costs $0.10, or 0.1x the input price, and a 5-minute cache write costs 1.25x. The Batch API gives a 50% discount, so the batch price is $0.50 and $2.50.

The Claude models overview lists a 200,000-token context window, 64,000 maximum output tokens, and a reliable knowledge cutoff of February 2025. Haiku 4.5 uses extended thinking, not effort levels. Claude web search costs $10 per 1,000 searches.

Bar chart of input and output prices per 1M tokens for GPT-6 Luna, DeepSeek V4.1 Flash, Gemini 3.8 Flash, and Claude Haiku 4.5, with a box for DeepSeek off-peak prices, Gemini 2027 prices, and cache reads
On output, Haiku 4.5 costs 10x Luna and Gemini 3.8 Flash costs 7.5x Luna. The right box holds the prices that change by hour and date.

Specs and prices side by side

This table uses numbers from the official pages, read on 23 September 2026. Prices are in USD per 1M tokens on standard processing.

ItemGPT-6 LunaDeepSeek V4.1 FlashGemini 3.8 FlashClaude Haiku 4.5
API IDgpt-6-lunadeepseek-flashgemini-3.8-flashclaude-haiku-4-5-20251001
Input / output$0.10 / $0.50$0.30 / $1.20 (off-peak $0.15 / $0.60)$0.75 / $3.75 (2027: $1.50 / $7.50)$1 / $5
Cache read$0.01$0.006 (off-peak $0.003)$0.075$0.10
Context / max output1,050,000 / 128,0001M / 384,0001,048,576 / 65,536200,000 / 64,000
Non-text inputImagesImagesImages, video, audio, PDFImages
Reasoning6 levels, default mediumThinking and non-thinking, default thinkinglow, medium, highExtended thinking
Bulk discountBatch and Flex 50%Off-peak 50%Batch and Flex 50%Batch 50%
Model weightsClosedOpen, MIT licenseClosedClosed

The model weights row comes from the DeepSeek V4.1 Flash page on Artificial Analysis: 552 billion total parameters and 16 billion active parameters. So you can run DeepSeek on your own server, but that size needs several large GPUs.

Capability: independent tests on the same harness

Vendor benchmarks are hard to compare because their methods differ. So we use Artificial Analysis, which tests every model on the same harness. The tested variants are Luna max, DeepSeek max, Gemini high, and Haiku 4.5 reasoning.

Index score and cost per task

On the Intelligence Index v4.3.2, Gemini 3.8 Flash scores 41, DeepSeek 39, Luna 37, and Haiku 4.5 17. Haiku 4.5 launched in October 2025, so Haiku is the oldest model in this table.

The cost per index task is much wider: Luna $0.07, Haiku $0.21, DeepSeek $0.27, and Gemini $1.24. Haiku uses the fewest output tokens, 78 million for the full index, but its token price is 10x Luna.

Scatter plot of Intelligence Index score against cost per task on a logarithmic cost axis for GPT-6 Luna, DeepSeek V4.1 Flash, Gemini 3.8 Flash, and Claude Haiku 4.5
Luna, DeepSeek, and Gemini sit in the 37 to 41 point range. Their cost per task runs from $0.07 to $1.24.

Profile per field

The index score hides differences per field. The next table holds 8 evaluations that sit closest to business work. The numbers come from the Artificial Analysis evaluation pages on 23 September 2026.

EvaluationLunaDeepSeekGeminiHaiku 4.5
Terminal-Bench 4.0 (agentAI agentAn AI program that performs work steps by itself, for example reading a message, drafting a reply, and recording the result.Open the glossary coding)12.6%26.8%19.7%0.0%
GDPval-AA v2.1 (real work, Elo)1,3671,6001,412719
AutomationBench-AA (SaaS flows)53.2%68.9%59.9%3.2%
AA-LCR v1.1 (long context)83.3%84.0%81.3%74.3%
SciCode (scientific code)54.6%51.9%56.6%42.2%
MMMU-Pro (images)76%77%86%59%
AA-Omniscience accuracy44%46%55%18%
Hallucination when it does not know (low = good)77%96%55%27%
Output speed (tokens per second)157227276109
Time per index task (minutes)5.34.94.12.2
Color table of the capability profile of 4 models on 11 Artificial Analysis measures, with green cells for the strongest value and red cells for the weakest value in each row
DeepSeek leads agent coding, real work, SaaS flows, and long context. Gemini leads images, knowledge, and speed. Luna leads cost per task, and Haiku leads low hallucination.

For 𝜏³-Banking, which tests a bank customer service agent, Artificial Analysis lists only 2 of the 4 models so far. Gemini 3.8 Flash scores 44.9%, and Haiku 4.5 scores 9.3%.

Accuracy is not the same as honesty

AA-Omniscience separates 2 things: how many answers are right, and how often a model guesses when it does not know. DeepSeek is right on 46% of questions, but answers wrongly on 96% of the questions it does not know. Haiku 4.5 is right on only 18%, but admits it does not know most often.

Two side-by-side bar charts: AA-Omniscience accuracy and hallucination rate for 4 models, with a note on how to read both numbers
Gemini has the top accuracy. Haiku has the lowest hallucination. DeepSeek almost always guesses.

For a customer bot, the hallucination number matters more than the index score. One invented answer about price or warranty can cost a sale. So add a knowledge base and a "say you do not know" rule on every model, not only on the weak ones.

What you need

  • API access to at least 2 vendors, or 1 gateway that forwards to several vendors.
  • A list of 20 real tasks from 1 workflow, with the correct answers.
  • 1 code validator per task: a JSON schemaSchemaExtra description inside page code that tells a search engine what the page is, for example an article, a service, or a question and answer.Open the glossary, a recompute, or a list of facts the model may state.
  • A log per call that stores model, effort, tokens, time, validator result, and cost.
  • A written decision on personal data: which data may leave Indonesia, and to which vendor.
  • A rupiah exchange rate for reports. We use the Bank Indonesia JISDOR rate of 22 September 2026: Rp17,883 per USD.

Step 1: Map each task to 1 candidate model

Start from the task, not from the model. Answer the 4 questions in the diagram from the top, then stop at the first "yes".

A 4-question decision tree: multimodal input leads to Gemini 3.8 Flash, short factual answers to customers lead to Claude Haiku 4.5, agents with tools lead to DeepSeek V4.1 Flash, and high volume that code can check leads to GPT-6 Luna
Each right-hand box holds 1 number that supports the pick. If every answer is "no", test the 2 cheapest candidates.
  1. Does the input hold images, audio, or video? Pick Gemini 3.8 Flash. Of these 4 models, only Gemini accepts audio and video according to the official model pages.
  2. Short factual answers to customers? Pick Haiku 4.5. Its hallucination is the lowest, 27%, but its accuracy is only 18%, so add a knowledge base.
  3. Multi-step agent with tools? Pick DeepSeek V4.1 Flash if the data may be processed abroad. If not, pick Gemini 3.8 Flash.
  4. High volume, and code can check the result? Pick Luna. Its cost per task is the lowest, $0.07.

Check: every task on your list of 20 gets 1 candidate. This split is a Rama Digital recommendation that comes from the Artificial Analysis numbers and the official pages above.

Step 2: Calculate the monthly cost in rupiah

Multiply the tokens per month by the price per token. The next simulation uses dummy data, the same tokens for every model, no cache, and no reasoning tokens. The goal is to see the list-price gap.

Workload per monthLunaDeepSeekGeminiHaiku 4.5
CS replies: 20,000 chats, 120M input, 20M outputRp393,426Rp1,072,980Rp2,950,695Rp3,934,260
Order extraction: 50,000 messages, 40M input, 7.5M outputRp138,593Rp375,543Rp1,039,449Rp1,385,933
Document summaries: 500 documents, 75M input, 1.5M outputRp147,535Rp434,557Rp1,106,511Rp1,475,348
Bar chart of monthly token cost in rupiah for 3 simulated workloads and 4 models, with a box for DeepSeek off-peak and Gemini 2027 prices on CS replies
At equal tokens, Haiku 4.5 costs 10x Luna. The right box shows the CS cost when DeepSeek runs off-peak and when Gemini uses the 2027 price.

Two variants change this table. CS replies on DeepSeek at off-peak hours drop to Rp536,490. CS replies on Gemini at the 2027 price rise to Rp5,901,390.

Check: recompute your numbers with the reasoning tokens from your log. The full formula with cache and the long-context premium is in our guide to API cost per task.

Step 3: Test every candidate with the same prompt

Send the exact same prompt to every candidate, then grade the result with code. We did this on 23 September 2026 through an internal gateway with each model on default settings. Each task ran 2 times.

  • Invoice math: 3 items, a discount on 1 item, 11% VAT, and an admin fee outside VAT. Correct answer: Rp7,249,250.
  • JSON extraction: a WhatsApp chat that corrects the quantity from 24 to 30 pieces.
  • CS reply: 3 sentences at most, the exact price, no emoji, and 1 closing question.
  • Rupiah code: a formatRupiah function tested on 7 cases, including negative numbers.
  • Fake feature: a question about a Meta Business SuiteBusiness ManagerThe place where Meta holds your business ad assets: ad accounts, Pages, pixels, and people access.Open the glossary feature that does not exist. A model passes if it states that the feature does not exist.
Result table of the Rama Digital test on 23 September 2026 for 4 models and 5 tasks, plus a follow-up panel for invoice math on default settings and high effort
Gemini passes 10 of 10. Luna, DeepSeek, and Haiku fail invoice math on default settings, then pass 3 of 3 at high effort.
ModelInvoiceJSONCS replyCodeFake featureTotalMedian time
GPT-6 Luna0/22/22/22/21/27/102.8 seconds
DeepSeek V4.1 Flash0/22/22/22/22/28/101.2 seconds
Gemini 3.8 Flash2/22/22/22/22/210/106.5 seconds
Claude Haiku 4.50/22/21/22/22/27/102.2 seconds

The most important finding is in the invoice task. On default settings, DeepSeek ran with 0 reasoning tokens and answered with random totals between Rp8.9 million and Rp14 million. With reasoning_effort: "high", DeepSeek used 192 to 277 reasoning tokens and passed 3 of 3.

Luna and Haiku show the same pattern: 0 of 5 on default settings, 3 of 3 at high effort. Gemini thinks by default with 532 to 714 reasoning tokens, so Gemini passes 5 of 5 without a settings change.

Two small notes. Haiku failed 1 CS reply because the greeting "Halo kak!" became a 4th sentence. Luna invented steps for the fake feature in 1 of 2 runs, while the other 3 models refused both times.

The limit of this test: our gateway uses subscription connections, not a direct paid API, so response times include gateway delay. Code blocks from Haiku came back unclosed through the gateway, and we graded the code itself. We think this is a gateway effect, not a model effect.

Step 4: Set up routing and fallback

One model for every task is always too expensive or too weak. Set up a rule router that reads the task class from Step 1, then a validator that decides when a task moves up.

A 7-point flow diagram: task arrives, rule router, cheap model, validator, send on pass, retry at high effort on fail, then escalate to an upper-class model or a person
Point 4 decides every branch. Each task keeps 1 log row with the 6 fields in the bottom box.
  1. Send the task to the candidate from Step 1 with default effort.
  2. Run the validator. If it passes, send the result.
  3. If it fails, retry 1 time at high effort. Our test shows that this step fixed every invoice math failure.
  4. If it fails again, escalate to an upper-class model or to a person.

Check: the log shows the share of tasks that move up. If more than 10% of tasks move up, that cheap candidate is wrong for that task class. The 10% threshold is a Rama Digital recommendation, not a vendor number.

Step 5 (optional): Schedule price and version rechecks

Cheap model prices change faster than large model prices. Put the dates from the diagram on the team calendar, then redo Step 2 each time a date passes.

Price and version timeline from October 2025 to January 2027: Haiku 4.5 launch, DeepSeek V4.1 and Gemini 3.8, GPT-6 Luna launch, a recheck date, and the end of the Gemini promo
The date with the most impact is 1 January 2027. On that date the Gemini 3.8 Flash input and output prices double.

Check: a reminder for 1 December 2026 is on the calendar, so you have 1 month to move Gemini traffic if needed. Luna launch details are in our GPT-6 Luna price and spec notes.

Worked example: 1 day at a dummy T-shirt shop

A simulation with dummy data. A fictional T-shirt shop receives 6 tasks on Thursday, 1 October 2026. Tokens are assumptions, and costs use standard prices at Rp17,883 per USD.

Time WIBIncoming taskRouter decisionWhat the log recordsResult and cost
08.15WhatsApp order: 30 navy polo shirts, ship to SlemanExtraction, GPT-6 Luna, default effort800 input tokens, 150 output tokens, JSON passes the schemaOrder enters the system, Rp2.77
09.40Question: "can I pay cash on delivery?"Factual CS, Haiku 4.5 with a knowledge base1,500 input, 250 output, price matches the catalogReply sent, Rp49.18
10.05Calculate an invoice with 3 items and 11% VATArithmetic, GPT-6 Luna, high effort300 input, 400 output including reasoning, recompute matchesInvoice sent, Rp4.11
13.30Photo of a transfer receiptImage, Gemini 3.8 Flash1,500 input, 200 output, amount matches the invoicePayment recorded, Rp33.53
14.10Complaint about a wrong sizeRisky CS, Haiku 4.5 with a knowledge baseHaiku promises a refund outside the rules, the validator rejects it, the task moves to an adminAdmin replies, Rp53.65 for 1 call
22.00Batch of 500 product descriptionsHigh volume without personal data, DeepSeek off-peak200,000 input, 150,000 output, 8 descriptions retried500 descriptions ready, Rp2,146

The total token cost for that day is Rp2,289. The cheapest decision is to move the batch to 22.00, because the DeepSeek off-peak price is half. The safest decision is to hand the complaint to an admin after the validator rejects the refund promise.

Checklist before you move traffic

  1. Write 20 real tasks per workflow with correct answers. Owner: process owner. Evidence: a dated task file.
  2. Write 1 code validator per task. Owner: developer. Evidence: passing validator tests.
  3. Test every candidate with the same prompt at default and high effort. Owner: developer. Evidence: a log per call.
  4. Calculate the cost per successful task in rupiah from the log. Owner: finance team. Evidence: a cost sheet.
  5. Decide which data may go to the DeepSeek API. Owner: business owner. Evidence: a personal data decision note.
  6. Set up routing, 1 retry at high effort, and escalation. Owner: developer. Evidence: router configuration.
  7. Schedule DeepSeek batches outside WIB peak hours. Owner: operations team. Evidence: a cron schedule.
  8. Set price reminders for 23 October and 1 December 2026. Owner: operations team. Evidence: a calendar invite.
  9. Stop testing when 1 candidate passes 19 of 20 tasks at the lowest cost per successful task. If no candidate passes 18 of 20, use an upper-class model for that task class.

When to pick which model

This table sums up the facts above per model. Each "pick" and "avoid" cell comes from a number we already stated with its source.

ModelPick whenAvoid whenSupporting facts
GPT-6 LunaHigh volume, code can check the result, context below 272,000 tokensArithmetic on default effort, or factual answers without a knowledge baseCost per task $0.07, AA-LCR 83.3%, hallucination 77%
DeepSeek V4.1 FlashAgents and SaaS flows, off-peak batches, or you need open weightsPersonal data through the API, or factual answers without a knowledge baseAutomationBench 68.9%, GDPval 1,600, hallucination 96%, MIT license
Gemini 3.8 FlashImages, audio, video, general knowledge, and speedA tight budget after 1 January 2027Index 41, MMMU-Pro 86%, 276 tokens per second, cost per task $1.24
Claude Haiku 4.5Short answers that must be careful, with a knowledge baseMulti-step agents, arithmetic without extended thinking, or context above 200,000 tokensHallucination 27%, accuracy 18%, 2.2 minutes per task, AutomationBench 3.2%

The DeepSeek privacy policy states that DeepSeek processes and stores personal data in the People's Republic of China. If your internal rules forbid that, run the DeepSeek open weights on your own server or use another model for customer data.

Rama Digital recommendation: make Luna the default lane for checkable tasks, Gemini 3.8 Flash the multimodal lane, and Haiku 4.5 the customer answer lane. This holds when you have code validators and a knowledge base. Pick DeepSeek when your task is an agent without personal data, mainly at off-peak hours.

Frequently asked questions

Which model is the cheapest per task? GPT-6 Luna. Artificial Analysis lists $0.07 per index task for Luna, $0.21 for Haiku 4.5, and $0.27 for DeepSeek. Luna also has the lowest token price, $0.10 input and $0.50 output.

Why is Claude Sonnet 5 not in this comparison? Sonnet 5 costs $2 input and $10 output, 2x the Haiku 4.5 price and 20x the Luna output price. Haiku 4.5 is the Claude model with the price closest to the other 3 models, so Sonnet 5 belongs to the price class above.

Is DeepSeek Flash safe for customer data? The DeepSeek privacy policy states that personal data is processed and stored in the People's Republic of China. Check your personal data law and your contracts. The alternative is the DeepSeek open weights on your own server.

Why do cheap models fail invoice math? On default settings, Luna, DeepSeek, and Haiku answer without enough reasoning for multi-step arithmetic. In our test, high effort fixed all 3 to 3 of 3.

Will the Gemini 3.8 Flash price go up? Yes. The Gemini API pricing page lists $0.75 input and $3.75 output through 31 December 2026, then $1.50 and $7.50 from 1 January 2027.

Is Claude Haiku 4.5 still worth using? Yes, for short answers that must be careful and fast. Haiku 4.5 has a 27% hallucination rate, the lowest of the 4 models, and 2.2 minutes per index task. For multi-step agents, pick DeepSeek or Gemini, because Haiku scores only 3.2% on AutomationBench.

Next step

The limit that still applies: independent scores and our test are only a start, because your prompts, data, and risks differ. Test 20 real tasks before you move traffic or sign a budget.

If you want us to map the model lanes for your workflow, open AI Workflow Audit. To talk it through first, pick a 30-minute first consultation slot.

Sources