Direct answer: use GPT-6 Luna for high-volume tasks that code can check, and DeepSeek V4.1 Flash for agents with tools. Use Gemini 3.8 Flash for images, audio, and video, and ClaudeClaude ProThe paid subscription for the Claude AI assistant from Anthropic. It unlocks connectors to outside tools.Open the glossary Haiku 4.5 for short answers that must be careful. Luna, DeepSeek, and Gemini sit within 4 points on the index score, but their cost per task spans 18x.
Main condition: this split holds when you measure the cost per successful task and turn on reasoning for arithmetic. In our test, Luna, DeepSeek, and Haiku failed invoice math 0 of 5 times on default settings and passed 3 of 3 at high effort. Limit: the Gemini 3.8 Flash price doubles on 1 January 2027, and the DeepSeek privacy policy says it stores personal data in China.
On 23 September 2026 we read the official pricing pages of OpenAI, DeepSeek, Google, and Anthropic. We also read 4 model pages and 9 evaluation pages on Artificial Analysis. On the same day we tested all 4 models with the same prompts and dummy data: 40 calls across 5 tasks, then 21 follow-up calls.
The prices look alike, the real cost does not
These 4 models often appear on "cheap model" lists. Output prices per 1M tokens sit between $0.50 and $5. That 10x gap looks small until you multiply it by millions of messages per month.
From the Claude family, we picked Haiku 4.5 because its price sits closest to the other 3 models. Claude Sonnet 5 costs $2 input and $10 output, so Sonnet 5 belongs to the price class above.
The second problem is less visible. A model that writes many reasoning tokens can cost more per task than a model with a higher token price. So we compare 3 things: price per token, cost per task from an independent test, and our own test.
For the class above, read our comparison of GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5. This article covers the volume class: the models you use for thousands of tasks per day.
How each vendor bills
All 4 vendors bill per token, but their extra rules differ. These extra rules make 2 models with the same price produce different bills.
GPT-6 Luna
The GPT-6 Luna model page lists $0.10 input, $0.01 cached input, $0.125 cache write, and $0.50 output per 1M tokens. A prompt above 272,000 input tokens costs 2x for input and 1.5x for output on the full request. Batch and Flex cost 50% of the standard price.
Luna has 6 effort levels: none, low, medium, high, xhigh, and max. The default is medium. The context window is 1,050,000 tokens, maximum output is 128,000 tokens, and the knowledge cutoff is 18 May 2026.
DeepSeek V4.1 Flash
The DeepSeek pricing page uses the model name deepseek-flash for DeepSeek-V4.1-Flash. The peak price is $0.30 input and $1.20 output. The off-peak price is half: $0.15 and $0.60.
Peak hours are 01:00 to 04:00 and 06:00 to 10:00 UTC, Monday to Friday. In Western Indonesia Time (WIB), that is 08.00 to 11.00 and 13.00 to 17.00. So DeepSeek peak hours fall exactly on Indonesian office hours.

The old name deepseek-v4-flash is still accepted, but that model is retired. DeepSeek serves those requests with V4.1 Flash at the Flash price. The DeepSeek APIAPIThe official door 2 systems use to exchange data, without anybody copying it by hand.Open the glossary accepts the OpenAI and Anthropic formats, so you only change the base URL and the model name.
Gemini 3.8 Flash
The Gemini API pricing page lists $0.75 input and $3.75 output through 31 December 2026. From 1 January 2027, the price becomes $1.50 and $7.50. Batch and Flex cost 50%, and Priority costs 1.8x.
The Gemini 3.8 Flash model page lists text, image, video, audio, and PDF input, with limits of 1,048,576 input tokens and 65,536 output tokens. Thinking supports low, medium, and high. The minimal level returns an error.
Gemini has a free tier for 3.8 Flash. But Google states that it uses free-tier content to improve its products. Grounding with Google Search gives 5,000 free requests per month, then $14 per 1,000 requests.
Claude Haiku 4.5
The Claude pricing page lists Haiku 4.5 at $1 input and $5 output per 1M tokens. A cache read costs $0.10, or 0.1x the input price, and a 5-minute cache write costs 1.25x. The Batch API gives a 50% discount, so the batch price is $0.50 and $2.50.
The Claude models overview lists a 200,000-token context window, 64,000 maximum output tokens, and a reliable knowledge cutoff of February 2025. Haiku 4.5 uses extended thinking, not effort levels. Claude web search costs $10 per 1,000 searches.

Specs and prices side by side
This table uses numbers from the official pages, read on 23 September 2026. Prices are in USD per 1M tokens on standard processing.
| Item | GPT-6 Luna | DeepSeek V4.1 Flash | Gemini 3.8 Flash | Claude Haiku 4.5 |
|---|---|---|---|---|
| API ID | gpt-6-luna | deepseek-flash | gemini-3.8-flash | claude-haiku-4-5-20251001 |
| Input / output | $0.10 / $0.50 | $0.30 / $1.20 (off-peak $0.15 / $0.60) | $0.75 / $3.75 (2027: $1.50 / $7.50) | $1 / $5 |
| Cache read | $0.01 | $0.006 (off-peak $0.003) | $0.075 | $0.10 |
| Context / max output | 1,050,000 / 128,000 | 1M / 384,000 | 1,048,576 / 65,536 | 200,000 / 64,000 |
| Non-text input | Images | Images | Images, video, audio, PDF | Images |
| Reasoning | 6 levels, default medium | Thinking and non-thinking, default thinking | low, medium, high | Extended thinking |
| Bulk discount | Batch and Flex 50% | Off-peak 50% | Batch and Flex 50% | Batch 50% |
| Model weights | Closed | Open, MIT license | Closed | Closed |
The model weights row comes from the DeepSeek V4.1 Flash page on Artificial Analysis: 552 billion total parameters and 16 billion active parameters. So you can run DeepSeek on your own server, but that size needs several large GPUs.
Capability: independent tests on the same harness
Vendor benchmarks are hard to compare because their methods differ. So we use Artificial Analysis, which tests every model on the same harness. The tested variants are Luna max, DeepSeek max, Gemini high, and Haiku 4.5 reasoning.
Index score and cost per task
On the Intelligence Index v4.3.2, Gemini 3.8 Flash scores 41, DeepSeek 39, Luna 37, and Haiku 4.5 17. Haiku 4.5 launched in October 2025, so Haiku is the oldest model in this table.
The cost per index task is much wider: Luna $0.07, Haiku $0.21, DeepSeek $0.27, and Gemini $1.24. Haiku uses the fewest output tokens, 78 million for the full index, but its token price is 10x Luna.

Profile per field
The index score hides differences per field. The next table holds 8 evaluations that sit closest to business work. The numbers come from the Artificial Analysis evaluation pages on 23 September 2026.
| Evaluation | Luna | DeepSeek | Gemini | Haiku 4.5 |
|---|---|---|---|---|
| Terminal-Bench 4.0 (agentAI agentAn AI program that performs work steps by itself, for example reading a message, drafting a reply, and recording the result.Open the glossary coding) | 12.6% | 26.8% | 19.7% | 0.0% |
| GDPval-AA v2.1 (real work, Elo) | 1,367 | 1,600 | 1,412 | 719 |
| AutomationBench-AA (SaaS flows) | 53.2% | 68.9% | 59.9% | 3.2% |
| AA-LCR v1.1 (long context) | 83.3% | 84.0% | 81.3% | 74.3% |
| SciCode (scientific code) | 54.6% | 51.9% | 56.6% | 42.2% |
| MMMU-Pro (images) | 76% | 77% | 86% | 59% |
| AA-Omniscience accuracy | 44% | 46% | 55% | 18% |
| Hallucination when it does not know (low = good) | 77% | 96% | 55% | 27% |
| Output speed (tokens per second) | 157 | 227 | 276 | 109 |
| Time per index task (minutes) | 5.3 | 4.9 | 4.1 | 2.2 |

For 𝜏³-Banking, which tests a bank customer service agent, Artificial Analysis lists only 2 of the 4 models so far. Gemini 3.8 Flash scores 44.9%, and Haiku 4.5 scores 9.3%.
Accuracy is not the same as honesty
AA-Omniscience separates 2 things: how many answers are right, and how often a model guesses when it does not know. DeepSeek is right on 46% of questions, but answers wrongly on 96% of the questions it does not know. Haiku 4.5 is right on only 18%, but admits it does not know most often.

For a customer bot, the hallucination number matters more than the index score. One invented answer about price or warranty can cost a sale. So add a knowledge base and a "say you do not know" rule on every model, not only on the weak ones.
What you need
- API access to at least 2 vendors, or 1 gateway that forwards to several vendors.
- A list of 20 real tasks from 1 workflow, with the correct answers.
- 1 code validator per task: a JSON schemaSchemaExtra description inside page code that tells a search engine what the page is, for example an article, a service, or a question and answer.Open the glossary, a recompute, or a list of facts the model may state.
- A log per call that stores model, effort, tokens, time, validator result, and cost.
- A written decision on personal data: which data may leave Indonesia, and to which vendor.
- A rupiah exchange rate for reports. We use the Bank Indonesia JISDOR rate of 22 September 2026: Rp17,883 per USD.
Step 1: Map each task to 1 candidate model
Start from the task, not from the model. Answer the 4 questions in the diagram from the top, then stop at the first "yes".

- Does the input hold images, audio, or video? Pick Gemini 3.8 Flash. Of these 4 models, only Gemini accepts audio and video according to the official model pages.
- Short factual answers to customers? Pick Haiku 4.5. Its hallucination is the lowest, 27%, but its accuracy is only 18%, so add a knowledge base.
- Multi-step agent with tools? Pick DeepSeek V4.1 Flash if the data may be processed abroad. If not, pick Gemini 3.8 Flash.
- High volume, and code can check the result? Pick Luna. Its cost per task is the lowest, $0.07.
Check: every task on your list of 20 gets 1 candidate. This split is a Rama Digital recommendation that comes from the Artificial Analysis numbers and the official pages above.
Step 2: Calculate the monthly cost in rupiah
Multiply the tokens per month by the price per token. The next simulation uses dummy data, the same tokens for every model, no cache, and no reasoning tokens. The goal is to see the list-price gap.
| Workload per month | Luna | DeepSeek | Gemini | Haiku 4.5 |
|---|---|---|---|---|
| CS replies: 20,000 chats, 120M input, 20M output | Rp393,426 | Rp1,072,980 | Rp2,950,695 | Rp3,934,260 |
| Order extraction: 50,000 messages, 40M input, 7.5M output | Rp138,593 | Rp375,543 | Rp1,039,449 | Rp1,385,933 |
| Document summaries: 500 documents, 75M input, 1.5M output | Rp147,535 | Rp434,557 | Rp1,106,511 | Rp1,475,348 |

Two variants change this table. CS replies on DeepSeek at off-peak hours drop to Rp536,490. CS replies on Gemini at the 2027 price rise to Rp5,901,390.
Check: recompute your numbers with the reasoning tokens from your log. The full formula with cache and the long-context premium is in our guide to API cost per task.
Step 3: Test every candidate with the same prompt
Send the exact same prompt to every candidate, then grade the result with code. We did this on 23 September 2026 through an internal gateway with each model on default settings. Each task ran 2 times.
- Invoice math: 3 items, a discount on 1 item, 11% VAT, and an admin fee outside VAT. Correct answer: Rp7,249,250.
- JSON extraction: a WhatsApp chat that corrects the quantity from 24 to 30 pieces.
- CS reply: 3 sentences at most, the exact price, no emoji, and 1 closing question.
- Rupiah code: a
formatRupiahfunction tested on 7 cases, including negative numbers. - Fake feature: a question about a Meta Business SuiteBusiness ManagerThe place where Meta holds your business ad assets: ad accounts, Pages, pixels, and people access.Open the glossary feature that does not exist. A model passes if it states that the feature does not exist.

| Model | Invoice | JSON | CS reply | Code | Fake feature | Total | Median time |
|---|---|---|---|---|---|---|---|
| GPT-6 Luna | 0/2 | 2/2 | 2/2 | 2/2 | 1/2 | 7/10 | 2.8 seconds |
| DeepSeek V4.1 Flash | 0/2 | 2/2 | 2/2 | 2/2 | 2/2 | 8/10 | 1.2 seconds |
| Gemini 3.8 Flash | 2/2 | 2/2 | 2/2 | 2/2 | 2/2 | 10/10 | 6.5 seconds |
| Claude Haiku 4.5 | 0/2 | 2/2 | 1/2 | 2/2 | 2/2 | 7/10 | 2.2 seconds |
The most important finding is in the invoice task. On default settings, DeepSeek ran with 0 reasoning tokens and answered with random totals between Rp8.9 million and Rp14 million. With reasoning_effort: "high", DeepSeek used 192 to 277 reasoning tokens and passed 3 of 3.
Luna and Haiku show the same pattern: 0 of 5 on default settings, 3 of 3 at high effort. Gemini thinks by default with 532 to 714 reasoning tokens, so Gemini passes 5 of 5 without a settings change.
Two small notes. Haiku failed 1 CS reply because the greeting "Halo kak!" became a 4th sentence. Luna invented steps for the fake feature in 1 of 2 runs, while the other 3 models refused both times.
The limit of this test: our gateway uses subscription connections, not a direct paid API, so response times include gateway delay. Code blocks from Haiku came back unclosed through the gateway, and we graded the code itself. We think this is a gateway effect, not a model effect.
Step 4: Set up routing and fallback
One model for every task is always too expensive or too weak. Set up a rule router that reads the task class from Step 1, then a validator that decides when a task moves up.

- Send the task to the candidate from Step 1 with default effort.
- Run the validator. If it passes, send the result.
- If it fails, retry 1 time at high effort. Our test shows that this step fixed every invoice math failure.
- If it fails again, escalate to an upper-class model or to a person.
Check: the log shows the share of tasks that move up. If more than 10% of tasks move up, that cheap candidate is wrong for that task class. The 10% threshold is a Rama Digital recommendation, not a vendor number.
Step 5 (optional): Schedule price and version rechecks
Cheap model prices change faster than large model prices. Put the dates from the diagram on the team calendar, then redo Step 2 each time a date passes.

Check: a reminder for 1 December 2026 is on the calendar, so you have 1 month to move Gemini traffic if needed. Luna launch details are in our GPT-6 Luna price and spec notes.
Worked example: 1 day at a dummy T-shirt shop
A simulation with dummy data. A fictional T-shirt shop receives 6 tasks on Thursday, 1 October 2026. Tokens are assumptions, and costs use standard prices at Rp17,883 per USD.
| Time WIB | Incoming task | Router decision | What the log records | Result and cost |
|---|---|---|---|---|
| 08.15 | WhatsApp order: 30 navy polo shirts, ship to Sleman | Extraction, GPT-6 Luna, default effort | 800 input tokens, 150 output tokens, JSON passes the schema | Order enters the system, Rp2.77 |
| 09.40 | Question: "can I pay cash on delivery?" | Factual CS, Haiku 4.5 with a knowledge base | 1,500 input, 250 output, price matches the catalog | Reply sent, Rp49.18 |
| 10.05 | Calculate an invoice with 3 items and 11% VAT | Arithmetic, GPT-6 Luna, high effort | 300 input, 400 output including reasoning, recompute matches | Invoice sent, Rp4.11 |
| 13.30 | Photo of a transfer receipt | Image, Gemini 3.8 Flash | 1,500 input, 200 output, amount matches the invoice | Payment recorded, Rp33.53 |
| 14.10 | Complaint about a wrong size | Risky CS, Haiku 4.5 with a knowledge base | Haiku promises a refund outside the rules, the validator rejects it, the task moves to an admin | Admin replies, Rp53.65 for 1 call |
| 22.00 | Batch of 500 product descriptions | High volume without personal data, DeepSeek off-peak | 200,000 input, 150,000 output, 8 descriptions retried | 500 descriptions ready, Rp2,146 |
The total token cost for that day is Rp2,289. The cheapest decision is to move the batch to 22.00, because the DeepSeek off-peak price is half. The safest decision is to hand the complaint to an admin after the validator rejects the refund promise.
Checklist before you move traffic
- Write 20 real tasks per workflow with correct answers. Owner: process owner. Evidence: a dated task file.
- Write 1 code validator per task. Owner: developer. Evidence: passing validator tests.
- Test every candidate with the same prompt at default and high effort. Owner: developer. Evidence: a log per call.
- Calculate the cost per successful task in rupiah from the log. Owner: finance team. Evidence: a cost sheet.
- Decide which data may go to the DeepSeek API. Owner: business owner. Evidence: a personal data decision note.
- Set up routing, 1 retry at high effort, and escalation. Owner: developer. Evidence: router configuration.
- Schedule DeepSeek batches outside WIB peak hours. Owner: operations team. Evidence: a cron schedule.
- Set price reminders for 23 October and 1 December 2026. Owner: operations team. Evidence: a calendar invite.
- Stop testing when 1 candidate passes 19 of 20 tasks at the lowest cost per successful task. If no candidate passes 18 of 20, use an upper-class model for that task class.
When to pick which model
This table sums up the facts above per model. Each "pick" and "avoid" cell comes from a number we already stated with its source.
| Model | Pick when | Avoid when | Supporting facts |
|---|---|---|---|
| GPT-6 Luna | High volume, code can check the result, context below 272,000 tokens | Arithmetic on default effort, or factual answers without a knowledge base | Cost per task $0.07, AA-LCR 83.3%, hallucination 77% |
| DeepSeek V4.1 Flash | Agents and SaaS flows, off-peak batches, or you need open weights | Personal data through the API, or factual answers without a knowledge base | AutomationBench 68.9%, GDPval 1,600, hallucination 96%, MIT license |
| Gemini 3.8 Flash | Images, audio, video, general knowledge, and speed | A tight budget after 1 January 2027 | Index 41, MMMU-Pro 86%, 276 tokens per second, cost per task $1.24 |
| Claude Haiku 4.5 | Short answers that must be careful, with a knowledge base | Multi-step agents, arithmetic without extended thinking, or context above 200,000 tokens | Hallucination 27%, accuracy 18%, 2.2 minutes per task, AutomationBench 3.2% |
The DeepSeek privacy policy states that DeepSeek processes and stores personal data in the People's Republic of China. If your internal rules forbid that, run the DeepSeek open weights on your own server or use another model for customer data.
Rama Digital recommendation: make Luna the default lane for checkable tasks, Gemini 3.8 Flash the multimodal lane, and Haiku 4.5 the customer answer lane. This holds when you have code validators and a knowledge base. Pick DeepSeek when your task is an agent without personal data, mainly at off-peak hours.
Frequently asked questions
Which model is the cheapest per task? GPT-6 Luna. Artificial Analysis lists $0.07 per index task for Luna, $0.21 for Haiku 4.5, and $0.27 for DeepSeek. Luna also has the lowest token price, $0.10 input and $0.50 output.
Why is Claude Sonnet 5 not in this comparison? Sonnet 5 costs $2 input and $10 output, 2x the Haiku 4.5 price and 20x the Luna output price. Haiku 4.5 is the Claude model with the price closest to the other 3 models, so Sonnet 5 belongs to the price class above.
Is DeepSeek Flash safe for customer data? The DeepSeek privacy policy states that personal data is processed and stored in the People's Republic of China. Check your personal data law and your contracts. The alternative is the DeepSeek open weights on your own server.
Why do cheap models fail invoice math? On default settings, Luna, DeepSeek, and Haiku answer without enough reasoning for multi-step arithmetic. In our test, high effort fixed all 3 to 3 of 3.
Will the Gemini 3.8 Flash price go up? Yes. The Gemini API pricing page lists $0.75 input and $3.75 output through 31 December 2026, then $1.50 and $7.50 from 1 January 2027.
Is Claude Haiku 4.5 still worth using? Yes, for short answers that must be careful and fast. Haiku 4.5 has a 27% hallucination rate, the lowest of the 4 models, and 2.2 minutes per index task. For multi-step agents, pick DeepSeek or Gemini, because Haiku scores only 3.2% on AutomationBench.
Next step
The limit that still applies: independent scores and our test are only a start, because your prompts, data, and risks differ. Test 20 real tasks before you move traffic or sign a budget.
If you want us to map the model lanes for your workflow, open AI Workflow Audit. To talk it through first, pick a 30-minute first consultation slot.
Sources
- OpenAI: GPT-6 Luna model page
- DeepSeek API Docs: Models & Pricing
- DeepSeek: Privacy Policy
- Google: Gemini Developer API pricing
- Google: Gemini 3.8 Flash model page
- Anthropic: Claude pricing
- Anthropic: Claude models overview
- Artificial Analysis: Intelligence Index v4.3.2
- Artificial Analysis: GPT-6 Luna, DeepSeek V4.1 Flash, Gemini 3.8 Flash, Claude 4.5 Haiku
- Bank Indonesia: JISDOR rate




