Gemini 4 Argon: Million-Token Output, API Pricing and Limited Access

Google announced Gemini 4 Argon on September 30 for long software engineering tasks, professional work and cybersecurity defense. Its million-token output limit stands out, but access remains limited. For developers, the practical questions are whether it can finish complex work and what that work will cost.
Checked October 1, 2026, UTC+8. This is an analysis of official information, not a hands-on Argon review. Google announcement
Access comes before integration
The initial rollout serves trusted defenders through Fairwind. Paid API customers and Google AI Ultra subscribers are planned next; buying Ultra today does not establish immediate access. No broad release date is given. Fairwind access rules
At the time of checking, the Gemini API model directory did not list a public Argon model ID. Account eligibility, regions, input limits and rate limits still need confirmation before integration. A usable API example would require those details, so none is supplied here.
A larger output budget needs checkpoints
The advertised limit rises from 64K to 1M output tokens. It is not an input context-window specification or a promise of a million tokens of visible answer. Output-limit announcement
More generation room could help a task continue through changes across files, debugging or migrations. It does not establish that the resulting changes are correct. A useful workflow still needs version control, tests and review.
Break a long assignment into inspectable stages: a proposed plan, a small batch of changes, then test results and unresolved issues. Judge the model by verified deliverables, not how much text it produces. Checkpoints also make failures easier to diagnose and reduce the work lost when a task must stop.
Coding results vary by benchmark
These are Google-published scores, not tests performed for this article. Official comparison table
| Benchmark | Gemini 4 Argon | GPT-6 Astra | Claude Opus 5.5 |
|---|---|---|---|
| DeepSWE v1.1 | 77.9% | 74.1% | 74.2% |
| FrontierSWE v2 | 55.0% | 65.5% | 62.3% |
| Terminal-bench 4.0 | 57.4% | 58.2% | 66.4% |
Argon leads these models on DeepSWE, while Astra leads on FrontierSWE and Opus on Terminal-bench. The same page reports 51.3% on AutomationBench.
These tests use different tasks and completion criteria. Their percentages cannot be merged into a single success rate or transferred directly to a production repository. The full methodology document was not verified here, so this article makes no assumptions about equal harnesses, retry counts or reasoning budgets.
For your own evaluation, keep the work fixed: a known bug, a change spanning modules and a document that needs factual verification. Compare elapsed time, total cost, test outcomes and human corrections. This connects model selection to the work you actually need done. See the site’s guide to model leaderboards and benchmark metrics.
Budget for both pricing periods
Prices below are US dollars per million tokens. Announcement and pricing footnote
| Usage | Introductory price | After introduction |
|---|---|---|
| Input | US$2 | US$4 |
| Output | US$10 | US$20 |
Cached input has a stated 95% discount. Applying it to the introductory input price gives US$0.10 per million, an arithmetic derivation. No introductory end date is specified.
A hypothetical task using one million ordinary input tokens and 100,000 billable output tokens would cost about US$3 at introductory rates or US$6 afterward. This is arithmetic, not a measured invoice. It excludes additional tool or storage fees; the billing classification of reasoning tokens still needs confirmation from API documentation.
For agents, count usage across the whole task. Repeated background information, tool results and retries can add up even when each individual request looks small. Cache hits and an early stopping policy matter to the final bill.
What developers can prepare now
Before access arrives, three preparations are useful:
- Save a baseline. Use bug fixes, refactoring and documentation tasks with explicit acceptance criteria to record current-model performance.
- Set task limits. Define cost, time and retry budgets before permitting extended execution.
- Make permissions explicit. Run changes and tests in isolated environments, with clear approval points for production deployment, deletion and external actions.
The larger output allowance makes Argon worth following. Access details and project-level results still determine whether it fits a workflow. Once eligible, start with a small, testable assignment before expanding its responsibilities.