TokenAcornTokenAcorn

Subscribe to model updates

Enter your email and we'll notify you when model prices, rankings, or new releases change.

Live Updates

Model News

Track the latest AI model releases, price changes, and industry updates

Latest1 weeks ago
View Source

xAI Introduces Grok 4.6: Frontier-Tier Reasoning for Long-Running Agents and Complex Coding Workflows

[August 2026]

xAI has officially introduced its newest flagship model, Grok 4.6 (grok-4.6). Engineered with an emphasis on long-running agentic workflows, software development, and multi-step interactive project execution, Grok 4.6 combines multimodal input capabilities with adjustable reasoning controls.

A central element of this release is an aggressive API pricing structure designed to maintain low token costs during repetitive agentic loops, accompanied by a context-tiered pricing framework for extended document and repository analysis.

Grok 4.6 API Pricing (Per Million Tokens)

Model Tier / Context BandPositionInput Price (per 1M Tokens)Cached Input Price (per 1M Tokens)Output Price (per 1M Tokens)
Grok 4.6 (Standard Band)Requests below 200,000 tokens: Flagship reasoning and coding execution$2.00$0.50$6.00
Grok 4.6 (Long-Context Band)Requests at or above 200,000 tokens (up to the 500K context limit)$4.00$1.00$12.00
Grok 4.6 Fast VariantLatency-optimized tier for time-sensitive subagent executions$4.00$1.00$12.00

Key Features and Technical Highlights

  • Long-Horizon Agent Capabilities: Grok 4.6 is fine-tuned through supplemental reinforcement learning runs focused on domain-specific programming, multi-file code editing, and autonomous tool usage. The model is built to sustain multi-turn tasks such as scaffolding applications, researching domain knowledge, and iterating across feedback loops.
  • Balanced Output Pricing: With standard output pricing set at $6.00 per million tokens, Grok 4.6 maintains a 3:1 output-to-input price ratio, helping lower operational costs for agentic tasks that require extensive code generation and verbose reasoning traces.
  • Configurable Reasoning Effort: Developers can regulate thinking depth via four reasoning levels: low, medium, high (default), and xhigh. This allows applications to balance token consumption against task complexity.
  • 500K Context Window & Caching: The model supports up to a 500,000-token context window with text and image input support. Prompt caching is supported via dedicated session headers, reducing cached input rates to $0.50 per million tokens for standard workloads.

Grok 4.6 is currently accessible through the xAI API, Grok Build, and Cursor, as well as deployment platforms including OpenRouter, Vercel, and Cloudflare.

2026年8月14日 03:59
1 weeks ago
View Source

Google Introduces Gemini 3.7 Flash: Intelligent Workhorse Model for Coding and Agents with 50% Introductory API Pricing

[August 2026]

Google has officially released Gemini 3.7 Flash (gemini-3.7-flash), its newest workhorse model optimized for software engineering, multi-step agentic workflows, and complex document processing. Arriving three weeks after Gemini 3.6 Flash, the model incorporates algorithmic reasoning improvements and enhanced tool-calling capabilities designed to minimize failed agent loops and improve instruction adherence.

To facilitate adoption across high-volume production systems, Google has introduced a temporary 50% price reduction across the Gemini API through the end of 2026.

Gemini 3.7 Flash API Pricing (Per Million Tokens)

Model Name / Rate PeriodPositionInput Price (Uncached / 1M)Cached Input Price (per 1M Tokens)Output Price (per 1M Tokens)
Gemini 3.7 Flash (Introductory)Valid through December 31, 2026: Workhorse model for coding, web UI generation, and agentic orchestration$0.75$0.075$3.75
Gemini 3.7 Flash (Standard)Effective January 1, 2027: Standard enterprise and developer rate$1.50$0.15$7.50

Key Features and Technical Highlights

  • Software Engineering & Roadblock Self-Correction: Gemini 3.7 Flash demonstrates notable gains in multi-turn debugging, issue resolution, and schema compliance. The model achieves 65.3% on DeepSWE v1.1 and 43.6% on FrontierCode 1.1 Main, with improved error-recovery mechanisms during autonomous execution.
  • Web Development & UI Fidelity: The model improves front-end layout generation from visual mocks, screenshots, and design systems, scoring a 1588 Elo on Arena.ai's WebDev Arena to deliver higher design adherence and 1:1 visual parity.
  • Document Comprehension & Enterprise Workflows: Gemini 3.7 Flash enhances reasoning over dense multimodal inputs, achieving 34.0% on the GDP.pdf document evaluation and 30.4% on AutomationBench for complex business process execution.
  • Context Capacity & Configurable Thinking: Supporting a 1,048,576-token (1M) context window and up to 65,536 output tokens, the model provides configurable thinking levels (low, medium, and high) to give developers fine-grained control over reasoning latency and token expenditure.

Gemini 3.7 Flash is currently rolling out across the Gemini API via Google AI Studio, Vertex AI, Gemini Enterprise Agent Platform, and as the default model within Google Antigravity.

2026年8月13日 05:12
3 weeks ago
View Source

OpenAI Announces Major API Price Cuts for GPT-5.6 Model Series

SAN FRANCISCO — July 30, 2026 — OpenAI today announced significant price reductions across its recently launched GPT-5.6 model family, cutting API costs by up to 80%. The updated rates stem from major infrastructure efficiency gains and autonomous system optimizations.

Summary of Pricing Changes

  • GPT-5.6 Luna (Lightweight & Cost-Sensitive Workloads): 80% Reduction

    • Input: $0.20 per 1 million tokens (down from $1.00)
    • Output: $1.20 per 1 million tokens (down from $6.00)
  • GPT-5.6 Terra (Balanced Everyday Work): 20% Reduction

    • Input: $2.00 per 1 million tokens (down from $2.50)
    • Output: $12.00 per 1 million tokens (down from $15.00)
  • GPT-5.6 Sol (Flagship Model): Base Pricing Unchanged

    • Input: $5.00 per 1 million tokens
    • Output: $30.00 per 1 million tokens
    • New Feature: OpenAI introduced a new Fast Mode for Sol, offering up to 2.5× faster responses at twice the base API cost, replacing the previous Priority Processing feature.

Availability & Subscription Impact

  • API Users: The price cuts take effect immediately across all standard API endpoints.
  • ChatGPT Work & Codex Subscribers: While subscription plan fees remain the same, tasks running on Terra and Luna will now consume significantly fewer quota credits, enabling users to complete more work under existing monthly allocations.

About OpenAI
OpenAI is an AI research and deployment company dedicated to ensuring that artificial general intelligence benefits all of humanity.


2026年7月31日 03:27
4 weeks ago
View Source

Anthropic Introduces Claude Opus 5: Frontier-Level Intelligence for Autonomous Agents and Advanced Coding

[July 2026]

Anthropic has officially released Claude Opus 5 (claude-opus-5), its newest flagship model engineered for long-horizon agentic workflows, software development, and complex knowledge work. Designed to deliver performance approaching Anthropic's top-tier Claude Fable 5 at half the operating cost, Opus 5 emphasizes token efficiency and precise developer control over reasoning effort.

A key focus of this release is providing a predictable cost structure for enterprise workloads alongside flexible execution modes for latency-sensitive subagent tasks.

Claude Opus 5 API Pricing (Per Million Tokens)

Model NamePositionInput Price (Uncached / 1M)Cache Hit Price (per 1M Tokens)Output Price (per 1M Tokens)
Claude Opus 5Standard Tier: Flagship reasoning model for long-running agentic tasks, coding, and analysis$5.00$0.50$25.00
Claude Opus 5 (Fast Mode)Research Preview: Speed-optimized execution running at ~2.5x default speed$10.00$1.00$50.00

Key Technical Features and Architecture Updates

  • Effort Ladder & Adaptive Thinking: Reasoning ("thinking") is enabled by default. Developers can control model depth using an effort parameter ranging from low to max. This allows workloads to scale computing effort up for critical software architecture tasks or down to conserve tokens on routine executions.
  • Expanded Token Capacity: Opus 5 features a standard 1-million-token context window and supports up to 128,000 output tokens per request, enabling continuous multi-file refactoring and long-form document generation.
  • Dynamic Tool Selection: Introduces beta support for mid-conversation tool changes. Developers can add or remove API tools between interaction turns without invalidating or resending the prompt cache.
  • Lower Prompt Caching Threshold: The minimum threshold for prompt caching on Opus 5 has been reduced from 1,024 to 512 tokens. Cached input tokens receive a 90% discount ($0.50 per 1M tokens), lowering the cost of short subagent prompts and persistent system instructions.

Claude Opus 5 is currently available across the Claude API, Claude Code, and subscription plans (Claude Pro and Claude Max), as well as via Amazon Bedrock, Google Cloud, and Microsoft Foundry.

2026年7月26日 19:40
7月21日
View Source

Google Expands Gemini Family: Introduces Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber with Specialized API Pricing

[July 2026]

Google has officially introduced three new additions to its Gemini model family: the flagship workhorse Gemini 3.6 Flash, the lightweight Gemini 3.5 Flash-Lite, and the domain-specialized Gemini 3.5 Flash Cyber. Designed to scale high-volume agentic workflows, software development, and specialized cybersecurity operations, these new releases focus on token efficiency, execution speed, and cost control across operational scales.

Gemini Flash Series API Pricing (Per Million Tokens)

Model NamePositionInput Price (Uncached / 1M)Input Price (Cached / 1M)Output Price (per 1M Tokens)
Gemini 3.6 FlashMainline workhorse model optimized for coding, knowledge work, and complex agentic tasks$1.50$0.15$7.50
Gemini 3.5 Flash-LiteUltra-fast, low-cost model tuned for high-throughput execution, subagent fan-out, and document processing$0.30$0.03$2.50
Gemini 3.5 Flash CyberSpecialized cybersecurity model fine-tuned for vulnerability discovery and patch generationEnterprise / Partner Access (Integrated in CodeMender)N/AEnterprise / Partner Access

Key Features and Operational Highlights

  • Token Efficiency & Reduced Output Costs: Gemini 3.6 Flash delivers notable efficiency improvements, using approximately 17% fewer output tokens on standard evaluation sets and requiring fewer reasoning steps for multi-turn tasks compared to Gemini 3.5 Flash. Output pricing has been reduced from $9.00 to $7.50 per million tokens.
  • High-Throughput Performance: Gemini 3.5 Flash-Lite provides the lowest latency and cost profile in the 3.5 generation, reaching speeds of approximately 350 output tokens per second for high-volume agentic search and rapid execution tasks.
  • Specialized Cybersecurity Workflows: Gemini 3.5 Flash Cyber is designed specifically for security operations. Paired with Google’s CodeMender agent platform, it focuses on identifying, validating, and generating remediation patches for software vulnerabilities.
  • Context Caching Structure: Both 3.6 Flash and 3.5 Flash-Lite offer a 90% discount on cached input tokens ($0.15/1M and $0.03/1M respectively), reducing operational costs for applications requiring long-context maintenance or repeated prompt templates.

The Gemini 3.6 Flash and Gemini 3.5 Flash-Lite models are now generally available through Google AI Studio, Vertex AI, and partner developer integrations.

2026年7月21日 19:44
7月16日
View Source

Moonshot AI Introduces Kimi K3: A 2.8-Trillion Parameter Open-Weight Reasoning Model with Context-Aware API Pricing

[July 2026]

Moonshot AI has officially introduced its next-generation flagship model, Kimi K3 . Built with 2.8 trillion total parameters, Kimi K3 is an open-weight, native multimodal reasoning model designed to tackle complex agentic workflows, advanced software engineering, and deep knowledge analysis .

A central aspect of this release is the deployment of Kimi K3 on the Moonshot API Platform with an optimized pricing structure that heavily rewards context reusability, helping developers manage costs for long-form, multi-turn reasoning workflows .

Kimi K3 API Pricing (Per Million Tokens)

Model NamePositionInput Price (Cache Miss)Input Price (Cache Hit)Output Price
Kimi K3Flagship 2.8T multimodal reasoning model with "always-on" thinking mode$3.00$0.30$15.00

Key Pricing and Technical Features

  • Context Caching Discount: To accommodate dense reasoning and long-context applications, Kimi K3 integrates a Context Caching mechanism . When an input query hits the cache, the cost drops by 90% to $0.30 per million tokens, allowing cost-effective deployment of long-context agents and multi-turn conversational systems [2].
  • Architectural Milestones: Kimi K3 introduces Kimi Delta Attention (KDA)—a hybrid linear attention mechanism—alongside Attention Residuals . These architectural developments allow the model to manage up to a 1-million-token context window with stable compute scaling and controlled latency .
  • Efficient Mixture-of-Experts: Utilizing a Stable LatentMoE architecture, the model scales to 896 total experts but activates only 16 experts per token . This design allows Kimi K3 to offer deep, high-capacity reasoning capabilities while keeping the active computational load relatively lightweight during inference .
  • Open-Weight Commitment: Following its initial developer preview, Moonshot AI plans to open-source the weights of Kimi K3 on July 27, 2026, allowing the broader research community to study and build upon its multimodal reasoning foundation .

The Kimi K3 model is currently available in developer preview through the Kimi API Platform and is rolling out to selected developer integrations .

2026年7月16日 22:55
7月9日
View Source

OpenAI Introduces GPT-5.6 Model Series: Sol, Terra, and Luna with Tiered Pricing Structure

[July 2026]

OpenAI has officially introduced its new generation of models, the GPT-5.6 series. Designed to accommodate different use cases and budgetary needs, the series consists of three distinct model tiers: the flagship GPT-5.6 Sol, the balanced GPT-5.6 Terra, and the lightweight GPT-5.6 Luna.

A central aspect of this release is the introduction of a structured, tiered API pricing scheme aimed at managing developer costs across different operational scales.

GPT-5.6 Series API Pricing (Per Million Tokens)

Model NamePositionInput Price (per 1M Tokens)Output Price (per 1M Tokens)
GPT-5.6 SolFlagship model optimized for advanced reasoning and complex tasks$5.00$30.00
GPT-5.6 TerraBalanced model offering a compromise between capability and cost$2.50$15.00
GPT-5.6 LunaLightweight model optimized for fast responses and budget efficiency$1.00$6.00

Key Pricing Features

  • Fixed Cost Ratio: For all three models in the series, the price of output tokens is set at exactly six times the price of input tokens. This uniform ratio is designed to help developers more easily project API costs for long-form generation and agentic workflows.
  • Cost Efficiency Tiers: The mid-tier Terra model represents a 50% cost reduction compared to the flagship Sol model. Meanwhile, the lightweight Luna model operates at one-fifth of the cost of Sol, providing a lower-cost option for high-frequency, lower-complexity tasks.

The GPT-5.6 series models are currently rolling out in developer preview through the OpenAI API and selected integrations, such as GitHub Copilot.

2026年7月9日 19:07
6月16日
View Source

Claude Sonnet 5 Released

Anthropic has released Claude Sonnet 5, designed to bring advanced agentic capabilities to the fast and cost-efficient Sonnet model line. Here is a summary of its key features:

  • Autonomous Agentic Capabilities: Built to act as a more capable agent, Sonnet 5 can formulate plans, interact with developer tools (such as browsers and terminals), and run autonomously over longer durations than previous Sonnet models.
  • Performance and Cost Efficiency: It narrows the gap between the mid-tier Sonnet class and flagship-class models. It achieves performance close to Claude Opus 4.8, while retaining the lower pricing structure and faster processing speeds of the Sonnet family.
  • Step Up from Sonnet 4.6: The model introduces noticeable improvements over its predecessor, Sonnet 4.6, in core areas including logical reasoning, multi-file codebase navigation, tool usage, and general knowledge work.
  • Safety and Risk Mitigation: System card evaluations indicate a lower overall rate of undesirable behaviors compared to Sonnet 4.6. To align with safety protocols, its cybersecurity execution capabilities are intentionally restricted relative to top-tier models (like Mythos 5), reducing potential deployment risks.
  • New Tokenizer: It utilizes an upgraded tokenizer that generates roughly 30% more tokens for the same text, improving processing efficiency and context handling across various tasks.
2026年6月16日 18:00

News content is curated and published by administrators from official vendor channels.