Claude Sonnet 4.6
Claude Sonnet 4.6 is a verified AI model profile covering official specifications, pricing or access, capabilities, practical use…
GPT-5.6 Luna is OpenAI's cost-optimized GPT-5.6 model with 1.05M context, 128K output, tool use and pricing from $0.20/M input.
GPT-5.6 Luna is OpenAI's lowest-cost GPT-5.6 model for high-volume automation, extraction, classification, lightweight agents and production workloads.
GPT-5.6 Luna is the lowest-cost member of OpenAI’s GPT-5.6 family, designed for workloads where volume, latency and predictable API spend matter more than extracting the maximum possible reasoning performance from every request.
OpenAI positions Luna as the cost-sensitive tier beneath GPT-5.6 Terra and Sol. That positioning matters: Luna is not simply an older or stripped-down model. It keeps the modern GPT-5.6 API surface—including reasoning controls, vision input and the Responses API tool ecosystem—while targeting high-throughput tasks such as classification, extraction, routing, summarization, support automation and sub-agent work. For teams building large pipelines, this means the same application architecture can often route routine requests to Luna and escalate difficult cases to Terra, Sol or Astra.
The model has a 1.05-million-token context window and can generate up to 128,000 output tokens. It accepts text and image input and produces text. Reasoning effort can be configured from none through several higher levels, allowing developers to trade latency and token use for stronger reasoning on individual requests. Function calling and structured outputs are supported, and the Responses API can connect Luna to web search, file search, code interpreter, hosted shell, computer use, MCP and other OpenAI tools. Its reliable knowledge cutoff is February 16, 2026, so web-connected tools are still important for fresh facts.
In real systems, Luna is particularly attractive as the inexpensive worker in a routed architecture. A support system can let Luna classify intent, retrieve files and draft routine answers while escalating ambiguous cases. A document pipeline can use Luna to extract fields from thousands of pages and return validated JSON. An agentic application can assign routine sub-tasks to Luna while reserving expensive frontier models for planning or final review. Because it shares the modern OpenAI tool interface, teams do not need a separate integration just to use the cheaper tier.
GPT-5.6 Luna is active on the OpenAI API. It is not a fine-tuning model, but it supports streaming, functions and structured outputs. It is intended for production rather than a temporary preview. OpenAI also lists very high rate-limit ceilings at upper usage tiers, which aligns with the model’s high-volume positioning.
Standard pricing is $0.20 per million input tokens, $0.02 per million cached input tokens and $1.20 per million output tokens. That price can be dramatically lower than frontier tiers when a workflow processes large amounts of text. However, OpenAI applies higher rates when the prompt exceeds 272K input tokens, so teams using the full 1.05M context should model long-context costs separately instead of assuming the base price applies to every request.
Choose Luna when the task is repetitive, structured, high-volume or easy to verify: extraction, classification, transformation, support triage, metadata generation, simple coding assistance and agent sub-tasks are good examples. Choose Terra when mistakes become more expensive, Sol for complex professional reasoning and Astra for the hardest end-to-end work. A mixed-model router will usually produce a better cost-quality balance than using one premium model for everything.
Luna’s main trade-off is intelligence per request. It can use tools and reasoning, but difficult coding, ambiguous planning and high-stakes analysis can expose capability gaps relative to larger GPT-5.6 tiers. Its knowledge is also static without external tools. Large contexts can increase both cost and failure surface, and structured outputs or tool calls should still be validated before downstream systems act on them.
Classify these support tickets into our approved categories and return JSON.Extract the requested fields from these documents and validate missing values.Use the provided tools to resolve this routine customer-service workflow.$0.20 per 1M input tokens, $0.02 per 1M cached input tokens and $1.20 per 1M output tokens. Requests above 272K input tokens use higher long-context rates.
The most useful GPT-5.6 model for scale: use Luna when throughput and cost matter more than maximum frontier reasoning.