Anthropic's New AI Model Targets Coding, Enterprise Work

Anthropic has released Claude Opus 4.6, introducing a million-token context window and automated agent coordination features as the AI company seeks to expand beyond software development into broader enterprise applications.

The San Francisco-based firm said the model improves performance on coding tasks, financial analysis, and document processing compared to its predecessor. Anthropic positioned the release as strengthening its position in enterprise AI workflows, an increasingly crowded market where it competes directly with OpenAI and Google.

"We're focused on building the most capable, reliable, and safe AI systems," an Anthropic spokesperson said. "Opus 4.6 is even better at planning, helping solve the most complex coding tasks."

The release comes three days after OpenAI launched a desktop application for its Codex AI coding system, underscoring the rapid pace of competition in AI development tools. Anthropic said in November that Claude Code, its coding product, reached $1 billion in annualized revenue six months after general availability.

Extended Context and Agent Coordination

Opus 4.6 supports up to one million tokens of context in beta on Anthropic's developer platform, a substantial increase from the 200,000-token limit of earlier Opus versions. The expansion allows the model to process larger codebases and longer documents without splitting tasks across multiple requests.

The company also introduced agent teams in Claude Code as a research preview, allowing multiple AI agents to work simultaneously on segmented portions of a project. Scott White, Anthropic's head of product, compared the feature to coordinating a human team working in parallel.

Anthropic said Opus 4.6 addresses context degradation, a common problem where AI performance declines as conversations lengthen. On a retrieval benchmark that hides information in large text volumes, Opus 4.6 scored 76% compared to 18.5% for its Sonnet 4.5 model.

The model supports outputs of up to 128,000 tokens. Anthropic introduced adaptive thinking, which allows the model to determine when to apply deeper reasoning, and four effort settings that developers can adjust to balance performance, speed, and cost.

Benchmark Performance

Anthropic reported that Opus 4.6 leads on Terminal-Bench 2.0, an evaluation of AI agents completing command-line tasks, with a 65.4% score under maximum-effort settings. The Terminal-Bench project's public leaderboard shows separate entries for Opus 4.6, with a score of 62.9% under one configuration.

On GDPval-AA, a benchmark measuring performance on professional tasks across finance, legal, and other domains, Anthropic said Opus 4.6 outperforms OpenAI's GPT-5.2 by approximately 144 Elo points, a gap that corresponds to a roughly 70% win rate in direct comparisons. Artificial Analysis, which maintains the GDPval-AA leaderboard, describes the evaluation framework in its methodology documentation.

Anthropic also cited results from BrowseComp, an OpenAI benchmark for browsing agents that measures the ability to locate hard-to-find information across 1,266 questions that require persistent web navigation.

Safety Testing and Cybersecurity Measures

Anthropic said Opus 4.6 underwent extensive safety evaluations, including tests for deception, sycophancy, and cooperation with potential misuse. The company's system card reports the model showed low rates of problematic behaviors while achieving the lowest rate of over-refusals among recent Claude models.

The company developed six cybersecurity probes to detect harmful uses of the model's enhanced capabilities. Anthropic said it is using Opus 4.6 to identify and patch vulnerabilities in open-source software as part of defensive cybersecurity efforts.

"Agents have tremendous potential for positive impacts in work, but it's important that agents continue to be safe, reliable, and trustworthy," the spokesperson said, referring to a framework Anthropic published outlining core principles for agent development.

Product Integrations and Pricing

Anthropic released Claude in PowerPoint as a research preview for paid subscribers, building on existing integrations with Excel. The PowerPoint tool reads layouts, fonts, and slide templates to generate presentations, the company said.

White said Anthropic has observed the use of Claude Code expanding beyond software engineers to product managers, financial analysts, and workers in other fields. The company cited deployments at Uber, Salesforce, Accenture, Spotify, and other enterprises.

Opus 4.6 is available on claude.ai and through the Claude API under the identifier claude-opus-4-6. Pricing remains $5 per million input tokens and $25 per million output tokens. Premium pricing of $10 per million input tokens and $37.50 per million output tokens applies when prompts exceed 200,000 tokens using the million-token context window. The model is also available through Amazon Bedrock and Google Cloud Vertex AI.

The release arrives as OpenAI's GPT-5.3-Codex began rolling out through GitHub Copilot, according to GitHub's changelog. GitHub described GPT-5.3-Codex as OpenAI's latest agentic coding model and outlined availability for Copilot Pro, Business, and Enterprise users.

For more information, visit the Anthropic site.

About the Author

John K. Waters is the editor in chief of a number of Converge360.com sites, with a focus on high-end development, AI and future tech. He's been writing about cutting-edge technologies and culture of Silicon Valley for more than two decades, and he's written more than a dozen books. He also co-scripted the documentary film Silicon Valley: A 100 Year Renaissance, which aired on PBS.  He can be reached at [email protected].

Featured

  • lock symbol with quantum bits in dynamic motion

    Microsoft Accelerates Focus on Quantum-Safe Security

    Microsoft is speeding up its quantum-safe security timeline, saying advances in quantum computing and new federal requirements have pushed post-quantum cryptography from a future planning issue into an immediate engineering priority.

  • robot hand holding stacks of coins

    Designing AI Systems for Financial Aid

    Financial aid offices have been slow to adopt AI, risking technological stagnation at a critical early student touchpoint. Systematic AI integration can improve student experiences and strengthen institutional positioning.

  • large cloud icon with abstract code  and interconnected polygons

    Research: Enterprise AI Workloads Are Tipping Toward Private Cloud

    Broadcom's 2026 private cloud report says enterprise AI is moving from experimentation into production, with private cloud emerging as the preferred deployment environment for AI inference among surveyed organizations.

  • woman surrounded by virtual hologram icons

    Beyond AI Adoption: Designing Learning for an Age of Abundant Intelligence

    Higher education was designed for a world in which access to knowledge, expertise, feedback, mentorship, and authentic learning experiences were inherently scarce. By making many forms of intelligence increasingly abundant, AI is inherently redefining the existing paradigm.