Quick Answer: Claude Code pricing is a monthly plan per developer, 15% to 20% cheaper billed annually on Pro and Team, plus API-rate usage past the allowance. Enterprise meters every token. For a team, usage usually decides the real Claude Code cost, not the plan, per Anthropic's published costs as of 28 September 2026.
Every paid Claude plan includes Claude Code, but a plan only buys an allowance. Once developers run long agent sessions on large repositories, tokens matter more than the seat.
How much does Claude Code cost per developer?
Per developer, Claude Code costs a seat plus usage. On Anthropic's pricing page, as of 28 September 2026, Pro costs $20 a month billed monthly or $17 billed annually, and Max starts at $100. A Team standard seat costs $25 billed monthly and a premium seat, with five times the usage, $125; both are 20% cheaper billed annually. Enterprise charges the annual standard-seat rate plus all usage at API rates.
Usage is where budgets move. Anthropic's guide to managing Claude Code costs reports, as of 28 September 2026, an enterprise average of about $13 per developer per active day and $150 to $250 per developer per month, with 90% of users under $30 a day. For 10 developers, that's $1,500 to $2,500 a month; for 25, $3,750 to $6,250.
On Pro, Max and Team plans, Claude Code shares one usage pool with Claude chat, reset on a rolling five-hour window. Past it, optional usage credits bill at API rates.
Plans and billing as documented by Anthropic on 28 September 2026; links in the text.
| Plan | Who it's for | Claude Code | How usage is billed | Spend control |
|---|---|---|---|---|
| Pro | One developer, short sessions | Included | Plan limits, then usage credits | Personal monthly credit limit |
| Max 5x or 20x | One developer, all day | Included | 5x or 20x Pro usage per session | Personal monthly credit limit |
| Team standard or premium | Teams of 2 to 150 | Included | Allowance for each seat; premium gets 5x standard | Organization, group and member limits |
| Enterprise | Large organizations | Included | Seat plus usage at API rates | User and organization limits |
| Claude Console (API key) | Pay-as-you-go teams | Included | Every token billed | Workspace spend limits |
What drives Claude Code spend for a team beyond seats?
Beyond seats, tokens drive the bill: which model runs, how much context each turn carries, and how long agents work.
On Anthropic's API pricing table, as of 28 September 2026, Opus 5.5 costs $4 per million input tokens and five times that per million output tokens. Sonnet 5 costs $2 and $10, and Haiku 4.5 costs $1 and $5.
| Driver | Effect on the bill | Control |
|---|---|---|
| Default model | Opus 5.5 costs twice Sonnet 5 per token | Default to Sonnet; keep Opus for hard problems |
| Context size | Every turn resends the files and history in context | Clear between tasks; compact long sessions |
| Agent teams | Each teammate runs its own context window | Keep teams small; shut teammates down when done |
| Heavy users | A few developers can outspend the rest | Member spend limits and a weekly review |
| US-only inference | A 1.1x multiplier on token prices | Use it only where residency rules require it |
| Fast mode | Twice the standard price on Opus 5.5 | Allow it by exception |
Anthropic traces unexpectedly high API spend mostly to uncleared long sessions and Opus left as the default.
How do teams keep Claude Code spend under control?
Teams control Claude Code spend with hard caps, model defaults and per-developer reporting; the controls depend on how each developer signs in.
- Team and Enterprise plans: admins set spend limits at the organization, group or member level. On Team, the seat allowance is the default ceiling, so turn on usage credits only with those limits in place.
- Claude Console: set workspace spend limits, and read per-user spend in the Console dashboard or the Claude Code Analytics API.
- A cloud provider account: caps live in that provider's budget tools, and per-user numbers come from OpenTelemetry export or a gateway.
Set the default model in managed settings so nobody starts on Opus by accident, and teach developers to /clear between unrelated tasks.
If developers reach Claude through several accounts, an LLM gateway gives one place for team keys, budgets and logs. Anthropic documents routing Claude Code through a gateway, but not to non-Claude models.
How does Claude Code pricing compare with other AI coding tools?
Most AI coding tools now price the same way: a plan for each user with an included usage allowance, then metered usage past it. They differ in how the allowance is pooled. Pricing models below are as documented by each vendor on 28 September 2026.
- GitHub Copilot Business and Enterprise: each granted seat adds monthly GitHub AI Credits. Per GitHub's billing docs, credits pool across the billing entity, and extra usage bills per credit only if an admin allows it.
- Cursor Teams: a monthly price for each user. On Cursor's pricing page, every plan includes a set amount of model usage, with on-demand usage billed in arrears.
- OpenAI Codex: included in ChatGPT plans with usage limits; Plus and Pro users can buy extra credits, and anyone can use an API key at standard API rates.
- Claude Code: each seat's allowance is shared with Claude chat, while Enterprise meters all usage at API rates.
Teams that need on-premise AI code assistance have a different shortlist, covered in Cursor alternatives for on-premise AI coding. For the hosted field, see Claude Code alternatives.
When does a self-hosted coding assistant cost less?
A self-hosted coding assistant costs less when usage is heavy and steady, the team is large, and you already run GPU capacity, because owned inference is mostly a fixed cost.
- Usage intensity: a team near Anthropic's average is cheap to serve hosted; a team running long daily agent sessions is not.
- Hardware: existing GPU servers and an operations team change the math.
- Model fit: open-weight models handle completions and routine edits, but hard multi-file changes may need a frontier model. Claude Code itself runs only Claude models.
- Residency: Anthropic charges a 1.1x multiplier for US-only inference, and some policies forbid external inference entirely.
Teams with data residency requirements often compare a self-hosted assistant with Copilot and Claude Code on cost and control. For the Copilot side, see how GitHub Copilot handles data residency. Count self-hosting's hidden costs too: patching, upgrades and on-call time.
How do you estimate Claude Code spend before a rollout?
Start from a two-week pilot, not a list price; Anthropic also advises baselining a small pilot group first. This worksheet turns the pilot into a budget:
- Seats: count developers per tier; give premium seats only where pilot usage justifies them.
- Pilot: run 5 to 10 developers on real work, logging usage per person.
- Baseline: take the median cost per active day.
- Scale: multiply the median by active days per month and by developers.
- Buffer: price your top 10% of users separately at their own pilot rate.
- Modifiers: adjust for model mix, US-only inference and fast mode.
- Caps: set member and organization limits at the forecast, then review them weekly.
Monthly budget equals seats, plus median daily cost times active days times developers, plus the buffer.
What mistakes should you avoid when budgeting for Claude Code?
The costly mistakes hide usage until the invoice:
- Budgeting seats only. Usage past the allowance is where budgets slip.
- No member limits. One developer running agent teams all day can drain an organization cap.
- Opus as the default. It costs twice Sonnet 5 per token, and subagents that inherit the session's model use it too.
- No usage logs. Without OpenTelemetry or a gateway, you can't see who spends what.
- Ignoring data terms. Anthropic lists Pro and Max as opt-out for model training; Team and Enterprise don't train on your content by default.
How Origins AI's Coding Tool approaches cost for on-premise teams
Origins AI (originshq.com) makes the Origins AI Coding Tool, an enterprise self-hosted AI coding assistant and LLM gateway that runs on-premise, in your own AWS, Azure or GCP account, hybrid, or air-gapped. According to its product page, the gateway tracks every LLM request by team, project and engineer, and enforces team token quotas and model restrictions.
The company says the gateway works with Anthropic, OpenAI, open-weight models or your own model, so hosted and local traffic share one budget view. In on-premise and air-gapped modes, no source code leaves your network; in hybrid mode, the code context submitted to a hosted model does.
The company does not publish a rate card. It says it handles implementation, integration and support, and its product page offers a demo that scopes a pilot starting with the gateway. Origins AI reports that the gateway can be live and routing traffic from your IDE plugins within a week.
Talk to an engineer
Comparing Claude Code's bill with running a coding assistant on your own servers? Bring the spend worksheet above and book a call with an engineer.
Written by Apoorva Kumar, Co-Founder & CEO, Origins AI.


